Video style conversion method and device, electronic equipment and storage medium
By performing frame decomposition, key element identification, and style conversion on the video, combined with audio synchronization and adjustment of image differences, the problem of unsatisfactory animation style conversion effects was solved, achieving high-precision and smooth animation video generation.
Patent Information
- Application Number
- CN202511379442.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-12-30
AI Technical Summary
Existing anime style conversion technologies struggle to capture the differences in color expression, line drawing, and lighting effects among various anime styles, resulting in unsatisfactory conversion effects and a poor user experience.
The video to be converted is decomposed into consecutive frame images. Key elements are obtained through content recognition. Style conversion is performed based on the style features of the target anime style to generate initial anime images. These images are then combined with dubbing audio to form an initial anime video. The video is adjusted according to the differences between multiple frames to generate the target anime video, avoiding jumps between image frames and style inconsistencies.
It significantly improves the precision and smoothness of animation style transitions, meets users' personalized needs, and ensures overall stylistic harmony and visual coherence in videos.
Smart Images

Figure CN121235898A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of animation image processing technology, and in particular to video style conversion methods, apparatus, electronic devices and storage media. Background Technology
[0002] With the rapid development of digital entertainment and animation culture, users' demand for personalized and diversified visual experiences is increasing. Especially with the widespread dissemination of short videos, film and television clips, and personally created videos, users hope to convert their own videos or favorite film and television clips into anime style to obtain a richer viewing experience.
[0003] Currently, related anime style transfer technologies typically perform overall style transfer on videos directly, generally only supporting the transfer of a single anime style and lacking the ability to perform fine-grained processing on videos.
[0004] However, due to the significant differences in color expression, line drawing, and lighting effects among different animation styles, the overall style transfer process often fails to fully capture these differences, resulting in an unsatisfactory style transition effect and a poor user experience. Summary of the Invention
[0005] To address or partially address the problems existing in related technologies, this application provides a video style conversion method, apparatus, electronic device, and storage medium, which can improve the precision of animation style conversion and meet users' personalized needs.
[0006] The first aspect of this application provides a method for style transfer of a video, comprising: Obtain the video to be converted, the dubbing audio, and the target anime style corresponding to the video to be converted; The video to be converted is decomposed into consecutive frame images, each frame image including key elements obtained through content recognition; Based on the target style characteristics of the target anime style, the key elements are style-transformed to generate several initial anime images; The dubbing audio is combined with the several frames of initial animation images to generate an initial animation video with synchronized audio and video. Based on the image differences between multiple initial animation images, the initial animation video is adjusted to generate the target animation video.
[0007] In one instance, the key elements include character elements and environmental elements. The process involves style conversion of the key elements based on the target style characteristics of the target anime style to generate several initial anime images, including: Extract the spatial structure between the character elements and the environment elements; Obtain the target style features corresponding to the target anime style from the preset anime style library; Analyze the target style features to obtain character style features and environment style features; Keeping the spatial structure unchanged, the character style features are mapped to the character elements, and the environment style features are mapped to the environment elements to generate several initial animation images.
[0008] In one instance, the image differences include inter-frame deviations, and adjusting the initial anime video based on the image differences between multiple initial anime images to generate the target anime video includes: Adjacent initial anime images are used as the first anime image and the second anime image, and the second anime image is the next frame image of the first anime image; Calculate the inter-frame deviation between the first animation image and the second animation image, and locate the jump region in the second animation image where the inter-frame deviation exists; The transition region is fine-tuned to generate a fine-tuned target animation image, and this process continues until the initial animation image has been traversed to generate the target animation video.
[0009] In one example, fine-tuning the transition region to generate the fine-tuned target animation image includes: The first animation image and the second animation image are superimposed to obtain an intermediate frame; The intermediate frames are filled into the transition regions of the second animation image to generate the fine-tuned target animation image.
[0010] In one example, fine-tuning the transition region to generate the fine-tuned target animation image includes: Calculate the forward optical flow and backward optical flow of the first animation image and the second animation image; wherein, the forward optical flow is used to characterize the pixel motion vector from the first animation image to the second animation image, and the backward optical flow is used to characterize the pixel motion vector from the second animation image to the first animation image; Based on the forward optical flow and the backward optical flow, pixel remapping is performed on the first animation image and the second animation image to generate intermediate frames; The intermediate frames are filled into the transition regions of the second animation image to generate the fine-tuned target animation image.
[0011] In one instance, the image differences include style deviations, and adjusting the initial anime video based on the image differences of the initial anime image to generate the target anime video includes: For the aforementioned initial animation images, the style deviation of each initial animation image is calculated, and the corresponding initial animation image is used as the second animation image, and the previous frame of the second animation image is used as the first animation image. Adjust the second animation image to generate the adjusted target animation image, until the initial animation image has been traversed, and generate the target animation video.
[0012] In one instance, adjusting the second anime image to generate the adjusted target anime image includes: Based on the style vector of the first anime image, the second anime image is re-rendered to generate the adjusted target anime image.
[0013] A second aspect of this application provides a video style conversion device, comprising: The conversion data acquisition module is used to acquire the video to be converted, the dubbing audio, and the target anime style corresponding to the video to be converted; The video decomposition module is used to decompose the video to be converted into consecutive frame images, each frame image including key elements obtained through content recognition; The style conversion module is used to perform style conversion on the key elements based on the target style features of the target animation style, and generate several frames of initial animation images; The video synthesis module is used to combine the dubbing audio with the several frames of initial animation images to generate an initial animation video with synchronized audio and video. The video adjustment module is used to adjust the initial animation video based on the image differences between multiple initial animation images to generate the target animation video.
[0014] A third aspect of this application provides an electronic device, comprising: Processor; and A memory that stores executable code, which, when executed by the processor, causes the processor to perform the method described above.
[0015] A fourth aspect of this application provides a computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the method described above.
[0016] The fifth aspect of this application provides a computer program product comprising computer instructions that, when executed by a processor, implement the method described above.
[0017] The technical solution provided in this application may include the following beneficial results: In this application, the following steps are taken: First, the video to be converted, the dubbing audio, and the target anime style corresponding to the video to be converted are obtained. Then, the video to be converted is decomposed into consecutive frame images, each frame including key elements obtained through content recognition. Based on the target style features of the target anime style, style conversion is performed on the key elements to generate several initial anime images. The dubbing audio is combined with these initial anime images to generate an initial anime video with synchronized audio and video. Finally, based on the image differences between the multiple initial anime images, the initial anime video is adjusted to generate the target anime video.
[0018] Compared with related technologies, the technical solution of this application has the following advantages: On the one hand, the video to be converted is decomposed into consecutive frame images, and key elements obtained through content recognition are used to ensure accurate style conversion of key elements based on the target style characteristics of the target animation style. On the other hand, by combining the dubbing audio with several frames of initial animation images after conversion, an initial animation video with synchronized audio and video is generated. When image differences are detected between multiple frames of initial animation images, the initial animation video is further adjusted using the image differences to generate a target animation video with a coordinated overall style. This effectively avoids problems such as jumps, flickering, or style inconsistencies that may exist between adjacent image frames, significantly improves the precision and smoothness of animation style conversion, and meets the personalized needs of users.
[0019] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0020] The above and other objects, features and advantages of this application will become more apparent from the more detailed description of exemplary embodiments thereof in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments thereof.
[0021] Figure 1 This is a schematic flowchart illustrating a video style conversion method according to an embodiment of this application; Figure 2 This is another schematic flowchart illustrating a video style transfer method according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a video style conversion device according to an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device shown in an embodiment of this application. Detailed Implementation
[0022] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to make this application more thorough and complete, and to fully convey the scope of this application to those skilled in the art.
[0023] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0024] It should be understood that although the terms "first," "second," "third," etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0025] With the rapid development of digital entertainment and animation culture, users' demand for personalized and diversified visual experiences is increasing. Especially with the widespread dissemination of short videos, film and television clips, and personally created videos, users hope to convert their own videos or favorite film and television clips into anime style to obtain a richer viewing experience.
[0026] Currently, related anime style transfer technologies typically perform overall style transfer on videos directly, generally only supporting the transfer of a single anime style and lacking the ability to perform fine-grained processing on videos.
[0027] However, due to the significant differences in color expression, line drawing, and lighting effects among different animation styles, the overall style transfer process often fails to fully capture these differences, resulting in an unsatisfactory style transition effect and a poor user experience.
[0028] Among the related technologies, there are problems such as unsatisfactory animation style conversion effects and poor user experience.
[0029] To address the aforementioned issues, this application provides a video style conversion method that can improve the precision of animation style conversion and meet users' personalized needs.
[0030] The technical solutions of the embodiments of this application are described in detail below with reference to the accompanying drawings.
[0031] See Figure 1 , Figure 1 This is a schematic flowchart illustrating a video style transfer method according to an embodiment of this application.
[0032] Step 101: Obtain the video to be converted, the dubbing audio, and the target anime style corresponding to the video to be converted.
[0033] In this embodiment of the application, the data content obtained as input includes the video to be converted, the dubbing audio corresponding to the video to be converted, and the target anime style for the style features to be converted.
[0034] Optionally, the video to be converted refers to the video file uploaded by the user. After obtaining the video to be converted, its validity can be checked. For example, it can check whether the video file can be decoded normally, whether the video format conforms to the preset format type, whether the video duration and video resolution are within the preset range, etc. If any abnormalities are found, a prompt will be issued to the user, prompting the user to re-upload a valid video file.
[0035] Dubbing audio refers to audio files that users actively upload, or audio files obtained from existing audio resource libraries or third-party API resources. Dubbing audio typically corresponds to the video to be converted, used to achieve audio-visual synchronization. For example, dubbing audio can be background music or narration describing the video to be converted.
[0036] The target anime style refers to the anime style information selected from the preset anime style library, or it can be user-defined. Each anime style has unique stylistic characteristics, with differences in color tone, line thickness, texture style, and rendering method. For example, anime style A is characterized by bright and saturated colors and simple, smooth lines, while anime style B emphasizes soft colors and delicate lines.
[0037] As an example, in the style conversion interface, if a user selects anime style A, the style features corresponding to anime style A will be used for style conversion processing. If the user does not directly select an anime style from the interface, but instead enters a natural language description such as "I want the video to present a warm and healing atmosphere, with soft and delicate visuals," then the system can parse the user's input description and match anime style B that meets the user's needs from the preset anime style library. After the user confirms, the style features corresponding to anime style B will be used for style conversion processing. If the user neither selects a target anime style nor enters a text description in the style conversion interface, then by analyzing the rhythmic and emotional features of the dubbing audio, the system can automatically match the most suitable anime style C from the preset anime style library and use the style features corresponding to anime style C for style conversion processing, without user intervention.
[0038] It is worth noting that this application can process multiple target anime styles simultaneously. Taking the user's selection of two target anime styles as an example, anime style A corresponds to the first X frames of the image, and anime style B corresponds to the first Y frames of the image. During the style conversion process, according to the characteristics of each target style, the key elements of their respective frame segments are style converted respectively, thereby generating images with the characteristics of anime style A in the first X frames and the characteristics of anime style B in the last Y frames, so that the same video to be converted can present different anime styles at different time periods.
[0039] Step 102: Decompose the video to be converted into consecutive frame images, each frame image including key elements obtained through content recognition.
[0040] In this embodiment, the input video to be converted is decomposed into consecutive frame images in chronological order, and the key elements of each frame image are obtained through content recognition, thereby obtaining an image that corresponds one-to-one with the time information of the video to be converted, ensuring that the video can be accurately restored during subsequent synthesis.
[0041] Optionally, each frame of the image will be accompanied by metadata information, including frame number, timestamp, resolution, etc. After decomposing into consecutive frame images, image enhancement operations can be performed on each frame, such as adjusting brightness, enhancing contrast, and denoising, to improve the quality of each frame.
[0042] Content recognition refers to the process of analyzing the features of each frame of an image to automatically identify and extract key elements. These elements include people, the environment, and spatial structures within the image. These key elements ensure that the converted consecutive frames retain the semantics and spatial structure of the video being converted.
[0043] As an example, content recognition can be achieved in ways including, but not limited to, object detection and semantic segmentation, to separate human and environmental elements in each frame of an image.
[0044] Step 103: Based on the target style characteristics of the target animation style, perform style conversion on key elements to generate several initial animation images.
[0045] In this embodiment, based on the target style features of the target anime style, style conversion is performed on each key element in each frame of the image to generate an initial anime image that preserves the semantics and spatial structure of the video to be converted.
[0046] Optionally, style transfer refers to converting key elements into features that match the target style. For example, converting the lines of character elements into line art corresponding to the target anime style, mapping the colors of environmental elements to the color tone of the target anime style, and adding texture effects to the overall image. The final output of the initial anime image is consistent with the target anime style.
[0047] The initial anime image refers to the intermediate result image obtained after style conversion. It contains the semantic content and temporal information of the video to be converted, and presents the visual effect of the target anime style to meet the user's personalized preferences and aesthetic needs.
[0048] Step 104: Combine the dubbing audio with several frames of initial animation images to generate an initial animation video with synchronized audio and video.
[0049] In this embodiment, the dubbing audio is combined with several frames of initial animation images according to the time information to ensure that the dubbing audio and the initial animation images are synchronized to obtain the initial animation video.
[0050] Optionally, the initial animation video refers to the video data formed by synthesizing the initial animation images obtained through frame-by-frame style conversion with the dubbing audio.
[0051] As an example, for longer videos to be converted, to reduce computational load, keyframes can be selected first. For instance, one frame can be selected every five frames as a candidate keyframe. If the average absolute error of adjacent frames exceeds a set threshold, that frame is marked as a keyframe. After selection, style conversion or frame interpolation is performed on the keyframes, and then style optimization is performed on the non-keyframes between the keyframes through interpolation. This ensures the overall continuity and style consistency of the video while significantly reducing computational costs.
[0052] Step 105: Adjust the initial animation video based on the image differences between multiple initial animation images to generate the target animation video.
[0053] In this embodiment of the application, after obtaining the initial animation video, the target animation video can be generated by calculating the image differences between multiple frames of the initial animation images and adjusting the initial animation video according to the image differences.
[0054] Optionally, image differences refer to the changes in pixel values, colors, lines, or style vectors between adjacent or non-adjacent frames, used to determine problems such as inter-frame jumps or local style inconsistencies.
[0055] A target animation video refers to an animation video that has been adjusted based on the differences in multiple frames of images. The overall style of the target animation video is consistent and the visuals are coherent.
[0056] In this embodiment, the process involves acquiring the video to be converted, the dubbing audio, and the target anime style corresponding to the video; decomposing the video to be converted into consecutive frame images, each frame including key elements obtained through content recognition; performing style conversion on the key elements based on the target style features of the target anime style to generate several initial anime images; combining the dubbing audio with several initial anime images to generate an initial anime video with synchronized audio and video; and adjusting the initial anime video according to the image differences between multiple initial anime images to generate the target anime video.
[0057] Compared with related technologies, the technical solution of this application has the following advantages: On the one hand, the video to be converted is decomposed into consecutive frame images, and key elements obtained through content recognition are used to ensure accurate style conversion of key elements based on the target style characteristics of the target animation style. On the other hand, by combining the dubbing audio with several frames of initial animation images after conversion, an initial animation video with synchronized audio and video is generated. When image differences are detected between multiple frames of initial animation images, the initial animation video is further adjusted using the image differences to generate a target animation video with a coordinated overall style. This effectively avoids problems such as jumps, flickering, or style inconsistencies that may exist between adjacent image frames, significantly improves the precision and smoothness of animation style conversion, and meets the personalized needs of users.
[0058] Figure 2 This is another schematic diagram of a video style transfer method shown in an embodiment of this application. Figure 2 relatively Figure 1 The technical solutions of the embodiments of this application are described in more detail.
[0059] Step 201: Obtain the video to be converted, the dubbing audio, and the target anime style corresponding to the video to be converted.
[0060] In this embodiment, users can upload a video to be converted through the style conversion interface, and select the dubbing audio to be played synchronously and at least one target anime style to be converted.
[0061] In this application, the process first checks whether the video to be converted can be decoded normally, whether the video format is a preset format type, and whether the video duration and resolution are within a preset range. If all the checks meet the requirements, the video to be converted can be marked as qualified input data. Similarly, the audio format and content of the dubbing audio can be checked. If the dubbing audio is selected from an existing audio material library, the check step can be skipped.
[0062] Users can choose one or more anime styles as their target anime style, such as Japanese anime style, American anime style, Chinese style anime style, and Q version anime style.
[0063] Step 202: Decompose the video to be converted into consecutive frame images, each frame image including key elements obtained through content recognition.
[0064] In this embodiment of the application, after receiving the video to be converted input by the user, the video to be converted is decomposed into continuous frame images through video parsing technology, and the decomposed images are preprocessed. Then, the preprocessed images are subjected to content recognition to automatically detect and mark key elements such as people and environment elements.
[0065] In this application, the preprocessing includes at least image enhancement processing for each frame of the image, removing noise from the image, highlighting details, and adjusting each frame of the image to a set size to avoid resolution inconsistencies.
[0066] It is worth noting that the human element in this application refers to the facial features, body, posture, and clothing of the human figure in the image, while the environmental element refers to the scene features other than the human figure, including the background, buildings, natural scenery, and other objects.
[0067] As an example, the cv2.VideoCapture interface of OpenCV (Open Source Computer Vision Library) or the FFmpeg batch export command can be used to split the video to be converted into continuous image frames, such as 30 frames per second, and save them as an ordered sequence according to timestamps to ensure that the frame number is consistent with the time order.
[0068] Step 203: Based on the target style characteristics of the target animation style, perform style conversion on key elements to generate several initial animation images.
[0069] In this embodiment of the application, it is necessary to extract the spatial structure between character elements and environment elements, then obtain the target style features corresponding to the target anime style from the preset anime style library, and parse the target style features to obtain character style features and environment style features. Keeping the spatial structure unchanged, the character style features are mapped to character elements, and the environment style features are mapped to environment elements to generate several frames of initial anime images.
[0070] Among them, the spatial structure between character elements and environment elements refers to the information on the positional relationship and relative proportion between characters and environment in the same frame image. The spatial structure is used to ensure that the geometric position of characters and environment does not shift significantly during the style conversion process, and the converted animation video can restore the composition of the video to be converted.
[0071] The target style features are a set of style parameters extracted from a pre-defined anime style library, including at least color tone, line thickness, texture style, and rendering method. By analyzing the target style features, character style features and environmental style features can be further derived.
[0072] Character style features are primarily used to depict a character's face, body, posture, and clothing, highlighting the character's individual characteristics. Environmental style features, on the other hand, are used to describe the shape or style of the background, buildings, natural scenery, and other non-human objects, presenting a visual effect consistent with the target anime style.
[0073] As an example, for each target anime style selected by the user, the improved FLUX.1Kontext model can be used to perform style transfer on each preprocessed frame of the image to generate an initial anime image that matches the target anime style.
[0074] Among them, FLUX.1 Kontext is an image generation model whose core architecture is a stream matching architecture.
[0075] During training, the model learns from a multimodal pre-training dataset of millions of high-quality images, covering dozens of core semantic categories, effectively improving its generalization ability across multiple scenarios and objects. Simultaneously, an adversarial filtering mechanism is introduced during training to remove non-compliant and sensitive content, ensuring the security and reliability of the generated results. Furthermore, the model employs a hierarchical guided distillation strategy, transforming expert annotations into learnable vector representations. This allows the model to maintain high descriptive optimization capabilities and instruction compliance even in zero-shot instruction scenarios, enabling it to generate images conforming to the expected style based on user input without requiring additional training samples.
[0076] In practical use, the FLUX.1 Kontext model can achieve the following: 1) Context awareness: It can combine image context information to generate anime-style images that meet user needs more accurately. Especially when replacing or modifying the style of local areas, it automatically preserves the overall composition of the image, ensuring that the replacement and editing of local areas have a high success rate and avoiding affecting the integrity of other areas.
[0077] 2) 3D rotational position encoding: By encoding the three-dimensional spatial position of elements in the image, it simulates the spatial understanding ability of human designers and can accurately capture the spatial structure and semantic relationship between people, environment and other visual elements, so that the style-transformed image remains natural and harmonious.
[0078] 3) Multimodal fusion mechanism: The user-input text description and image are uniformly mapped to the same semantic space. The end-to-end conversion of "text description - image generation" is realized through semantic flow, so that the model can understand text and image information at the same time, thereby generating high-fidelity anime-style images that conform to user instructions.
[0079] Based on the aforementioned context-aware, 3D rotational position encoding, and multimodal fusion mechanisms, the improved FLUX.1 Kontext model can accurately understand the semantic relationships of images and user commands, thereby achieving high-fidelity and high-precision image processing results. For example, the improved FLUX.1 Kontext model can not only create visual effects based on the user's input text description and input image, but also generate images that meet the user's needs more accurately compared to traditional models that rely solely on text prompts. It also supports pixel-level fine-tuning, allowing for precise adjustments to facial expressions, background replacement, and color themes. Furthermore, the improved FLUX.1 Kontext model can introduce different visual styles into images, such as migrating from a realistic style to a cartoon rendering style, while maintaining high fidelity through multi-scale feature alignment technology.
[0080] This application can map the user-selected target anime style to each frame of the video to be converted with high fidelity, generating initial anime images with consistent style and rich detail.
[0081] Step 204: Combine the dubbing audio with several frames of initial animation images to generate an initial animation video with synchronized audio and video.
[0082] In this embodiment, the dubbing audio is aligned with the timeline of the video to be converted. The start and end times and playback position of the dubbing audio are adjusted according to the timeline alignment result to ensure that the audio content corresponding to each frame of the initial animation image is played synchronously with it. At the same time, the initial animation images are spliced into a video sequence in chronological order, and the dubbing audio is embedded in the video track to generate an initial animation video with synchronized audio and video.
[0083] The subsequent adjustments in this application mainly involve processing certain frames of the initial animation video for consistency based on the image differences of multiple initial animation images. This not only avoids disrupting the established audio-visual synchronization relationship but also helps maintain the overall coherence and stylistic unity of the video, ensuring smooth transitions and stylistic harmony in the target animation video.
[0084] Step 205: Adjust the initial animation video according to the inter-frame deviation between multiple initial animation images to generate the target animation video.
[0085] In this embodiment, image differences include inter-frame deviation, which refers to differences between adjacent image frames that cause the initial animation video to exhibit image jumps or discontinuities during playback. For example, abrupt changes in pixel values, color distribution, line thickness, texture, or brightness between adjacent frames can cause flickering, jumps, or uneven motion during video playback.
[0086] Inter-frame deviation can be calculated by comparing the style and pixel differences between adjacent frames to locate abnormal transition regions. When inter-frame deviation is detected between adjacent frames, the corresponding transition regions can be located and fine-tuned to improve the continuity and smoothness of the video.
[0087] As an optional example of the embodiments of this application, adjacent initial animation images can be used as the first animation image and the second animation image, and the second animation image is the next frame image of the first animation image. The inter-frame deviation between the first animation image and the second animation image is calculated, and the jump region where the inter-frame deviation exists in the second animation image is located. The jump region is fine-tuned to generate the fine-tuned target animation image, until the initial animation images are traversed and the target animation video is generated.
[0088] Inter-frame deviation includes style feature differences between adjacent frames and pixel-level differences.
[0089] Style feature differences are used to quantify the changes in color tone, line thickness, and texture style between adjacent frames. They are mainly based on style features such as color, line, and texture to calculate the differences between adjacent frames and determine whether abrupt changes have occurred between adjacent frames based on the relationship between the difference value and a preset threshold.
[0090] As an example, for color differences, the mean differences (ΔH, ΔS, ΔV) of the HSV (Hue, Saturation, Value) channels of the first and second anime images can be calculated. If the mean difference is greater than a first set threshold, the second anime image is determined to have a color abrupt change problem. Alternatively, the Bartholomew's distance between the color histograms of the first and second anime images can be calculated. If the Bartholomew's distance is greater than a second set threshold, the color distribution of the second anime image is determined to have a significant difference. For line and edge differences, the Canny edge detection algorithm can be used to extract the edge masks of the first and second anime images and calculate the intersection-union ratio (IU / R) of the edge pixels. If the IU / R is less than a third set threshold, the line structure of the second anime image is determined to have changed too much, such as suddenly changing from thick lines to thin lines.
[0091] Pixel-level difference is used to determine the pixel changes between adjacent frames by directly comparing the pixel values of adjacent frames. It mainly calculates the absolute difference of each pixel, generates a grayscale difference map based on the absolute difference, and determines whether there are differences between adjacent frames based on the relationship between the absolute difference and a preset threshold.
[0092] As an example, the absolute difference between the first anime image and the second anime image is calculated. If the absolute difference is greater than a fourth set threshold, it is determined that the second anime image has a pixel-level jump.
[0093] After calculating the inter-frame deviation between the first and second anime images using the above method, it is necessary to locate the transition region in the second anime image where the inter-frame deviation exists.
[0094] The transition region refers to a local area where abnormal changes occur due to style feature differences or pixel-level differences between adjacent frames exceeding a preset threshold. This application can accurately mark the local areas that need fine-tuning, and make targeted fine-tuning of the local areas to improve the accuracy of style transfer.
[0095] As an optional example of an embodiment of this application, the process of fine-tuning the transition region to generate a fine-tuned target animation image, until the initial animation image has been traversed, at least includes: superimposing the first animation image and the second animation image to obtain an intermediate frame, and filling the transition region of the second animation image with the intermediate frame to generate the fine-tuned target animation image. Alternatively, the forward and backward optical flows of the first and second animation images are calculated, and based on the forward and backward optical flows, pixel remapping is performed on the first and second animation images to generate an intermediate frame, and the intermediate frame is filled with the transition region of the second animation image to generate the fine-tuned target animation image.
[0096] In this context, overlay refers to a frame blending method that mixes adjacent frames in a certain proportion to generate a supplementary frame. This involves fusing the first and second animation images to obtain an intermediate frame used to fill in transition areas. Frame blending can improve animation smoothness to some extent, with relatively low computational cost.
[0097] For scenes with complex motion, this application uses optical flow to infer the trajectory of pixel movement based on previous and next frames, and then generates new interpolated frames. Optical flow can handle fast-moving objects more accurately, and is especially suitable for scenes containing objects with no motion blur or objects moving in front of a generally static background, thereby ensuring the smoothness of video motion and visual naturalness.
[0098] In practical applications, this application calculates the forward and backward optical flow of the first and second anime images using the optical flow method. The forward optical flow is used to characterize the pixel motion vector from the first anime image to the second anime image, and the backward optical flow is used to characterize the pixel motion vector from the second anime image to the first anime image.
[0099] In this application, the corresponding positions of pixels in the first anime image and the corresponding positions in the second anime image can be determined based on the forward optical flow, thereby constructing a mapping relationship. Similarly, the corresponding positions of pixels in the second anime image and the corresponding positions in the first anime image can be determined based on the backward optical flow, thereby constructing a mapping relationship. Based on the constructed mapping relationship, pixel remapping is performed on the first and second anime images to generate intermediate frames.
[0100] By filling the transition regions of the second animation image with the generated intermediate frames, the transition regions are replaced while the non-transition regions remain unchanged, thus generating a finely tuned target animation image.
[0101] As an example, the forward and backward optical flows of adjacent first and second anime images are calculated to describe the motion direction and distance of pixels from the first anime image to the second anime image, and vice versa. A mapping relationship is constructed, and pixels from adjacent frames are remapped to the other frame, thereby moving pixels from the first and second anime images to an intermediate time position, and then merging them to generate an intermediate frame. Furthermore, during the generation of the intermediate frame, this application can also identify occlusion regions through bidirectional optical flow consistency checks, reducing the weight of occluded frames.
[0102] For example, the first animation image is denoted as It, and the second animation image is denoted as It+1. The forward optical flow (It→It+1) and the backward optical flow (It+1→It) are calculated separately. Based on the forward optical flow, the It pixel is mapped to the position at time t+Δt, and based on the backward optical flow, the It+1 pixel is mapped to the position at time t+Δt. When generating intermediate frames, assuming the pixel movement is uniform, the position of the intermediate frame (e.g., between t and t+1, i.e., t+0.5) can be interpolated using half of the optical flow vector. For example, the forward and backward optical flows can be scaled by half to estimate the pixel position at the intermediate time. The pixel of frame t is moved by half the distance along the forward optical flow direction, and the pixel of frame t+1 is moved by half the distance along the backward optical flow direction, thus obtaining two candidate frames. These two candidate frames are then combined into the final intermediate frame using a weighted average or fusion strategy.
[0103] Since motion may not be perfectly linear and optical flow estimation may contain errors, artifacts may appear in the synthesized intermediate frames. Therefore, this application can combine post-processing methods such as bidirectional optical flow consistency checking, pixel weight adjustment, and hole filling to improve the smoothness and naturalness of the generated frames.
[0104] Furthermore, during this process, temporal filtering can be applied to the image to reduce flicker, and spatial filtering can be applied to the image to smooth edges and artifacts.
[0105] Step 206: Adjust the initial animation video based on the style deviation between multiple initial animation images to generate the target animation video.
[0106] In the embodiments of this application, image differences include style deviation, which refers to the phenomenon that a certain frame of an image has an inconsistent style when the initial animation video is played. For example, most frames of the initial animation video are warm-toned, but a certain frame has a distinctly cool tone, or the outline of the characters in the preceding and following frames suddenly changes from thick lines to thin lines, etc.
[0107] Style deviation can be addressed by quantifying the similarity of the style vectors of each initial anime image frame. Initial anime images with low similarity can be directly marked as anomalous anime images, and the difference between the anomalous anime image and the previous frame image can be calculated. As an optional example of this embodiment, for several initial anime images, the style deviation of each initial anime image is calculated, and the corresponding initial anime image is used as the second anime image, while the previous frame image of the second anime image is used as the first anime image. The second anime image is adjusted to generate an adjusted target anime image, until all initial anime images have been traversed, resulting in a target anime video. Further, based on the style vector of the first anime image, the second anime image is re-rendered to generate the adjusted target anime image.
[0108] As an example, style deviation can be determined by calculating style vectors such as vector distance, color difference, or line difference, and further based on the relationship between the style vectors and preset thresholds. For vector distance, Euclidean distance can be used to measure the similarity of style vectors of adjacent frames in a high-dimensional feature space. The smaller the distance value, the closer the styles are. If the distance value exceeds the fourth preset threshold, it is determined that there is a significant style difference. Distance values exceeding the fourth preset threshold are identified as style deviations, and the corresponding initial animation image is used as the second animation image, and the previous frame of the second animation image is used as the first animation image. For color difference, the HSV channel difference of the main color can be calculated. If the hue angle difference between two frames exceeds the fifth preset threshold, it indicates that there may be color inconsistency. Alternatively, the cross-entropy of the color histogram can be used to quantify the overall color distribution difference; the larger the value, the more inconsistent the color distribution. For line difference, the thickness variation of edge lines is statistically analyzed, and the standard deviation of line thickness in adjacent frames is calculated. If the standard deviation exceeds the sixth preset threshold, it indicates a sudden change in line style and style inconsistency.
[0109] As another example, for the second anime image where a style deviation is detected, its previous frame image is used as a reference, i.e., the first anime image. The style vector of the first anime image is calculated, and the second anime image is re-rendered based on the style vector of the first anime image to generate the adjusted target anime image. This process is repeated until all initial anime images have been traversed, thereby generating a target anime video with a unified style.
[0110] In practical applications, users can view the conversion effect through the style conversion interface. If they are not satisfied with a certain style or voice-over, they can reselect the style or adjust the voice-over parameters and re-execute the corresponding conversion and compositing steps until a satisfactory target animation video is generated.
[0111] To facilitate understanding of the video style transfer method in this application embodiment, a detailed explanation is provided below with reference to a complete example.
[0112] 1) The user uploads a 2-minute video of a daily life scene through the style conversion interface, with a frame rate of 25 frames per second. After receiving the video, video parsing technology is used to split it into 3000 independent frames according to the frame rate, and each frame is saved as a JPEG file, ensuring that the frame number is consistent with the time sequence.
[0113] 2) Process the extracted 3000 frames. First, enhance the sharpness and remove noise to make the images clearer, then adjust all images to the standard size of 800×600.
[0114] 3) The user selected "Japanese anime style" and "Q version anime style" as the target anime style from the multiple anime style library.
[0115] 4) Use a pre-trained deep convolutional neural network to extract the target style features of "Japanese anime style" and "Q version anime style" respectively.
[0116] Japanese anime style: delicate lines, soft colors, realistic character proportions, and background textures that closely resemble reality.
[0117] Q-version anime style: rounded lines, bright colors, large heads and small bodies for characters, and cute and rounded shapes for background objects.
[0118] 5) Utilizing the FLUX.1Kontext model, receive preprocessed image frames and target style features, and allow the user to set the first 1000 frames to use a Japanese style and the next 2000 frames to use a chibi style. Then, perform the conversion based on key elements (such as characters, backgrounds, and objects). The first 1000 frames of images are designed to depict figures with delicate detail and soft colors, while the textures of trees and houses present a typical Japanese style.
[0119] For the last 2000 frames, the characters are made cute, the colors are vibrant, and the background buildings and trees have rounded lines.
[0120] 6) The converted 3000 frames are stitched together sequentially to form a preliminary animated video. The generated voice-over audio is synchronized with the preliminary animated video to ensure that the dialogue matches the characters' lip movements and that the sound effects match the actions, thus creating an animated video with voice-over.
[0121] 7) Frame consistency processing: Calculate the differences between adjacent frames, and fine-tune the first 1000 frames of Japanese style and the last 2000 frames of Q version style images to eliminate flickering and frame skipping.
[0122] 8) Style consistency processing: After inspection, it was found that some Q version style frames had large color deviations. They were adjusted uniformly to ensure overall style coordination, while maintaining consistency in the Japanese style.
[0123] 9) Double-check the voice-over and visuals to ensure they are perfectly synchronized before generating the target animation video.
[0124] 10) Show the target anime video to users. Users can provide feedback after watching.
[0125] For example, if a user feels that the sound effects in the Q-style section could be richer, they can input a description of the adjustment through the style conversion interface, which will automatically trigger video adjustment commands, add dubbing audio that meets the user's needs, and then output the video again. Once the user is satisfied, they can download the video.
[0126] In this embodiment, the following steps are taken: The video to be converted, the dubbing audio, and the target anime style corresponding to the video to be converted are obtained; the video to be converted is decomposed into consecutive frame images, each frame including key elements obtained through content recognition; based on the target style features of the target anime style, style conversion is performed on the key elements to generate several initial anime images; the dubbing audio is combined with several initial anime images to generate an initial anime video with synchronized audio and video; the initial anime video is adjusted according to the inter-frame deviation between multiple initial anime images to generate the target anime video; and / or, the initial anime video is adjusted according to the style deviation between multiple initial anime images to generate the target anime video.
[0127] The technical solution of this application can not only perform high-fidelity style conversion on key elements based on the target anime style, so that the generated anime video is consistent with the target anime style, but also adjust the image through inter-frame deviation and style deviation to avoid problems such as flickering, abrupt changes or style inconsistency, thereby improving the precision of anime style conversion and meeting the personalized needs of users.
[0128] Corresponding to the aforementioned application function implementation method embodiments, this application also provides a video style conversion device, electronic device, and corresponding embodiments.
[0129] See Figure 3 , Figure 3 This is a schematic diagram of the structure of a video style conversion device shown in an embodiment of this application.
[0130] The conversion data acquisition module 301 is used to acquire the video to be converted, the dubbing audio, and the target anime style corresponding to the video to be converted. The video decomposition module 302 is used to decompose the video to be converted into consecutive frame images, each frame image including key elements obtained through content recognition; Style conversion module 303 is used to perform style conversion on key elements based on the target style features of the target animation style, and generate several frames of initial animation images; The video synthesis module 304 is used to combine the dubbing audio with several frames of initial animation images to generate an initial animation video with synchronized audio and video. The video adjustment module 305 is used to adjust the initial animation video based on the image differences between multiple initial animation images to generate the target animation video.
[0131] As an optional example of an embodiment of this application, the key elements include character elements and environment elements, and the style conversion module 303 is used for: Extract the spatial structure between character elements and environmental elements; Retrieve the target style features corresponding to the target anime style from the preset anime style library; Analyze the target style characteristics to obtain character style characteristics and environmental style characteristics; While keeping the spatial structure unchanged, character style features are mapped to character elements, and environmental style features are mapped to environmental elements to generate several initial animation images.
[0132] As an optional example of an embodiment of this application, the image difference includes inter-frame deviation, and the video adjustment module 305 includes: The adjacent frame image determination submodule is used to use adjacent initial animation images as the first animation image and the second animation image, and the second animation image is the next frame image of the first animation image; The inter-frame deviation calculation submodule is used to calculate the inter-frame deviation between the first animation image and the second animation image, and to locate the jump region in the second animation image where the inter-frame deviation exists. The jump region fine-tuning submodule is used to fine-tune the jump region and generate the fine-tuned target animation image until the initial animation image has been traversed and the target animation video is generated.
[0133] As an optional example of an embodiment of this application, the jump region fine-tuning submodule is used for: The first and second anime images are superimposed to obtain the intermediate frame; Fill the transition area of the second animation image with the intermediate frame to generate the fine-tuned target animation image.
[0134] As an optional example of an embodiment of this application, the jump region fine-tuning submodule is used for: Calculate the forward optical flow and backward optical flow of the first animation image and the second animation image; wherein, the forward optical flow is used to characterize the pixel motion vector from the first animation image to the second animation image, and the backward optical flow is used to characterize the pixel motion vector from the second animation image to the first animation image; Based on forward and backward optical flow, pixel remapping is performed on the first and second animation images to generate intermediate frames; Fill the transition area of the second animation image with the intermediate frame to generate the fine-tuned target animation image.
[0135] As an optional example of an embodiment of this application, the image difference includes style deviation, and the video adjustment module 305 includes: The style deviation calculation submodule is used to calculate the style deviation of each initial animation image for a number of initial animation images, and to use the corresponding initial animation image as the second animation image, and the previous frame of the second animation image as the first animation image. The image adjustment submodule is used to adjust the second animation image, generate the adjusted target animation image, and continue until the initial animation image has been traversed to generate the target animation video.
[0136] As an optional example of an embodiment of this application, the image adjustment submodule is used for: Based on the style vector of the first anime image, the second anime image is re-rendered to generate the adjusted target anime image.
[0137] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated further here.
[0138] Figure 4 This is a schematic diagram of the structure of an electronic device shown in an embodiment of this application.
[0139] See Figure 4 The electronic device 400 includes a memory 410 and a processor 420.
[0140] The processor 420 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. Memory 410 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM may store static data or instructions required by processor 420 or other modules of the computer. Permanent storage devices may be read-write storage devices. Permanent storage devices may be non-volatile storage devices that retain stored instructions and data even when the computer is powered off. In some embodiments, permanent storage devices use mass storage devices (e.g., magnetic or optical disks, flash memory) as permanent storage devices. In other embodiments, permanent storage devices may be removable storage devices (e.g., floppy disks, optical drives). System memory may be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. System memory may store some or all of the instructions and data required by the processor during operation. Furthermore, memory 410 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and disks and / or optical disks may also be used. In some embodiments, memory 410 may include a removable storage device that is readable and / or writable, such as a laser disc (CD), a read-only digital multifunction optical disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra-high density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not contain carrier waves or transient electronic signals transmitted wirelessly or via wired connections.
[0141] The memory 410 stores executable code, which, when processed by the processor 420, can cause the processor 420 to execute part or all of the methods described above.
[0142] Furthermore, the method according to this application can also be implemented as a computer program or computer program product, which includes computer program code instructions for performing some or all of the steps in the method described above.
[0143] Alternatively, this application may be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium) storing executable code (or computer program or computer instruction code) that, when executed by a processor of an electronic device (or server, etc.), causes the processor to perform part or all of the steps of the methods described above according to this application.
[0144] This application also provides a computer program product, which includes computer instructions that, when executed by a processor, implement the method described above.
[0145] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A style transfer method of a video, characterized by, The method comprises: acquiring a to-be-converted video, dubbing audio, and a target animation style corresponding to the to-be-converted video; decomposing the to-be-converted video into continuous frame images, each frame image comprising key elements obtained through content recognition; performing style conversion on the key elements based on target style features of the target animation style to generate a plurality of initial animation frame images; combining the dubbing audio and the plurality of initial animation frame images to generate an initial animation video with audio-visual synchronization; adjusting the initial animation video according to image differences between the plurality of initial animation frame images to generate a target animation video.
2. The method of claim 1, wherein, The key elements comprise character elements and environment elements, and the performing style conversion on the key elements based on target style features of the target animation style to generate a plurality of initial animation frame images comprises: extracting a spatial structure between the character elements and the environment elements; acquiring target style features corresponding to the target animation style from a preset animation style library; analyzing the target style features to obtain character style features and environment style features; keeping the spatial structure unchanged, mapping the character style features to the character elements, and mapping the environment style features to the environment elements to generate a plurality of initial animation frame images.
3. The method of claim 1, wherein, The image differences comprise inter-frame deviations, and the adjusting the initial animation video according to image differences between the plurality of initial animation frame images to generate a target animation video comprises: taking adjacent initial animation frame images as a first animation frame image and a second animation frame image, and the second animation frame image being a next frame image of the first animation frame image; calculating an inter-frame deviation between the first animation frame image and the second animation frame image, and locating a jump region of the second animation frame image with the inter-frame deviation; fine-tuning the jump region to generate a fine-tuned target animation frame image, until the initial animation frame images are traversed to generate the target animation video.
4. The method of claim 3, wherein, The fine-tuning the jump region to generate a fine-tuned target animation frame image comprises: superimposing the first animation frame image and the second animation frame image to obtain an intermediate frame; filling the intermediate frame into the jump region of the second animation frame image to generate the fine-tuned target animation frame image.
5. The method of claim 3, wherein, The fine-tuning the jump region to generate a fine-tuned target animation frame image comprises: calculating forward optical flow and backward optical flow of the first animation frame image and the second animation frame image; wherein the forward optical flow is used to represent a pixel motion vector from the first animation frame image to the second animation frame image, and the backward optical flow is used to represent a pixel motion vector from the second animation frame image to the first animation frame image; performing pixel remapping on the first animation frame image and the second animation frame image based on the forward optical flow and the backward optical flow to generate an intermediate frame; filling the intermediate frame into the jump region of the second animation frame image to generate the fine-tuned target animation frame image.
6. The method of claim 1, wherein, The image differences comprise style deviations, and the adjusting the initial animation video according to image differences of the initial animation frame images to generate a target animation video comprises: For the several frames of initial animation images, the style deviation of each frame of initial animation image is calculated, and the corresponding initial animation image is taken as a second animation image, and the last frame image of the second animation image is taken as a first animation image; Adjusting the second animation image generates an adjusted target animation image, until the initial animation image is traversed, and the target animation video is generated.
7. The method of claim 6, wherein, The adjustment of the second animation image includes: Based on the style vector of the first animation image, the second animation image is re-rendered to generate the adjusted target animation image.
8. A video style conversion device, characterized in that, Including: A conversion data acquisition module is configured to acquire a to-be-converted video, a dubbing audio, and a target animation style corresponding to the to-be-converted video; A video decomposition module is configured to decompose the to-be-converted video into continuous frame images, and each frame image includes key elements obtained through content recognition; A style conversion module is configured to perform style conversion on the key elements based on target style features of the target animation style to generate several frames of initial animation images; A video synthesis module is configured to combine the dubbing audio and the several frames of initial animation images to generate an initial animation video with synchronized sound and picture; A video adjustment module is configured to adjust the initial animation video according to image differences between multiple frames of initial animation images to generate a target animation video.
9. An electronic device, comprising: Including: A processor; And A memory having executable code stored thereon, which, when executed by the processor, causes the processor to perform the method of any one of claims 1-7.
10. A computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the method of any one of claims 1-7.