Training of cartoon sketch image reconstruction network and reconstruction method and device thereof
By collecting and training a generative adversarial network, real-world movie data and virtual-world animation data are used to train a cartoon sketch image reconstruction network, which solves the problem of low efficiency in style transfer of video data and achieves efficient cartoon sketch style reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies are inefficient, have high production barriers, and offer limited filter effects when converting video data to cartoon or sketch styles, making it difficult to achieve the desired cartoon or sketch style.
We collect real-world movie data and virtual-world animation data, extract content sample and style sample image data, train a generative adversarial network to reconstruct a cartoon sketch image network, and reconstruct the image data into a cartoon sketch style through the generative adversarial network.
It improves the efficiency of creating cartoon sketch-style videos from video data, lowers the production threshold, and achieves efficient style conversion.
Smart Images

Figure CN115272057B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision, and in particular to a training method and device for reconstructing cartoon sketch images using a network. Background Technology
[0002] In scenarios such as short videos and advertisements, users create various types of video data. After recording the original video data, the video data is usually post-processed to improve the quality of the video data.
[0003] Due to certain business needs, some post-processing involves converting the style of video data to cartoon, sketch, or other styles. Currently, the most common post-processing involves adding filters to the video data to convert it to other styles, such as retro, film, sunset, etc.
[0004] However, filters typically adjust the color values of pixels and add other decorative elements, resulting in a relatively simple effect. It is also difficult to achieve styles such as cartoons and sketches by using multiple filters. If video data is designed according to styles such as cartoons and sketches, this will greatly increase the threshold for video data production, resulting in a significant increase in the time required for video data production and low efficiency. Summary of the Invention
[0005] This invention provides a training method and device for a cartoon sketch image reconstruction network, which aims to solve the problem of how to efficiently recreate the style of a cartoon sketch in an image.
[0006] According to one aspect of the present invention, a method for training a cartoon sketch image reconstruction network is provided, comprising:
[0007] Collect data from movies whose stories take place in the real world, and data from multiple animated films whose stories take place in the virtual world;
[0008] Extract multiple frames of image data from the movie data as content sample image data;
[0009] Filter the animation data in a sketch style from the multiple animation data sets;
[0010] Extract multiple frames of image data from the animation data that are in a sketch style, and use them as style sample image data;
[0011] Based on the content sample image data and the style sample image data, a generative adversarial network is trained to become a cartoon sketch image reconstruction network, which is used to reconstruct image data containing a cartoon sketch style.
[0012] According to another aspect of the present invention, an image reconstruction method is provided, comprising:
[0013] Load a cartoon sketch image reconstruction network trained by the method according to any embodiment of the present invention;
[0014] Obtain the original image data to be reconstructed;
[0015] The original image data is input into the cartoon sketch image reconstruction network and reconstructed into target image data containing a cartoon sketch style.
[0016] According to another aspect of the present invention, a video reconstruction method is provided, characterized in that it includes:
[0017] Load a cartoon sketch image reconstruction network trained by the method according to any embodiment of the present invention;
[0018] The content to be acquired is the original video data introducing the game, which contains multiple frames of original image data;
[0019] The original image data is input into the cartoon sketch image reconstruction network and reconstructed into target image data containing a cartoon sketch style;
[0020] The target image data is replaced in the original video data to obtain the target video data.
[0021] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0022] At least one processor; and
[0023] A memory communicatively connected to the at least one processor; wherein,
[0024] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the training method, image reconstruction method, or video reconstruction method of the cartoon sketch image reconstruction network according to any embodiment of the present invention.
[0025] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, the computer program being configured to cause a processor to execute and implement the training method, image reconstruction method, or video reconstruction method of the cartoon sketch image reconstruction network according to any embodiment of the present invention.
[0026] In this embodiment, data from movies whose stories take place in the real world and animation data from multiple animations whose stories take place in the virtual world are collected. Multiple frames of image data are extracted from the movie data as content sample image data. Animation data exhibiting a sketching style is selected from the multiple animation data. Multiple frames of image data are extracted from the sketching-style animation data as style sample image data. Based on the content sample image data and the style sample image data, a generative adversarial network (GAN) is trained into a cartoon sketching image reconstruction network. This network is used to reconstruct image data containing a cartoon sketching style. By selecting a sketching style based on the cartoon style of the animation data, and combining the two, a cartoon sketching style can be obtained. This is used to train the GAN, enabling the cartoon sketching image reconstruction network to reconstruct image data into a cartoon sketching style. Reconstructing the cartoon sketching style is a post-processing step, which maintains the threshold for video data production and reduces the time required for video data production, greatly improving the efficiency of producing cartoon sketching style video data.
[0027] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a flowchart of a training method for a cartoon sketch image reconstruction network according to Embodiment 1 of the present invention;
[0030] Figure 2 This is an example diagram of an animated character provided according to Embodiment 1 of the present invention;
[0031] Figure 3 This is a flowchart of an image reconstruction method provided according to Embodiment 2 of the present invention;
[0032] Figure 4A and Figure 4B This is an example diagram of a reconstructed cartoon sketch provided according to Embodiment 2 of the present invention;
[0033] Figure 5 This is a flowchart of a video reconstruction method provided in Embodiment 3 of the present invention;
[0034] Figure 6This is a schematic diagram of the structure of a training device for a cartoon sketch image reconstruction network according to Embodiment 4 of the present invention;
[0035] Figure 7 This is a schematic diagram of the structure of an image reconstruction device according to Embodiment 5 of the present invention;
[0036] Figure 8 This is a schematic diagram of the structure of a video reconstruction device according to Embodiment Six of the present invention;
[0037] Figure 9 This is a schematic diagram of the structure of an electronic device provided in Embodiment 7 of the present invention. Detailed Implementation
[0038] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0039] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0040] Example 1
[0041] Figure 1 This is a flowchart of a training method for a cartoon sketch image reconstruction network provided in Embodiment 1 of the present invention. This embodiment is applicable to training a cartoon sketch image reconstruction network that achieves a cartoon sketch style. This method can be executed by a training device for the cartoon sketch image reconstruction network, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0042] Step 101: Collect movie data where the story takes place in the real world and animation data where the story takes place in the virtual world.
[0043] On the one hand, multiple movie data can be collected through authorized use, publicly available datasets, self-recording, etc. Generally, each movie data tells a story within a limited time (such as 1-3 hours). In this embodiment, the stories of the movie data collected take place in the real world.
[0044] The real world can include real natural environments, real buildings, real people, animals, and so on.
[0045] On the other hand, multiple animation data can be collected through authorized use, public datasets, self-recording, etc. Each animation data tells a story that takes place in a virtual world. If a certain animation data belongs to a season of animation data, then the animation data contains multiple episodes with a relatively short duration (such as 10-30 minutes). If a certain animation data belongs to the form of OVA (Original Video Animation), then the animation data belongs to a single animation data with a relatively long duration (such as 1-3 hours).
[0046] Step 102: Extract multiple frames of image data from the movie data as content sample image data.
[0047] Movie data is a form of video data. In this embodiment, multiple frames of image data can be extracted from each movie data through sampling methods such as random sampling and uniform sampling. These frames are used as samples for training the cartoon sketch image reconstruction network. For the cartoon sketch image reconstruction network, the image data used as samples belongs to the source of content and can therefore be referred to as content sample image data.
[0048] In one sampling method, command-line tools, library files, etc., can be used to divide the movie data into multiple segments, denoted as movie segments, with independent scenes as the segmentation nodes. Each movie segment contains one or more independent scenes.
[0049] Furthermore, the detection methods include the following two:
[0050] 1. Threshold mode
[0051] For movie data with obvious scene boundaries, a threshold mode is applied. Each frame of image data is compared with the set black level. Based on the detection results, it is determined whether it is the boundary of a scene such as fade-in, fade-out, or cut to black, thereby dividing the movie data into various scenes.
[0052] 2. Content Format
[0053] For movie data that switches rapidly between scenes, a content mode is applied. Each frame of image data is compared, and the image data with significant content changes is selected as the segmentation node, thereby dividing the movie data into various scenes.
[0054] Generally, movie data containing a single scene can be divided into a movie segment. However, considering that some movie data containing a single scene is short, the scene can be merged with other adjacent scenes, thereby dividing movie data containing two or more connected scenes into a movie segment. This embodiment does not impose any restrictions on this.
[0055] In each movie clip, one frame of image data is extracted at each preset first time interval as content sample image data.
[0056] In this embodiment, the movie data is divided into movie segments (i.e., slices) according to the scene, and frames are extracted from the movie segments. Since the content in the same scene is relatively fixed, slicing and frame extraction can improve the uniformity of the sampled image data, thereby improving the performance of the cartoon sketch image reconstruction network.
[0057] Step 103: Filter out animation data with a sketch style from multiple animation data sets.
[0058] Generally, the same animation data is produced by the same team, and the style of the same animation data is relatively uniform. However, considering the many influencing factors in the production process, there may be some differences in style between different episodes of animation data, and there may be significant differences in style between different animation data. Not every animation data will present a sketch style as a whole. In this embodiment, on the basis that the animation data basically presents a typical cartoon style, the style differences of each animation data can be further subdivided, so as to filter out the animation data that presents a more obvious sketch style as a whole.
[0059] In one embodiment of the present invention, step 103 may include the following steps:
[0060] Step 1031: Extract multiple frames of image data from each animation data set as reference image data.
[0061] Animation data is a form of video data. In this embodiment, multiple frames of image data can be extracted from each animation data through sampling methods such as random sampling and uniform sampling, and these are referred to as reference image data.
[0062] In one sampling method, command-line tools, library files, etc., can be used to divide the animation data into multiple segments, called animation segments, with independent scenes as the dividing nodes. Each animation segment contains one or more independent scenes.
[0063] Furthermore, the detection methods include the following two:
[0064] 1. Threshold mode
[0065] For animation data with obvious scene boundaries, a threshold mode is applied. Each frame of image data is compared with the set black level. Based on the detection results, it is determined whether it is the boundary of a scene such as fade-in, fade-out, or cut to black, thereby dividing the animation data into various scenes.
[0066] 2. Content Format
[0067] For animation data that switches rapidly between scenes, a content mode is applied. Each frame of image data is compared, and the image data with significant content changes is identified as the segmentation node, thereby dividing the animation data into various scenes.
[0068] Generally, animation data containing a single scene can be divided into an animation segment. However, considering that some animation data containing a single scene is relatively short, the scene can be merged with other adjacent scenes, thereby dividing animation data containing two or more connected scenes into an animation segment. This embodiment does not impose any restrictions on this.
[0069] In each animation segment, one frame of image data is extracted at each preset second time interval as reference image data.
[0070] In this embodiment, the animation data is divided into animation segments (i.e. slices) according to the scene, and frames are extracted from the animation segments. Since the content in the same scene is relatively fixed, slicing and frame extraction can improve the uniformity of the sampled reference image data.
[0071] Step 1032: Identify the outline data representing the sketch style from the reference image data.
[0072] In the art production process of animation data, the outer contour and inner contour of the object's surface are usually drawn simultaneously, and the width of the contour must be flexibly controlled. Taking the character in animation data as an example, when drawing the outer contour of the character, the contour line is generally drawn thicker. When drawing the turning points of the character, the contour line is generally drawn thicker, while some details of the face and body are drawn thinner.
[0073] Outline data consists of lines representing the edges of an outline. It can represent the style of a sketch to some extent. In the art production of animation data, outline data can be generated based on perspective (using the angle between the model's normal vector and the view vector; the closer the angle is to perpendicular, the closer it is to the outline), geometry generation methods (double-pass rendering, where the first pass renders the front of the object and the second pass renders the back of the object, making the outline visible), and image processing (passing in depth and normal information in the form of textures and using edge detection algorithms to find edges).
[0074] In some regions, animation data tends to use geometry-based methods for outlining. The advantage of this outlining method compared to the other two methods is that the line width is easier for artists to control. In other regions, animation data often uses outlining data with varying thickness to represent the characteristics of different parts of the character. In some cases, vertex colors per object are introduced to control the details of the outlining data, while also ensuring that the thickness of the outlining data does not change with the camera's viewing distance.
[0075] For animation data produced in different ways, this embodiment can identify outline data from reference image data of each frame in order to evaluate the strength of the sketch style of the animation data as a whole.
[0076] In one embodiment of the present invention, step 1032 may further include the following steps:
[0077] Step 10321: Detect head data containing hair data in the reference image data.
[0078] There are significant differences between different animation data, ranging from mecha to superpowers, from ancient to modern times to fantasy worlds, and so on. There are many types of objects with outline data added to the animation data. In order to make a unified comparison of the outline data of different animation data, this embodiment selects the head data of various characters that are widely distributed in different animation data, especially the hair data.
[0079] In animation data, the characters are primarily part of the story's plot, and users' attention is mostly focused on them. Considering factors such as drawing techniques, etc. Figure 2 As shown, characters in animation data are mostly distinguished by head data (hair data represents hairstyle), clothing, etc. Therefore, when artists create animation data, the outline data on the head data (especially hair data) of each character is drawn in more detail. Hair data is mostly flat, solid-color areas with less interference from other elements. In addition, the overall color of hair data has a more obvious color difference from the outline data, making it particularly suitable for separating outline data.
[0080] In practical implementation, cartoon face detection networks such as ACFD (Asymmetric Codone Face Detection Algorithm) can be used to perform face detection in reference image data to obtain the original detection boxes that identify the face data.
[0081] The original detection box is expanded horizontally and vertically upwards to cover the hair data. The step size for expanding the original detection box horizontally to the left and right, and the step size for expanding the original detection box vertically, are generally empirical values. For example, if the width of the original detection box is W and the height is H, it can be expanded by 1 / 3W to the left horizontally, 1 / 3W to the left horizontally, and 1 / 2H vertically upwards. In this way, the hair data can be basically included.
[0082] If the expansion is completed, the data of the original detection box located after the expansion can be extracted to obtain the original head data containing hair data.
[0083] Step 10322: Perform amplification processing on the header data.
[0084] Generally, the stroke data is smaller in size than the entire head data. If compared at the original size, it would be too sensitive. Therefore, the head data can be enlarged. Since the stroke data is mostly solid color, it can be compared at a larger magnification. Even if there are obvious jagged edges in the overall image, causing image distortion, it will not affect the comparison of the stroke data.
[0085] In one example, we can determine the first size of the header data before magnification (width srcWidth, height srcHeight) and the second size of the header data after magnification (width dstWidth, height dstHeight), and calculate the ratio between the first and second sizes. The second size is larger than the first size, and the ratio between the first and second sizes is the magnification factor.
[0086] Round the product of the magnified header data coordinates (dstX, dstY) and the scale to obtain the original header data coordinates (srcX, srcY).
[0087] srcX=dstX*(srcWidth / dstWidth)
[0088] srcY=dstY*(srcHeight / dstHeight)
[0089] Assign color to the pixels located at the coordinates of the header data before magnification, and then to the pixels located at the coordinates of the header data after magnification.
[0090] In this example, the colors of the pixels in the header data before magnification are proportionally mapped to the pixels in the header data after magnification. This ensures that the outline data remains unchanged, and the calculation is simple and the operation is convenient.
[0091] Step 10323: Perform binarization on the magnified header data to distinguish between black and white.
[0092] Since the outline data is mostly black, binarization can be performed on the enlarged header data in both black and white dimensions.
[0093] In the actual implementation, the red component R, green component G, and blue component B of each pixel in the magnified header data can be queried.
[0094] If the red component R is less than or equal to the first threshold, the green component G is less than or equal to the first threshold, and the blue component B is less than or equal to the first threshold, then the pixel is set to black (i.e., 0).
[0095] If at least one of the following conditions is met: the red component R is greater than the first threshold, the green component G is greater than the first threshold, and the blue component B is greater than the first threshold, then the pixel is set to white (i.e., 255).
[0096] Step 10324: Perform erosion processing on the binarized header data.
[0097] Step 10325: Perform dilation processing on the eroded header data.
[0098] The binarized header data may contain some noise. In this case, erosion processing can be performed on the binarized header data. Erosion processing enhances and expands areas with small gray values (visually darker), which can be used to remove bright noise, reduce the impact of noise on the statistics of the outline data, and reduce errors.
[0099] After erosion, the header data will shrink to a certain extent. At this time, dilation (erode) can be performed on the eroded header data. Dilation enhances and expands the areas with large gray values (visually brighter), mainly used to connect regions with similar colors or intensities (i.e. connected regions).
[0100] Step 10326: Detect black pixels in the expanded head data to obtain outline data representing the sketch style.
[0101] Pixels representing black (i.e., 0) are detected in the expanded head data to obtain stroke data that characterizes the sketch style.
[0102] Step 10327: Correct the outline data using at least one of the area or coordinates.
[0103] In practical applications, elements such as hair, eyebrows, eyes, and mouth in animated characters may also be black, which can interfere with the stroke data to some extent. Therefore, the stroke data can be corrected by analyzing factors such as the area and coordinates of the stroke data and using at least one of the area and coordinates.
[0104] In one example, for each stroke data belonging to an independent connected region, the area of the stroke data is counted (which can be equivalent to the number of pixels).
[0105] If the area is less than or equal to the second threshold, it means that the area of the outline data is small and the outline data is more reliable, so the outline data is retained.
[0106] If the area is greater than the second threshold, it means that the area of the outline data is large and may belong to hair data, so the outline data is filtered out.
[0107] In another example, when querying head data, the query records a region consisting of facial key points that represent facial features (such as eyebrows, eyes, mouth, etc.).
[0108] For each outline data belonging to an independent connected region, the coordinates of the outline data are compared with the region.
[0109] If the stroke data is outside the area and the stroke data is considered reliable, then the stroke data is retained.
[0110] If the outline data is located within a region, and the outline data may belong to facial feature data, then the outline data will be filtered out.
[0111] Of course, the above-described method for correcting outline data is merely an example. In implementing this embodiment, other methods for correcting outline data can be set according to actual circumstances, and this embodiment does not impose any limitations on this. Furthermore, besides the above-described method for correcting outline data, those skilled in the art can also employ other methods for correcting outline data as needed, and this embodiment does not impose any limitations on this either.
[0112] Step 1033: Configure a score representing the strength of the outline data for each animation.
[0113] Generally, stronger stroke data is characterized by larger length, larger maximum width, and darker color. Therefore, this embodiment analyzes the stroke data of each animation based on one or more features representing intensity, quantifies it, and obtains a score representing the strength of the stroke data.
[0114] In one embodiment of the present invention, step 1033 may include the following steps:
[0115] Step 10331: For each animation data, query the character represented in the animation data by the header data.
[0116] For each animation data, when detecting head data, the head data can be labeled with the character's ID. That is, if the head data of an existing character is detected, the head data can be mapped to the character's ID. If the head data of an unknown character is detected, a new ID can be configured for the unknown character, and the head data can be mapped to the character's ID, thereby realizing the mapping of each head data to each character in the animation data.
[0117] Step 10332: For the same character, calculate the average number of pixels in the outline data.
[0118] For the same character (i.e., the same ID), the number of pixels in each stroke data can be counted, and the average value of that number can be calculated.
[0119] Step 10333: Query the n characters that represent the animation data.
[0120] In this embodiment, n (n is a positive integer) characters can be selected from the animation data based on factors such as plot and popularity, and used as representatives of each character in the animation data.
[0121] In one filtering method, a variable can be configured for each role, denoted as the typical value, which is initially set to 0.
[0122] Queries the frequency of a character's appearance in various scenes (i.e. animation clips) within the animation data.
[0123] If the frequency of a certain character is greater than the third threshold, it means that the character appears frequently and plays an important role in the individual plot of the scene. It can be used as a representative of the scene, and the typical value of the character is incremented by one.
[0124] After traversing all scenes, the typical values of each character are sorted, and the n characters with the highest typical values are selected as the n characters representing the animation data. This method is simple to calculate, and the selected n characters play a relatively important role in the overall plot of all scenes, ensuring the typicality of these n characters. Users' attention will mostly be focused on these n characters, thereby ensuring the accuracy of the evaluation outline data.
[0125] Step 10334: Combine the average values of the n characters into a score representing the strength of the outline data.
[0126] In this embodiment, the average values of n characters can be merged into a score representing the strength of the outline data in a linear or nonlinear manner.
[0127] Taking a linear approach as an example, weights can be assigned to n roles respectively. The weights are positively correlated with the typical value, that is, the larger the typical value, the higher the weight, and vice versa.
[0128] The score representing the strength of the outline data is obtained by adding the product of the average value and the weight of the n roles.
[0129] Step 1034: Mark the k highest-scoring animation data as sketch-style animation data.
[0130] In this embodiment, the scores of each animation data can be sorted, and the k animation data with the highest scores (k is a positive integer) can be marked as sketch-style animation data.
[0131] Step 104: Extract multiple frames of image data from the sketch-style animation data as style sample image data.
[0132] In this embodiment, multiple frames of image data can be extracted from each animation data with a sketch style by sampling methods such as random sampling and uniform sampling. These images are used as samples to train the cartoon sketch image reconstruction network. For the cartoon sketch image reconstruction network, the image data used as samples is the source of the style and can therefore be referred to as style sample image data.
[0133] Furthermore, if reference image data is extracted during the initial screening of animation data in a sketch-like style, this reference image data can be reused as style sample image data.
[0134] Step 105: Train the generative adversarial network into a cartoon sketch image reconstruction network based on the content sample image data and style sample image data.
[0135] In this embodiment, a Generative Adversarial Network (GAN) can be pre-constructed.
[0136] Generally, a Generative Adversarial Network (GAN) consists of a generator and a discriminator. The generator is responsible for generating content based on random vectors; in this embodiment, the content is image data, especially image data with a cartoon sketch style. The discriminator is responsible for determining whether the received content is real; the discriminator usually provides a probability representing the authenticity of the content.
[0137] The generator and discriminator can use different structures. For the function of processing image data, these structures are not limited to artificially designed neural networks, such as convolutional layers, fully connected layers, etc. They can also be neural networks optimized by model quantization methods, neural networks searched for the characteristics of cartoon sketch style by NAS (Neural Architecture Search) methods, etc. This embodiment does not impose any restrictions on them.
[0138] Generative adversarial networks can be classified into the following types based on the different structures of generators and discriminators:
[0139] DCGAN (Deep Convolutional Generative Adversarial Network), CGAN (Conditional Generative Adversarial Network), CycleGAN (Periodic Generative Adversarial Network), CoGAN (Coupled Generative Adversarial Network), ProGAN (Progressive Growth Generative Adversarial Network), WGAN (Wasserstein Generative Adversarial Network), SAGAN (Self-Attention Generative Adversarial Network), BigGAN (Large Generative Adversarial Network), StyleGAN (Style-Based Generative Adversarial Network).
[0140] The generator and discriminator are in an adversarial relationship. This adversarial relationship refers to the alternating training process of the generative adversarial network. Taking the generation of image data with a cartoon sketch style as an example, the generator generates some fake image data and some real image data, which are then fed to the discriminator for judgment. The discriminator learns to distinguish between the two, giving high scores to real image data (i.e., image data with a cartoon sketch style) and low scores to fake image data (i.e., image data without a cartoon sketch style). Once the discriminator can skillfully judge the existing image data, the generator is instructed to continuously generate better fake image data with the goal of obtaining high scores from the discriminator, until it can fool the discriminator. This process is repeated until the discriminator's prediction probability for any image data is close to 0.5, that is, it can no longer distinguish between real and fake image data, at which point training can stop.
[0141] In this embodiment, content sample image data and style sample image face data with cartoon sketch style are used as samples to train the generative adversarial network. The content sample image data is the source of the content, and the style sample image data is the source of the cartoon sketch style. The generative adversarial network is trained in this way, and the trained generative adversarial network is denoted as the cartoon sketch image reconstruction network, so that the cartoon sketch image reconstruction network can be used to reconstruct image data containing cartoon sketch style.
[0142] Furthermore, the samples used to train the generative adversarial network can be paired data, which can improve the performance of the generative adversarial network. However, this requires collecting real-world image data corresponding to the style sample image data. In reality, most style sample image data do not have corresponding real-world image data. Therefore, the generative adversarial network in this embodiment supports training using unpaired data, such as CycleGAN, StyleGAN, etc.
[0143] Taking the "Learning to Cartoonize Using White-box Cartoon Representations" network as an example, this network contains three modules that can divide the original image and style image into three representations:
[0144] 1. Surface characterization
[0145] Surface representations are extracted to represent smooth surfaces in image data. Given image data, weighted low-frequency components can be extracted, where color components and surface texture are preserved while edges, textures, and details are ignored. This can be used to achieve flexible and learnable feature representations of smooth surfaces.
[0146] 2. Structure representation
[0147] Structural representation effectively captures global structural information and sparse color blocks in cel-shaded cartoon style. It extracts segmented regions from the input image data and applies an adaptive coloring algorithm to each segmented region to generate a structural representation. This structural representation mimics the cel-shaded cartoon style, characterized by clear boundaries and sparse color blocks.
[0148] 3. Texture representation
[0149] Texture representations contain the details and edges of the drawn image. The input image data is converted into a single-channel intensity map, where color and brightness are removed, and relative pixel intensity is preserved. Texture representations guide the network to independently learn high-frequency texture details, excluding color and brightness patterns.
[0150] The style of image data output is controlled by balancing the weights of surface representation, structural representation, and texture representation.
[0151] In this embodiment, data from movies whose stories take place in the real world and animation data from multiple animations whose stories take place in the virtual world are collected. Multiple frames of image data are extracted from the movie data as content sample image data. Animation data exhibiting a sketching style is selected from the multiple animation data. Multiple frames of image data are extracted from the sketching-style animation data as style sample image data. Based on the content sample image data and the style sample image data, a generative adversarial network (GAN) is trained into a cartoon sketching image reconstruction network. This network is used to reconstruct image data containing a cartoon sketching style. By selecting a sketching style based on the cartoon style of the animation data, and combining the two, a cartoon sketching style can be obtained. This is used to train the GAN, enabling the cartoon sketching image reconstruction network to reconstruct image data into a cartoon sketching style. Reconstructing the cartoon sketching style is a post-processing step, which maintains the threshold for video data production and reduces the time required for video data production, greatly improving the efficiency of producing cartoon sketching style video data.
[0152] Example 2
[0153] Figure 3 This is a flowchart of an image reconstruction method provided in Embodiment 2 of the present invention. This embodiment is applicable to situations where image data is reconstructed to a cartoon sketch style based on a cartoon sketch image reconstruction network. This method can be executed by an image reconstruction device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 3 As shown, the method includes:
[0154] Step 301: Load the cartoon sketch image to reconstruct the network.
[0155] In a specific implementation, a cartoon sketch image reconstruction network can be pre-trained according to the method described in Embodiment 1 of the present invention, wherein the cartoon sketch image reconstruction network can be used to reconstruct image data containing a cartoon sketch style.
[0156] When applying the cartoon sketch image reconstruction network, the cartoon sketch image reconstruction network and its parameters are loaded into memory for execution.
[0157] Step 302: Obtain the original image data to be reconstructed.
[0158] Generally, cartoon sketch image reconstruction networks have a large structure and consume a lot of resources. They are usually deployed on the server side. The server side can encapsulate the cartoon sketch image reconstruction network into interfaces, plugins, etc., to provide services for reconstructing cartoon sketch styles to users on local area networks or public networks. Users can use clients or browsers to call the interfaces, plugins, etc. to transmit the image data to be reconstructed to the server side. For easy distinction, the image data to be reconstructed is referred to as the original image data.
[0159] Of course, if electronic devices such as personal computers and laptops have sufficient local resources to support the operation of the cartoon sketch image reconstruction network, then the cartoon sketch image reconstruction network can be loaded and run locally on the electronic device. In this case, the original image data of the cartoon sketch style to be reconstructed can be input through command line or other means.
[0160] Step 303: Input the original image data into the cartoon sketch image reconstruction network to reconstruct the target image data containing the cartoon sketch style.
[0161] In this embodiment, the original image data is input into the cartoon sketch image reconstruction network. The cartoon sketch image reconstruction network processes the original image data according to its structure and reconstructs the original image data into new image data containing the cartoon sketch style while maintaining the content of the original image data. This new image data is denoted as the target image data.
[0162] In one example, it will be as follows Figure 4A The original image data shown is input into a cartoon sketch image reconstruction network, and the reconstructed image is as follows. Figure 4B The target image data shown is as follows: Figure 4B The target image data shown is compared to, for example Figure 4A The original image data shown has a more cartoonish character design, highlighting the style of a sketch (especially the outline).
[0163] In this embodiment, a cartoon sketch image reconstruction network is loaded; the original image data to be reconstructed is obtained; and the original image data is input into the cartoon sketch image reconstruction network to reconstruct target image data containing a cartoon sketch style. When training the cartoon sketch image reconstruction network, a sketch style is selected based on the cartoon style presented in the animation data. Combining the two yields the cartoon sketch style, which is then used to train a generative adversarial network. This allows the cartoon sketch image reconstruction network to reconstruct image data into a cartoon sketch style. Reconstructing the cartoon sketch style is a post-processing step, which maintains the threshold for creating video data and reduces the time required for video data production, greatly improving the efficiency of creating video data with a cartoon sketch style.
[0164] Example 3
[0165] Figure 5 This is a flowchart of a video reconstruction method provided in Embodiment 3 of the present invention. This embodiment is applicable to situations where video data is reconstructed to a cartoon sketch style based on a cartoon sketch image reconstruction network. This method can be executed by a video reconstruction device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 5 As shown, the method includes:
[0166] Step 501: Load the cartoon sketch image to reconstruct the network.
[0167] In a specific implementation, a cartoon sketch image reconstruction network can be pre-trained according to the method described in Embodiment 1 of the present invention, wherein the cartoon sketch image reconstruction network can be used to reconstruct image data containing a cartoon sketch style.
[0168] When applying the cartoon sketch image reconstruction network, the cartoon sketch image reconstruction network and its parameters are loaded into memory for execution.
[0169] Step 502: Obtain the original video data that introduces the game.
[0170] In this embodiment, artists can create video data for the game to be promoted, and the content of the video data is used to introduce the game.
[0171] The types of games can include MOBA (Multiplayer Online Battle Arena), RPG (Role-playing game), SLG (Simulation Game), etc., and this embodiment does not limit them.
[0172] In practical implementation, the content of the raw video data can be divided into two main forms: game content and real-life storyline. The storyline can be further divided into the following categories:
[0173] 1. Pseudo-food sharing
[0174] The raw video data includes some food-related materials, which can attract users' attention. Secondly, it incorporates the gameplay of earning money to eat food, while also providing users with a clear goal for playing the game.
[0175] 2. Themes closely related to users' daily lives
[0176] The raw video data closely reflects users' current lifestyles, seamlessly integrating the game's selling points into various aspects of life, such as purchasing in-game items, eating, and buying snacks to earn money through the game. This type of material is relatively simple to produce, with simple scenes and low shooting difficulty. The first half of the material mainly focuses on dialogue between two people, while the second half features game integration segments.
[0177] 3. Situational drama
[0178] The raw video data contains material from situational dramas, and in some cases, celebrities wear costumes from the game to endorse the products. Some of the storylines are quite exaggerated in order to attract users' attention.
[0179] Generally, cartoon sketch image reconstruction networks have a large structure and consume a lot of resources. They are usually deployed on the server side. The server side can encapsulate the cartoon sketch image reconstruction network into interfaces, plugins, etc., to provide services for reconstructing cartoon sketch styles to users on local area networks or public networks. Users can use clients or browsers to call these interfaces, plugins, etc. to transmit the video data to be reconstructed to the cartoon sketch style to the server side. For easy distinction, the video data to be reconstructed to the cartoon sketch style is recorded as the original video data.
[0180] Of course, if electronic devices such as personal computers and laptops have sufficient local resources to support the operation of the cartoon sketch image reconstruction network, then the cartoon sketch image reconstruction network can be loaded and run locally on the electronic device. In this case, the original video data of the cartoon sketch style to be reconstructed can be input through command line or other means.
[0181] Step 503: Input the original image data into the cartoon sketch image reconstruction network to reconstruct the target image data containing the cartoon sketch style.
[0182] In the specific implementation, the original video data contains multiple frames of image data, referred to as the original image data. Each frame of the original image data is input into the cartoon sketch image reconstruction network. The cartoon sketch image reconstruction network processes the original image data according to its structure and reconstructs the original image data into new image data containing the cartoon sketch style while maintaining the content of the original image data, referred to as the target image data.
[0183] Step 504: Replace the original image data with the target image data in the original video data to obtain the target video data.
[0184] In the original video data, the target image data can be replaced with the corresponding original image data to obtain the target video data.
[0185] Afterwards, game-related advertising element data can be added to the target video data to obtain advertising video data. The advertising element data includes the logo (icon) of the platform used to distribute the target game, banner ads, EC (ending clip, which generally contains information about the target game (such as its name, the platform on which the target game is distributed), and so on.
[0186] Advertising video data is published on designated channels (such as news, short videos, novel reading, sports and health, etc.) so that when the client accesses the channel, the advertising video data is pushed to the client for playback. When the user is interested in the game, he downloads the game from the platform that distributes the game.
[0187] In this embodiment, a cartoon sketch image reconstruction network is loaded; the original video data, which introduces the game, is obtained, containing multiple frames of original image data; the original image data is input into the cartoon sketch image reconstruction network to reconstruct target image data containing a cartoon sketch style; the target image data replaces the original image data in the original video data to obtain the target video data. When training the cartoon sketch image reconstruction network, a sketch style is selected based on the cartoon style presented in the animation data. Combining the two yields the cartoon sketch style, which is used to train a generative adversarial network. This allows the cartoon sketch image reconstruction network to reconstruct image data into a cartoon sketch style. Reconstructing the cartoon sketch style is a post-processing step, maintaining the threshold for creating video data and reducing the time required for video data creation, thus greatly improving the efficiency of creating cartoon sketch style video data.
[0188] Example 4
[0189] Figure 6 This is a schematic diagram of the structure of a training device for a cartoon sketch image reconstruction network provided in Embodiment 4 of the present invention. Figure 6 As shown, the device includes:
[0190] The video data acquisition module 601 is used to acquire movie data where the story takes place in the real world and animation data where the story takes place in the virtual world.
[0191] The content sample image data extraction module 602 is used to extract multiple frames of image data from the movie data as content sample image data.
[0192] Animation data filtering module 603 is used to filter animation data in a sketch style from multiple animation data sets;
[0193] The style sample image data extraction module 604 is used to extract multiple frames of image data from the animation data that is in a sketch style, as style sample image data.
[0194] A generative adversarial network training module 605 is used to train the generative adversarial network into a cartoon sketch image reconstruction network based on the content sample image data and the style sample image data. The cartoon sketch image reconstruction network is used to reconstruct image data containing a cartoon sketch style.
[0195] In one embodiment of the present invention, the content sample image data extraction module 602 is further configured to:
[0196] The movie data is divided into multiple movie segments using independent scenes as segmentation nodes;
[0197] In each of the movie segments, one frame of image data is extracted at each preset first time interval as content sample image data.
[0198] In one embodiment of the present invention, the animation data filtering module 603 is further configured to:
[0199] Multiple frames of image data are extracted from each animation data set and used as reference image data.
[0200] Identify outline data representing the sketch style from the reference image data;
[0201] For each animation data segment, a score representing the strength of the outline data is configured;
[0202] The k highest-scoring animation data are marked as sketch-style animation data.
[0203] In one embodiment of the present invention, the animation data filtering module 603 is further configured to:
[0204] Using independent scenes as the dividing nodes, each animation data is divided into multiple animation segments;
[0205] In each of the animation segments, one frame of image data is extracted at each preset second time interval as reference image data.
[0206] In one embodiment of the present invention, the animation data filtering module 603 is further configured to:
[0207] Detect head data containing hair data in the reference image data;
[0208] The header data is amplified.
[0209] Binarization is performed on the magnified header data to distinguish between black and white;
[0210] The binarized header data is subjected to erosion processing;
[0211] Perform dilation processing on the eroded header data;
[0212] Black pixels are detected in the inflated head data to obtain outline data that characterizes the sketch style;
[0213] The outline data is corrected using at least one of the area and coordinates.
[0214] In one embodiment of the present invention, the animation data filtering module 603 is further configured to:
[0215] Face detection is performed in the reference image data to obtain the original detection bounding boxes that identify the face data;
[0216] The original detection frame is expanded along the horizontal and vertical directions respectively to cover the hair data;
[0217] Extract the data from the original detection box after expansion to obtain raw head data containing hair data.
[0218] In one embodiment of the present invention, the animation data filtering module 603 is further configured to:
[0219] Determine the first size of the head data before magnification and the second size of the head data after magnification;
[0220] Calculate the ratio between the first dimension and the second dimension;
[0221] The coordinates of the head data before magnification are obtained by taking the integer part of the product between the magnified coordinates and the scale.
[0222] The pixels located at the coordinates of the header data before magnification are assigned colors to the pixels at the coordinates of the header data after magnification.
[0223] In one embodiment of the present invention, the animation data filtering module 603 is further configured to:
[0224] Query the red, green, and blue components of each pixel in the enlarged header data;
[0225] If the red component is less than or equal to the first threshold, the green component is less than or equal to the first threshold, and the blue component is less than or equal to the first threshold, then the pixel is set to black.
[0226] If at least one of the following conditions is met: the red component is greater than the first threshold, the green component is greater than the first threshold, and the blue component is greater than the first threshold, then the pixel is set to white.
[0227] In one embodiment of the present invention, the animation data filtering module 603 is further configured to:
[0228] For each stroke data belonging to an independent connected region, calculate the area of the stroke data;
[0229] If the area is less than or equal to the second threshold, the outline data is retained;
[0230] If the area is greater than the second threshold, then the outline data is filtered out;
[0231] And / or,
[0232] The region composed of facial key points representing the five facial features, recorded during the query and detection of the head data;
[0233] For each stroke data belonging to an independent connected region, the coordinates of the stroke data are compared with the region;
[0234] If the outline data is located outside the area, the outline data is retained;
[0235] If the outline data is located in the area, then the outline data is filtered out.
[0236] In one embodiment of the present invention, the animation data filtering module 603 is further configured to:
[0237] For each animation data set, query the character represented by the header data within the animation data;
[0238] For the same character, the average number of pixels in the outline data is calculated.
[0239] Query the animation data to find the n characters that represent them;
[0240] The average values corresponding to the n characters are combined into a score representing the strength of the outline data.
[0241] In one embodiment of the present invention, the animation data filtering module 603 is further configured to:
[0242] Configure typical values for each of the aforementioned roles;
[0243] Query the frequency of the character appearing in various scenes of the animation data;
[0244] If the frequency of a certain role is greater than the third threshold, then the typical value of the role is incremented by one.
[0245] The n characters with the highest typical values are selected as the n characters represented by the animation data.
[0246] In one embodiment of the present invention, the animation data filtering module 603 is further configured to:
[0247] Each of the n roles is assigned a weight, and the weight is positively correlated with the typical value;
[0248] The product of the average value of the n characters and the weight is added together to obtain a score representing the strength of the outline data.
[0249] The training device for the cartoon sketch image reconstruction network provided in this embodiment of the invention can execute the training method for the cartoon sketch image reconstruction network provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the training method for the cartoon sketch image reconstruction network.
[0250] Example 5
[0251] Figure 7 This is a schematic diagram of an image reconstruction device provided in Embodiment 5 of the present invention. Figure 3 As shown, the device includes:
[0252] A reconstruction network loading module 701 is used to load a cartoon sketch image reconstruction network trained by the method according to any embodiment of the present invention;
[0253] The raw image data acquisition module 702 is used to acquire the raw image data to be reconstructed;
[0254] The target image data generation module 703 is used to input the original image data into the cartoon sketch image reconstruction network to reconstruct target image data containing cartoon sketch style.
[0255] The image reconstruction apparatus provided in this embodiment of the invention can execute the image reconstruction method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the image reconstruction method.
[0256] Example 6
[0257] Figure 8 This is a schematic diagram of a video reconstruction device provided in Embodiment Six of the present invention. Figure 3 As shown, the device includes:
[0258] A reconstruction network loading module 801 is used to load a cartoon sketch image reconstruction network trained by the method according to any embodiment of the present invention;
[0259] The raw video data acquisition module 802 is used to acquire raw video data containing game introduction content, wherein the raw video data contains multiple frames of raw image data;
[0260] The target image data generation module 803 is used to input the original image data into the cartoon sketch image reconstruction network to reconstruct target image data containing cartoon sketch style;
[0261] The target video data generation module 804 is used to replace the original image data with the target image data in the original video data to obtain target video data.
[0262] In one embodiment of the present invention, it further includes:
[0263] An advertising video data generation module is used to add game-related advertising elements to the target video data to obtain advertising video data.
[0264] The advertising video data publishing module is used to publish the advertising video data on a designated channel, so that when the client accesses the channel, the advertising video data is pushed to the client for playback.
[0265] The video reconstruction apparatus provided in this embodiment of the invention can execute the video reconstruction method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the video reconstruction method.
[0266] Example 7
[0267] Figure 9 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0268] like Figure 9 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0269] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0270] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as training methods for cartoon sketch image reconstruction networks, image reconstruction methods, or video reconstruction methods.
[0271] In some embodiments, the training method, image reconstruction method, or video reconstruction method for a cartoon sketch image reconstruction network can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the training method, image reconstruction method, or video reconstruction method for a cartoon sketch image reconstruction network described above can be performed. Alternatively, in other embodiments, processor 11 can be configured by any other suitable means (e.g., by means of firmware) to perform the training method, image reconstruction method, or video reconstruction method for a cartoon sketch image reconstruction network.
[0272] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0273] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0274] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0275] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0276] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0277] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0278] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0279] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A training method for a cartoon sketch image reconstruction network, characterized in that, include: Collect data from movies whose stories take place in the real world, and data from multiple animated films whose stories take place in the virtual world; Extract multiple frames of image data from the movie data as content sample image data; Multiple frames of image data are extracted from each animation data set and used as reference image data. Detect head data containing hair data in the reference image data; The header data is amplified. Binarization is performed on the magnified header data to distinguish between black and white; The binarized header data is subjected to erosion processing; Perform dilation processing on the eroded header data; Black pixels are detected in the inflated head data to obtain outline data that characterizes the sketch style; The outline data should be corrected using at least one of the area and coordinates. For each animation data set, query the character represented by the header data within the animation data; For the same character, calculate the average number of pixels in the outline data; Query the animation data to find the n characters that represent them; The average values corresponding to the n characters are linearly combined into a score representing the strength of the outline data; The k highest-scoring animation data are marked as sketch-style animation data; Extract multiple frames of image data from the animation data that are in a sketch style, and use them as style sample image data; Based on the content sample image data and the style sample image data, a generative adversarial network is trained to become a cartoon sketch image reconstruction network, which is used to reconstruct image data containing a cartoon sketch style.
2. The method according to claim 1, characterized in that, The step of extracting multiple frames of image data from the movie data as content sample image data includes: The movie data is divided into multiple movie segments using independent scenes as segmentation nodes; In each of the movie segments, one frame of image data is extracted at each preset first time interval as content sample image data.
3. The method according to claim 1, characterized in that, The step of extracting multiple frames of image data from each animation data set as reference image data includes: Using independent scenes as the dividing nodes, each animation data is divided into multiple animation segments; In each of the animation segments, one frame of image data is extracted at each preset second time interval as reference image data.
4. The method according to claim 1, characterized in that, The step of detecting head data containing hair data in the reference image data includes: Face detection is performed in the reference image data to obtain the original detection bounding boxes that identify the face data; The original detection frame is expanded along the horizontal and vertical directions respectively to cover the hair data; Extract the data from the original detection box after expansion to obtain raw head data containing hair data.
5. The method according to claim 1, characterized in that, The amplification process performed on the header data includes: Determine the first size of the head data before magnification and the second size of the head data after magnification; Calculate the ratio between the first dimension and the second dimension; The coordinates of the head data before magnification are obtained by taking the integer part of the product between the magnified coordinates and the scale. The pixels located at the coordinates of the header data before magnification are assigned colors to the pixels at the coordinates of the header data after magnification.
6. The method according to claim 1, characterized in that, The binarization process performed on the magnified header data to distinguish between black and white includes: Query the red, green, and blue components of each pixel in the enlarged header data; If the red component is less than or equal to the first threshold, the green component is less than or equal to the first threshold, and the blue component is less than or equal to the first threshold, then the pixel is set to black. If at least one of the following conditions is met: the red component is greater than the first threshold, the green component is greater than the first threshold, and the blue component is greater than the first threshold, then the pixel is set to white.
7. The method according to claim 1, characterized in that, The correction of the outline data by at least one of the area used and the coordinates includes: For each stroke data belonging to an independent connected region, calculate the area of the stroke data; If the area is less than or equal to the second threshold, the outline data is retained; If the area is greater than the second threshold, then the outline data is filtered out; And / or, The region composed of facial key points representing the five facial features, recorded during the query and detection of the head data; For each stroke data belonging to an independent connected region, the coordinates of the stroke data are compared with the region; If the outline data is located outside the area, the outline data is retained; If the outline data is located in the area, then the outline data is filtered out.
8. The method according to claim 1, characterized in that, The querying of the n characters represented in the animation includes: Configure typical values for each of the aforementioned roles; Query the frequency of the character appearing in various scenes of the animation data; If the frequency of a certain role is greater than the third threshold, then the typical value of the role is incremented by one. The n characters with the highest typical values are selected as the n characters represented by the animation data.
9. The method according to claim 8, characterized in that, The step of linearly fusing the average values corresponding to the n characters into a score representing the strength of the outline data includes: Each of the n roles is assigned a weight, and the weight is positively correlated with the typical value; The product of the average value of the n characters and the weight is added together to obtain a score representing the strength of the outline data.
10. An image reconstruction method, characterized in that, include: Load the cartoon sketch image reconstruction network trained by the method according to any one of claims 1-9; Obtain the original image data to be reconstructed; The original image data is input into the cartoon sketch image reconstruction network and reconstructed into target image data containing a cartoon sketch style.
11. A video reconstruction method, characterized in that, include: Load the cartoon sketch image reconstruction network trained by the method according to any one of claims 1-9; The content to be acquired is the original video data introducing the game, which contains multiple frames of original image data; The original image data is input into the cartoon sketch image reconstruction network and reconstructed into target image data containing a cartoon sketch style; The target image data is replaced in the original video data to obtain the target video data.
12. The method according to claim 11, characterized in that, Also includes: Add game-related advertising elements to the target video data to obtain advertising video data; The advertising video data is published on a designated channel so that when a client accesses the channel, the advertising video data is pushed to the client for playback.
13. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, which is executed by the at least one processor to enable the at least one processor to perform the training method of the cartoon sketch image reconstruction network according to any one of claims 1-9, the image reconstruction method according to claim 10, or the video reconstruction method according to any one of claims 11-12.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the training method of the cartoon sketch image reconstruction network according to any one of claims 1-9, the image reconstruction method according to claim 10, or the video reconstruction method according to any one of claims 11-12.
Citation Information
Patent Citations
Image binarization method and device, electronic device and storage medium
CN109325497A
A sketch generation method and device
CN109472838A
Real scene image cartoonalization processing method and device, computer equipment and storage medium
CN111696028A
Sketch style scene rendering method and device and storage medium
CN113935893A