A method and device for character stroke segmentation and vector reconstruction generation
Through ControlNet and stroke vector gallery technology, the problem of unable to generate stroke-by-stroke Chinese character vectors in the existing technology is solved, and efficient and standardized font design is achieved to meet the development needs of family fonts and variable fonts.
Patent Information
- Application Number
- CN202411869728.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-18
AI Technical Summary
The prior art is difficult to generate stroke-by-stroke Chinese character vector graphics, which cannot meet the needs of family fonts and variable font design, and cannot ensure the consistent order of key points, consistent direction of Bezier curves and standardized control point positions.
By using ControlNet to disassemble strokes, a stroke vector gallery is constructed, and vector images are optimized through affine transformation and differentiable rasterization technology to generate stroke-by-stroke font vectors that meet the needs of font design.
It greatly reduces the time and energy of designers in hand-made Chinese characters vector fonts, improves design efficiency and consistency, and the generated font reconstructed vector graphics meet professional requirements and is more standardized.
Smart Images

Figure CN119337821B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of the combination of computer vision and computer graphics, and specifically relates to a method and device for character font stroke segmentation and vector reconstruction generation. Background Art
[0002] In the traditional Chinese font design process, designers need to follow a unified font style and manually draw the vector glyphs of each Chinese character. This process is extremely time-consuming and labor-intensive due to the huge number of Chinese characters.
[0003] At present, deep learning technology has become an important means to improve design efficiency in the font industry. For example, the patent application with publication number CN116975344A discloses a Chinese font generation method and device based on Stable Diffusion. It trains a deep learning model, learns the style of specific font samples, and generates a set of Chinese character bitmaps of the same style, allowing designers to further outline curves on this basis, which can significantly reduce time costs and improve design efficiency.
[0004] In order to meet the needs of designers for further modification and the production of font families or variable fonts, in many cases, each stroke of each glyph in the Chinese font needs to be an independent path. However, existing methods cannot generate stroke-by-stroke Chinese character glyph vector maps, and usually only generate vector outlines of the entire Chinese character. In addition, the glyph vector maps required for family fonts and variable font design also require consistent key point order, consistent Bezier curve direction, and standardized control point positions. These problems cannot be solved using existing technologies. Summary of the invention
[0005] In view of the above, the present invention proposes a font stroke splitting and vector reconstruction generation method and device, which learns the stroke splitting method from the Chinese character glyphs pre-designed by the designer, establishes a stroke vector library, and in the subsequent vector glyph generation, disassembles the strokes from the automatically generated Chinese character bitmap, matches the stroke vector map, and optimizes the stroke-by-stroke vector map with the aforementioned Chinese character bitmap as a reference. Furthermore, with another related Chinese character bitmap as a reference, the generated vector map can also be adjusted to a new glyph with isomorphic outlines to meet the development needs of family fonts and variable fonts.
[0006] To achieve the above-mentioned purpose of the invention, an embodiment provides a method for character stroke segmentation and vector reconstruction generation, comprising the following steps:
[0007] After constructing sample data and using the sample data to train ControlNet, ControlNet is used to decompose the overall font bitmap into strokes to obtain a stroke decomposition bitmap, and the stroke decomposition bitmap is separated according to the strokes to obtain each stroke bitmap;
[0008] Construct a stroke vector map library, extract the stroke shape features of each stroke bitmap, match the stroke vector map with the stroke vector map library based on the stroke shape features, perform affine transformation on the stroke vector map to align its bounding box with the stroke bitmap, and then obtain the font reconstruction vector map;
[0009] For the stroke vector map or the font vector map, a differentiable rasterization method is used to make the stroke contour points closer to the stroke bitmap or the font overall bitmap, thereby completing the optimization of the font reconstruction vector map.
[0010] Preferably, constructing sample data and using the sample data to train ControlNet includes:
[0011] Prepare the fonts designed by the designer in advance, color all the strokes of each Chinese character with different colors, ensure that the strokes of the same color do not overlap, and generate a stroke decomposition bitmap. Introduce a random mechanism to generate multiple stroke decomposition bitmaps for the same Chinese character for data enhancement. The overall font bitmap of each Chinese character is compared with the corresponding colored stroke decomposition bitmap. Figure 1 One-to-one correspondence, forming image pairs, and introducing description text for the image pairs, forming a sample data set consisting of image pairs and description text, where the description text includes font style and weight;
[0012] The sample data is used to train ControlNet, with the overall font bitmap in the image pair as the input of ControlNet, the descriptive text as the guiding condition of ControlNet, and the colored stroke disassembly bitmap as the true value label of ControlNet. The ControlNet is supervised and trained to optimize the ControlNet parameters.
[0013] Preferably, when the stroke decomposition bitmap is separated according to the strokes to obtain each stroke bitmap, it is specifically separated according to color. First, the stroke decomposition bitmap is read, and then the red, green and blue channels are extracted respectively, and color threshold processing is performed on each channel to divide the pixels into a specific range; then, for each color threshold of the red, green and blue channels, a full-white image is initialized, and the pixels at the corresponding positions below the corresponding color threshold are set to black. The set black-and-white image is saved only when there are non-white pixels, and finally a stroke decomposition graph separated by multiple channels is obtained, and the connected contours are further extracted for each stroke decomposition graph to obtain a stroke bitmap stroke by stroke.
[0014] Preferably, constructing a stroke vector library includes:
[0015] The stroke shape features of each stroke vector map are extracted, and the stroke shape features are stored together with the Chinese character to which the stroke belongs, the bitmap path corresponding to the stroke, the stroke size, and the contour points to form a structured stroke vector map library.
[0016] Preferably, center distance, Hu distance, or neural network is used to extract the stroke shape features of each stroke vector map and stroke bitmap.
[0017] Preferably, matching the stroke vector map with the stroke vector map library based on the stroke shape features, and performing affine transformation on the stroke vector map so that its bounding box is aligned with the stroke bitmap includes:
[0018] According to the stroke shape features, multiple stroke vector maps with high similarity to each stroke bitmap are retrieved from the stroke vector map library, and the contour points of the stroke vector map are initially fitted to the stroke bitmap shape according to the position and size of the bounding box of the stroke bitmap using the affine method, thereby obtaining the font reconstruction vector map. The similarity can be measured using similarity indicators such as L1, L2, and MSE (mean square error).
[0019] Preferably, a differentiable rasterization method is used for the stroke vector map or the font vector map to make the stroke contour points closer to the stroke bitmap or the font overall bitmap, including:
[0020] Taking the overall font bitmap of the same weight generated by the diffusion model as the target, differentiable rasterization is used to optimize the font vector map, so that the stroke contour points are close to the overall font bitmap; or taking the stroke bitmap obtained by disassembling the overall font bitmap as the target, differentiable rasterization is used to optimize the stroke vector map, so that the stroke contour points are close to the stroke bitmap.
[0021] Preferably, the method also includes: using a diffusion model to generate a font overall bitmap of a target font weight based on the font overall bitmap, and performing vector reconstruction for the font overall bitmap of the target font weight, including using the font reconstruction vector map corresponding to the original font weight as a guide, and optimizing through differentiable rasterization to obtain the font reconstruction vector map corresponding to the target font weight.
[0022] Preferably, the vector reconstruction of the entire font bitmap of the target font weight further includes:
[0023] ControlNet is used to decompose the overall font bitmap of the target weight into strokes to obtain a stroke decomposition bitmap, and the stroke decomposition bitmap is separated according to the strokes to obtain each stroke bitmap. Based on the font reconstruction vector map of the original weight, each stroke bitmap of the target weight is matched with the font reconstruction vector map of the original weight through the center coordinates and aspect ratio of the bounding box, and then the adjustment is guided stroke by stroke to achieve adjustment and optimization of all stroke contour points to obtain the font reconstruction vector map of the target weight.
[0024] Specifically, according to the bounding box of the stroke bitmap, the coordinate system of the stroke vector image object matched by the stroke vector image library is rescaled and offset to align it with the coordinate system of the stroke bitmap obtained by stroke decomposition; the contour points of the matched stroke vector image object are rendered into an image stroke by stroke or whole character using a differentiable rasterizer, the loss between the rendered image and the stroke shape in the target stroke bitmap is calculated, and the stroke contour points are adjusted using an optimization method to make the matched stroke contour points closer to the target stroke shape.
[0025] To achieve the above-mentioned purpose of the invention, the embodiment of the present invention further provides a font stroke splitting and vector reconstruction generating device, comprising:
[0026] The font stroke decomposition module is used to construct sample data and train ControlNet with the sample data, then use ControlNet to decompose the overall font bitmap into strokes to obtain a stroke decomposition bitmap, and separate the stroke decomposition bitmap according to the strokes to obtain each stroke bitmap;
[0027] A vector reconstruction module is used to construct a stroke vector map library, extract the stroke shape features of each stroke bitmap, match the stroke vector map with the stroke vector map library based on the stroke shape features, perform affine transformation on the stroke vector, align its bounding box with the stroke bitmap, and then obtain the font reconstructed vector map;
[0028] The vector optimization module is used to use a differentiable rasterization method for the stroke vector map or the font vector map to make the stroke contour points closer to the stroke bitmap or the font overall bitmap, thereby completing the optimization of the font reconstruction vector map.
[0029] To achieve the above-mentioned purpose of the invention, an embodiment further provides a computing device, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned font stroke splitting and vector reconstruction generation method.
[0030] To achieve the above-mentioned purpose of the invention, the embodiment further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the above-mentioned font stroke splitting and vector reconstruction generation method is implemented.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] The present invention utilizes ControlNet to perform stroke disassembly, matching and reconstruction, and batch-generates font vectors that meet the requirements of font design, thereby greatly reducing the time and energy of designers in manually making Chinese character vector glyphs.
[0033] When performing vector reconstruction, the present invention performs matching optimization reconstruction based on the constructed stroke vector map library so that the obtained font reconstruction vector map meets the professional requirements of font design and is more standardized. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0035] Figure 1 is a flow chart of a method for font stroke splitting and vector reconstruction generation provided by an embodiment;
[0036] Figure 2 It is a flowchart of a method for font stroke splitting and vector reconstruction generation provided by an embodiment;
[0037] Figure 3 is a schematic diagram of font vector reconstruction provided by an embodiment;
[0038] Figure 4 is a schematic diagram of font vector reconstruction of multiple word weights provided by an embodiment;
[0039] Figure 5 It is a structural schematic diagram of a font stroke splitting and vector reconstruction generation device provided in an embodiment. DETAILED DESCRIPTION
[0040] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0041] The inventive concept of the present invention is to solve the pain points in the process of Chinese font design due to the huge number of Chinese characters and the time-consuming and labor-intensive manual drawing. The diffusion model is used to generate the overall bitmap of the Chinese character font of the target weight, and ControlNet and vector database technology are further used to disassemble the strokes and reconstruct them into stroke vector diagrams in which each stroke is an independent path. This technical route greatly reduces the workload of designers in manually outlining strokes, and improves design efficiency and consistency through the construction and reuse of stroke vector diagram libraries, thereby greatly improving the efficiency of Chinese font design while ensuring design quality.
[0042] like Figure 1 As shown, the embodiment provides a method for font stroke splitting and vector reconstruction generation, comprising the following steps:
[0043] S1, after constructing sample data and using the sample data to train ControlNet, ControlNet is used to decompose the overall font bitmap into strokes to obtain a stroke decomposition bitmap, and the stroke decomposition bitmap is separated according to the strokes to obtain each stroke bitmap.
[0044] The ControlNet used in the embodiment is a deep neural network structure that adds additional guiding conditions as input to the Stable Diffusion model to control the model output. Its training set consists of paired pictures and corresponding prompt words. The present invention constructs a sample data set with customized stroke overlap area annotations, and no additional annotation of stroke types is required when constructing the sample data set. ControlNet trained with sample data can predict the stroke division relationship and the shape of the overlapping area, extract strokes from the overall font bitmap, and convert them into vectors.
[0045] In the embodiment, constructing a sample data set and using the sample data set to train ControlNet includes:
[0046] Prepare the fonts designed in advance by the designer, color all the strokes of each Chinese character with different colors, ensure that the strokes of the same color do not overlap, and generate a stroke decomposition bitmap. Specifically, firstly, extract the vector outline information of each Chinese character by traversing the Chinese character font files in the specified directory; then, rasterize the complete Chinese character vector outline into a binary bitmap and save it as a reference image with a size of 512*512 pixels, where the pixel value of the text part is 0 and the background part is 255; then, for each stroke, generate an independent stroke bitmap according to its vector outline and save it as a bitmap, construct a Boolean intersection matrix to determine the overlap between the stroke bitmaps, and by comparing the stroke bitmaps, determine whether they have pixel overlap, that is, whether the strokes overlap, and store the results in the Boolean intersection matrix; finally, according to the intersection Boolean matrix, assign each stroke to a different level by level, give each level a unique color for easy distinction, and ensure that the strokes on the same level do not overlap;
[0047] In the specific implementation, six channel lists are set to represent the levels. In order to ensure that there are no more than six layers, check whether the number of layers exceeds six layers. If it exceeds, it is recorded as an error character and skipped. The stroke contours are concentrated in the first three channel lists, so the first three channel lists are set as the main level list, and the last three are the secondary level lists. The main level list and the secondary level list are sorted in descending order according to the number of stroke contours. Subsequently, the contours of different levels are randomly combined to generate four label images and save them. To achieve this, the elements in the main level and the secondary level are randomly arranged, and the two lists are spliced. The stroke contours of each level are colored and superimposed on the label image. Among them, the RGB color value of the first level is [170, 0,0], the RGB color value of the second level is [0, 170, 0], the RGB color value of the third level is [0, 0, 170], the RGB color value of the fourth level is [85, 0, 0], the RGB color value of the fifth level is [0, 85, 0], and the RGB color value of the sixth level is [0, 0, 85].
[0048] A random mechanism is introduced to generate multiple stroke decomposition bitmaps for the same Chinese character for data enhancement. The overall font bitmap of each Chinese character is compared with the corresponding colored stroke decomposition bitmap. Figure 1 One-to-one correspondence, forming an image pair and saving it, while introducing a description text for the image pair. A description text tag can be added to the file name of the image pair for identification, wherein the added description text includes font style and weight; thus forming a sample data set consisting of image pairs and description text;
[0049] The sample data is used to train ControlNet, with the overall font bitmap in the image pair as the input of ControlNet, the description text as the guiding condition of ControlNet, and the colored stroke disassembly bitmap as the true value label of ControlNet. The ControlNet is supervised and trained to optimize the ControlNet parameters. At the same time, the trained ControlNet can also be fine-tuned on a targeted sample data set containing only the fonts to be extracted.
[0050] In the embodiment, for the image data input to ControlNet, the pixel values of the image are uniformly adjusted to between 0 and 1, and the pixel value type of the image is converted to a 32-bit floating point type for preprocessing. In the process of training ControlNet, the number of training worker threads is specified to be 4, the batch size is 4, and the learning rate is 10. -5 During the training process, the loss value during the training process and the training results of each training cycle will be saved in the log, and the 10 pre-trained weights with the best effect will be saved according to the loss value obtained during training.
[0051] After the training is completed, ControlNet is used to decompose the overall font bitmap into strokes to obtain a stroke decomposition bitmap. Specifically, the overall font bitmap to be decomposed is input into the trained ControlNet, and the stroke decomposition bitmap is obtained through forward reasoning. Then, the stroke decomposition bitmap is separated according to the strokes to obtain each stroke bitmap. Since the strokes are distinguished by color, the stroke decomposition bitmap can be separated according to coloring to obtain the stroke bitmap of each stroke. The specific process is to traverse the directory of the stroke decomposition bitmap and perform the following processing on each image file in the directory:
[0052] Create a directory to save the processed image and read the image data; convert the image data into a one-dimensional array with each row representing a pixel value; use the np.unique() function to get the unique pixel values and their frequency of occurrence; extract the R, G, and B color channels of the image. The pixel values of the red, green, and blue color channels are graded as follows: values less than 25 are set to 0, values between 25 and 127 are set to 85, values between 127 and 213 are set to 170, and values greater than 213 are set to 255. Reassemble the processed red, green, and blue channels into a color image and save it. Create a new black-and-white image, process the image according to the levels of the red, green, and blue channels, and save the visualization results of each channel, that is, perform six times, set the pixels with red channel values less than or equal to 85, the pixels with green channel values less than or equal to 85, the pixels with blue channel values less than or equal to 85, the pixels with red channel values 170, the pixels with green channel values 170, and the pixels with blue channel values 170 to black, and save the other pixels as white, and save a total of six bitmaps.
[0053] S2, construct a stroke vector map library, extract the stroke shape features of each stroke bitmap, and match the stroke vector map with the stroke vector map library based on the stroke shape features, perform affine transformation on the stroke vector map to align its bounding box with the stroke bitmap, and then obtain the font reconstructed vector map.
[0054] In the embodiment, a stroke vector diagram library for stroke vector reconstruction is pre-constructed, and the specific construction process is: pre-define the method of extracting the stroke shape features of each stroke vector diagram, including center distance, Hu distance, or neural network, wherein the neural network may include autoencoder (AE), variational autoencoder (VAE), contrastive learning and other methods. When using a neural network, it is necessary to define and train the neural network, use the weights obtained by training the stroke image data set, retain the model weights of the encoder part, and use the same model during feature extraction, load the saved model weights for inference to extract the stroke shape features of the stroke image. Then, the stroke shape features of each stroke vector diagram are extracted using the above method, and the stroke shape features are stored together with the Chinese character to which the stroke belongs, the bitmap path corresponding to the stroke, the stroke size, and the contour points to form a stroke vector diagram library. To facilitate retrieval, the stroke vector diagram library uses vector database technology, so that when matching, it can be retrieved from it according to the similarity of the aforementioned feature vectors.
[0055] In the specific implementation, the font file designed in advance by the designer is converted into a JSON format file, which contains the outline information of the vector font, including the whole body unit number (UPM), the baseline and the coordinates of the outline points. Each file is read and processed one by one. For each file, its content is opened and read, parsed into a JSON data structure, and the whole body unit number (UPM), the baseline and the outline information are extracted. Each outline in the outline information is extracted and converted into a list containing point coordinate information. Then, each outline is rasterized and converted into an image in RGBA format. After the rasterization process is completed, the outline will be adjusted according to the target whole body unit number (UPM) to generate a standardized image. In this embodiment, the stroke shape feature can be a central moment, and for each outline, its standardized central moment is calculated and extracted. In addition, the pixel count of the image, that is, the number of white pixels in the image, is calculated to describe the stroke size. The bounding box of each outline is also calculated to describe the position and area of the outline in the image. For each processed outline, the generated standardized image will be saved and the corresponding image path will be recorded. These feature data and the standardized image path are added to the original JSON data structure, while the original contour data is deleted. After processing is completed, the updated JSON file, i.e. the stroke vector library, is saved to the specified directory.
[0056] In the embodiment, Figure 3 As shown, the stroke vector is reconstructed based on the stroke vector library. Specifically, for each stroke bitmap obtained in S1, the stroke shape features of the stroke bitmap are extracted using the same method as that of constructing the stroke vector library. Then, the stroke vector is matched with the stroke vector library based on the stroke shape features. An affine transformation is performed on the stroke vector to align its bounding box with the stroke bitmap, thereby obtaining the font reconstructed vector map.
[0057] In the specific implementation, the central moment feature of the font stroke is selected as the stroke shape feature for retrieval matching, and in order to improve the retrieval speed, the vector database is used to create an index based on the L2 distance of the center distance feature for fast search of similar stroke vector maps. In the data preparation stage, the stroke data is loaded, the Faiss feature index is constructed, and the necessary save directory is created. Specifically, the stroke vector map set in the specified directory is loaded first; then the image file is read, the connected area is extracted as a single stroke, the color is inverted and normalized, the central moment feature is calculated, and the result is stored in the dictionary; then the center distance feature is used to find the most similar candidate stroke vector map in the Faiss feature index, and the best candidate is further selected by L1 loss. Finally, the stroke bitmap in the font folder is processed, and the most similar stroke is obtained by matching with the stroke vector map library, and the matching result is saved. According to the stroke shape feature, multiple stroke vector maps with high similarity to each stroke bitmap are retrieved in the stroke vector map library, and the contour points of the stroke vector map are preliminarily fitted to the stroke bitmap shape according to the position and size of the bounding box of the stroke using the affine method.
[0058] S3, using a differentiable rasterization method for the stroke vector map or the font vector map, so that the stroke contour points are closer to the stroke bitmap or the font overall bitmap, and the font reconstruction vector is optimized.
[0059] In the embodiment, the font vector image may be generated by the designer when designing the glyph, or may be composed of stroke vector images in a stroke vector image library, and the specific details are not limited thereto. The font vector image may be used to reconstruct the vector of the entire font.
[0060] Specifically, a differentiable rasterization method is used, with the overall font bitmap of the same weight generated by the diffusion model as the target, and the font vector map is optimized using differentiable rasterization to guide the stroke contour points to be closer to the overall font bitmap, thereby obtaining the optimized font reconstruction vector map. It should be noted that this process optimization is for overall font bitmaps of other styles that have the same weight as the original overall font bitmap.
[0061] Specifically, a differentiable rasterization method is used, with the stroke bitmap obtained from the overall font bitmap as the target, to guide the contour points to further fit into the corresponding stroke bitmap shape, and all stroke contour points are adjusted and optimized to obtain an optimized font reconstructed vector map.
[0062] In the embodiment, a method for reconstructing multiple weights of font vectors of other weights of a current font is also provided, such as Figure 2 and Figure 4As shown, a diffusion model is used to generate a font overall bitmap of a target weight based on the font overall bitmap, and vector reconstruction is performed on the font overall bitmap of the target weight, including taking the font reconstruction vector map corresponding to the original font weight as a guide, and optimizing the font reconstruction vector map corresponding to the target weight through differentiable rasterization.
[0063] like Figure 4 As shown, the overall font bitmap of the target weight is vectorized and reconstructed, which also includes: using ControlNet to decompose the overall font bitmap of the target weight into strokes to obtain a stroke decomposition bitmap, and separating the stroke decomposition bitmap according to the strokes to obtain each stroke bitmap, taking the original weight font reconstruction vector map as a reference, matching each stroke bitmap of the target weight with the original weight font reconstruction vector map through the center coordinates and the aspect ratio of the bounding box, and then guiding and adjusting each stroke stroke by stroke to achieve adjustment and optimization of all stroke contour points, and obtaining the target weight font reconstruction vector map.
[0064] The method of the present invention uses an existing stroke-by-stroke Chinese font file to generate a sample data set with a stroke disassembly bitmap, and fine-tune and train ControlNet with it. Then, the trained ControlNet is applied to the pre-generated overall font bitmap, and the inferred coloring results are separated to obtain a stroke-by-stroke stroke bitmap; one or more feature extraction methods are selected to construct a stroke vector map library of an existing font; the stroke vector map library is searched according to the aforementioned stroke bitmap, and on the basis of the matched stroke vector map, the stroke vector map is pasted on the stroke bitmap through affine transformation and differentiable rasterization technology; the above operations are performed on each pre-generated stroke bitmap, and finally the font is vectorized and reconstructed to obtain a font reconstructed vector map, which can improve the font design efficiency and the professional standardization of the generated font vector map.
[0065] like Figure 5 As shown, the embodiment also provides a font stroke splitting and vector reconstruction generating device 50, including a font stroke splitting module 51, a vector reconstruction module 52, and a vector optimization module 53, wherein the font stroke splitting module 51 is used to construct sample data and use the sample data to train ControlNet, and then use ControlNet to decompose the overall font bitmap into strokes to obtain a stroke decomposition bitmap, and separate the stroke decomposition bitmap according to the strokes to obtain each stroke bitmap; the vector reconstruction module 52 is used to construct a stroke vector map library, extract the stroke shape features of each stroke bitmap, and match the stroke vector map with the stroke vector map library based on the stroke shape features, perform affine transformation on the stroke vector, align its bounding box with the stroke bitmap, and then obtain the font reconstructed vector map; the vector optimization module 53 is used to use a differentiable rasterization method for the stroke vector map or the font vector map to make the stroke contour points closer to the stroke bitmap or the overall font bitmap, thereby completing the optimization of the font reconstruction vector.
[0066] It should be noted that the font stroke splitting and vector reconstruction generating device provided in the above embodiment should be illustrated by the division of the above functional modules when performing font stroke splitting and vector reconstruction. The above functional distribution can be completed by different functional modules as needed, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the functions described above. In addition, the font stroke splitting and vector reconstruction generating device provided in the above embodiment and the font stroke splitting and vector reconstruction generating method embodiment belong to the same concept. The specific implementation process is detailed in the font stroke splitting and vector reconstruction generating method embodiment, which will not be repeated here.
[0067] Based on the same inventive concept, an embodiment further provides a computing device, including a memory and one or more processors, wherein executable code is stored in the memory, and when the one or more processors execute the executable code, the method for font stroke splitting and vector reconstruction generation is implemented, specifically including the following steps:
[0068] S1, after constructing sample data and using the sample data to train ControlNet, use ControlNet to decompose the overall font bitmap into strokes to obtain a stroke decomposition bitmap, and separate the stroke decomposition bitmap according to the strokes to obtain each stroke bitmap;
[0069] S2, constructing a stroke vector map library, extracting the stroke shape features of each stroke bitmap, and matching the stroke vector map with the stroke vector map library based on the stroke shape features, performing affine transformation on the stroke vector map to align its bounding box with the stroke bitmap, and then obtaining a font reconstruction vector map;
[0070] S3, using a differentiable rasterization method for the stroke vector map or the font vector map, so that the stroke contour points are closer to the stroke bitmap or the font overall bitmap, and the font reconstruction vector map is optimized.
[0071] The computing device provided in the embodiment, in addition to the processor and memory, also includes hardware required for other services such as internal bus, network interface, memory, etc. at the hardware level. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the font stroke splitting and vector reconstruction generation method described in S1-S3 above. Of course, in addition to the software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0072] Based on the same inventive concept, an embodiment further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the above-mentioned font stroke splitting and vector reconstruction generation method is implemented, which specifically includes the following steps:
[0073] S1, after constructing sample data and using the sample data to train ControlNet, use ControlNet to decompose the overall font bitmap into strokes to obtain a stroke decomposition bitmap, and separate the stroke decomposition bitmap according to the strokes to obtain each stroke bitmap;
[0074] S2, constructing a stroke vector map library, extracting the stroke shape features of each stroke bitmap, and matching the stroke vector map with the stroke vector map library based on the stroke shape features, performing affine transformation on the stroke vector map to align its bounding box with the stroke bitmap, and then obtaining a font reconstruction vector map;
[0075] S3, using a differentiable rasterization method for the stroke vector map or the font vector map, so that the stroke contour points are closer to the stroke bitmap or the font overall bitmap, and the font reconstruction vector map is optimized.
[0076] In the embodiment, computer-readable media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data.
[0077] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for character stroke segmentation and vector reconstruction generation, characterized in that: The following steps are involved: After constructing sample data and using the sample data to train ControlNet, ControlNet is used to decompose the overall font bitmap into strokes to obtain a stroke decomposition bitmap, and the stroke decomposition bitmap is separated according to the strokes to obtain each stroke bitmap; wherein, constructing sample data and using the sample data to train ControlNet includes: Prepare fonts designed in advance by designers, color all strokes of each Chinese character with different colors, ensure that strokes of the same color do not overlap, and generate stroke decomposition bitmaps. Introduce a random mechanism to generate multiple stroke decomposition bitmaps for the same Chinese character for data enhancement. The overall font bitmap of each Chinese character corresponds to the corresponding colored stroke decomposition bitmap to form an image pair. At the same time, introduce description text for the image pair to form a sample data set consisting of image pairs and description text, where the description text includes font style and weight; use sample data to train ControlNet, use the overall font bitmap in the image pair as the input of ControlNet, use the description text as the guiding condition of ControlNet, use the colored stroke decomposition bitmap as the true value label of ControlNet, and perform supervised training on ControlNet to optimize ControlNet parameters; construct a stroke vector map library, extract the stroke shape features of each stroke bitmap, and match the stroke vector map with the stroke vector map library based on the stroke shape features, perform affine transformation on the stroke vector map to align its bounding box with the stroke bitmap, and then obtain the font reconstruction vector map; For the stroke vector map or the font vector map, a differentiable rasterization method is used to make the stroke contour points closer to the stroke bitmap or the font overall bitmap, thereby completing the optimization of the font reconstruction vector map.
2. The method for character stroke segmentation and vector reconstruction generation according to claim 1, characterized in that: Build a stroke vector library, including: The stroke shape features of each stroke vector map are extracted, and the stroke shape features are stored together with the Chinese character to which the stroke belongs, the bitmap path corresponding to the stroke, the stroke size, and the contour points to form a stroke vector map library, in which the stroke features are extracted using center distance, Hu distance, or neural network.
3. The method for character stroke segmentation and vector reconstruction generation according to claim 1, characterized in that: Match the stroke vector map based on the stroke shape features and the stroke vector map library, and perform affine transformation on the stroke vector map to align its bounding box with the stroke bitmap, including: According to the stroke shape features, multiple stroke vector maps with high similarity to each stroke bitmap are retrieved in the stroke vector map library. The contour points of the stroke vector map are preliminarily fitted to the stroke bitmap shape using affine transformation according to the position and size of the bounding box of the stroke bitmap, thereby obtaining the font reconstructed vector map.
4. The method for character stroke segmentation and vector reconstruction generation according to claim 1, characterized in that: For the stroke vector map or font vector map, a differentiable rasterization method is used to make the stroke contour points closer to the stroke bitmap or the overall font bitmap, including: Taking the overall font bitmap of the same weight generated by the diffusion model as the target, differentiable rasterization is used to optimize the font vector map, so that the stroke contour points are close to the overall font bitmap; or taking the stroke bitmap obtained by disassembling the overall font bitmap as the target, differentiable rasterization is used to optimize the stroke vector map, so that the stroke contour points are close to the stroke bitmap.
5. The method for character stroke segmentation and vector reconstruction generation according to claim 1, characterized in that: Also includes: A diffusion model is used to generate a font bitmap of a target weight based on the font bitmap, and vector reconstruction is performed on the font bitmap of the target weight, including using the font reconstruction vector map corresponding to the original weight as a guide, and optimizing the font reconstruction vector map corresponding to the target weight through differentiable rasterization.
6. The method for character stroke segmentation and vector reconstruction generation according to claim 5, characterized in that: Vector reconstruction of the entire font bitmap of the target font weight is performed, including: ControlNet is used to decompose the overall font bitmap of the target weight into strokes to obtain a stroke decomposition bitmap, and the stroke decomposition bitmap is separated according to the strokes to obtain each stroke bitmap. Based on the font reconstruction vector map of the original weight, each stroke bitmap of the target weight is matched with the font reconstruction vector map of the original weight through the center coordinates and aspect ratio of the bounding box, and then the adjustment is guided stroke by stroke to achieve adjustment and optimization of all stroke contour points to obtain the font reconstruction vector map of the target weight.
7. A font stroke splitting and vector reconstruction generating device, characterized in that: include: The font stroke decomposition module is used to construct sample data and train ControlNet with the sample data, then use ControlNet to decompose the overall font bitmap into strokes to obtain a stroke decomposition bitmap, and separate the stroke decomposition bitmap according to the strokes to obtain each stroke bitmap; wherein, constructing sample data and training ControlNet with the sample data includes: The overall font bitmap corresponds to the corresponding colored stroke decomposition bitmap one by one to form an image pair, and at the same time, a description text is introduced for the image pair to form a sample data set consisting of image pairs and description texts, where the description text includes font style and weight; training, using the overall font bitmap in the image pair as the input of ControlNet, the description text as the guiding condition of ControlNet, and the colored stroke decomposition bitmap as the true value label of ControlNet, and supervised training of ControlNet to optimize ControlNet parameters; A vector reconstruction module is used to construct a stroke vector map library, extract the stroke shape features of each stroke bitmap, match the stroke vector map with the stroke vector map library based on the stroke shape features, perform affine transformation on the stroke vector, align its bounding box with the stroke bitmap, and then obtain the font reconstructed vector map; The vector optimization module is used to use a differentiable rasterization method for the stroke vector map or the font vector map to make the stroke contour points closer to the stroke bitmap or the font overall bitmap, thereby completing the optimization of the font reconstruction vector map.
8. A computing device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, they are used to implement the font stroke splitting and vector reconstruction generation method described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the method for splitting and reconstructing the font strokes and generating vectors according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Method and device for generating Chinese character library based on Stable Diffusion
CN116975344A
Method and device for generating small sample calligraphy fonts based on diffusion model
CN117669492A