Image style conversion method based on deformable convolution and accurate feature matching

By adopting deformable convolution and precise feature matching techniques in Chinese character image style conversion, the existing methods generate unnatural and structural distortion problems when dealing with complex stroke shapes, and high-quality Chinese character image style conversion is achieved.

CN120147112APending Publication Date: 2025-06-13XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510255716.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing Chinese character image style conversion method cannot be flexibly adjusted when processing stroke shapes with different scales and angles, resulting in unnatural fonts and distorted structures. At the same time, existing methods only focus on low-order statistics of feature distribution, and cannot accurately match complex feature distributions, affecting the generation effect.

Method used

Using an image style conversion method based on deformable convolution and precise feature matching, the content and style features are extracted by building a variable convolution network and an accurate feature distribution matching module, and feature alignment is achieved through the empirical cumulative distribution function to generate the final output image.

Benefits of technology

It significantly enhances the modeling ability of local features, achieves accurate feature alignment, avoids style distortion, and the generated images have natural and realistic visual effects, improving the quality and practicality of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147112A_ABST
    Figure CN120147112A_ABST
Patent Text Reader

Abstract

The invention discloses an image style conversion method based on deformable convolution and accurate feature matching in the technical field of image processing, and the method comprises the following steps: obtaining various styles of Chinese character images, constructing a style font data set, and dividing the style font data set into a training set and a test set; building a font image style conversion model; using the training set to train the font image style conversion model to optimize the model, and then using the test set to test the optimized model to obtain a Chinese character image style conversion model; and inputting the target content image and the style reference image into the Chinese character image style conversion model to obtain a converted target style Chinese character image. When the font image style conversion model is built, the modeling capability of the network for local features can be effectively enhanced by building the variable convolutional network, especially for the features which are irregularly and non-linearly distributed in space. Conversion between different font styles is achieved, and content consistency and style accuracy are kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image style conversion method based on deformable convolution and precise feature matching. Background Art

[0002] As the core carrier of Chinese culture, Chinese characters with their complex structures and diverse writing styles have irreplaceable value in cultural inheritance, artistic creation, and technological applications. The Chinese character image style conversion technology can, through the deep integration of computer vision and deep learning, convert one font style (such as regular script) into another style (such as running script or cursive script). This technology can restore endangered fonts, perform style transfer on damaged ancient books or inscriptions, restore the original brushstrokes and ink color changes of historical fonts, and prevent cultural heritages from being lost due to physical degradation. At the same time, it can also assist in personalized learning, automatically generate Chinese character examples in different calligraphy styles (such as the Yan style and the Liu style), help learners intuitively compare stroke features (such as pen tips and turning strengths), and quickly master writing norms. It can accurately transfer the calligraphy styles of famous masters (such as the running script brushstrokes of Wang Xizhi), and combine modern design requirements to generate fonts with both traditional charm and a sense of the times, expanding the boundaries of artistic expression.

[0003] Currently, the methods for font image style conversion mainly utilize traditional convolution operations. However, the traditional convolution operation has limited ability to model the geometric deformation of images. The fixed size of the convolution kernel makes it impossible to flexibly adjust for stroke shapes with different scales and angles, resulting in the generated fonts appearing unnatural in some scenarios and the structure of the generated fonts being distorted. Specifically, problems such as stroke adhesion and structural loss occur in the converted Chinese character images.

[0004] When dealing with the representation of style information, the existing font style conversion methods mainly rely on traditional feature distribution matching. Usually, it is assumed that font features conform to a Gaussian distribution, and the style conversion is achieved by matching the mean and standard deviation. However, the existing methods have significant limitations, that is, they only focus on the low-order statistics of the feature distribution. The feature distribution of actual data contains more complex statistical information. Simply aligning the mean and standard deviation cannot accurately match this complex distribution, especially when dealing with fonts with complex feature distributions, which is not flexible enough. This leads to the situation that when facing real font styles, the previous methods for processing two features may not be able to effectively maintain the overall form of the template style features, thus affecting the generation effect. Summary of the Invention

[0005] The purpose of the present invention is to provide an image style conversion method based on deformable convolution and precise feature matching, which can achieve the expansion of Chinese character font styles, can also generate fonts that conform to the target style for Chinese characters that cannot be directly converted, and at the same time ensure the consistency of content and the accuracy of style.

[0006] The technical solution adopted by the present invention is an image style conversion method based on deformable convolution and exact feature matching, including the following steps: Step 1: Obtain multiple style Chinese character images, construct a style font data set, and divide the style font data set into a training set and a test set; Step 2: Build a font image style conversion model; Step 3: Use the training set to train the font image style conversion model to optimize the model, and then use the test set to test the optimized model to obtain a Chinese character image style conversion model; Step 4: Input the target content image and the style reference image into the Chinese character image style conversion model to obtain the converted target style Chinese character image.

[0007] The characteristics of the present invention also lie in: In Step 1, the data ratio of the training set to the test set is 8:2.

[0008] The font image style conversion model includes a data set generation module, a content feature extraction network, a style feature extraction network, an exact feature distribution matching module, and an image decoder; The input content image and style image are adjusted to a unified resolution by the data set generation module; the content feature extraction network extracts the content features of the content image, including stroke position, angle, connection relationship between strokes, and geometric shape; the style feature extraction network extracts the style features of the style image; the exact feature distribution matching module fuses the content features and style features, and the image decoder restores the fused features to generate the final output image.

[0009] The content feature extraction network includes a residual block, a deformable convolution module, and an instance normalization module. The deformable convolution module includes three deformable convolution modules, namely dcn1, dcn2, and dcn3. An instance normalization layer is introduced after each deformable convolution module. The steps for the content feature extraction network to extract content are as follows: The input target image is subjected to content feature extraction through the residual block to obtain a content feature map; The content feature map passes through the dcn1 module to extract the shallow features of the image. After the shallow features are normalized by the instance normalization layer, they are input into the ReLU activation function. After non-linear transformation, the positive values are retained and the negative values are set to zero to obtain the output image 1, and the image 1 is copied as skip1; The image 1 then passes through the dcn2 module to extract the middle features of the image. After the middle features are normalized by the instance normalization layer, they are input into the ReLU activation function. After non-linear transformation, the positive values are retained and the negative values are set to zero to obtain the output image 2, and the image 2 is copied as skip2; Finally, Image 2 passes through the dcn3 module to extract the deep features of the image. After the deep features are normalized by the instance normalization layer, they are input into the ReLU activation function. After non-linear transformation, the positive values are retained and the negative values are set to zero, and finally the content features are output. F C 。

[0010] The residual block consists of two standard convolutional layers. The convolutional kernel size of each convolutional layer is 3×3, the stride is 1, the padding is 1, and the number of output channels is the same as the number of input channels. An instance normalization layer is introduced after each convolutional layer, and after the instance normalization layer, the ReLU activation function is applied.

[0011] The dcn1 module uses the following parameters: the input channel is 3, the output channel is 64, the convolutional kernel size is 7×7, the stride is 1, and the padding is 3; The dcn2 module uses the following parameters: the input channel is 64, the output channel is 128, the convolutional kernel size is 4×4, the stride is 2, and the padding is 1; The dcn3 module uses the following parameters: the input channel is 128, the output channel is 256, the convolutional kernel size is 4×4, the stride is 2, and the padding is 1.

[0012] The style feature extraction network adopts the VGG16 network architecture. After the style image passes through the convolutional layer and pooling layer of the VGG16 network architecture, the style feature map is obtained. After flattening the style feature map into a one-dimensional vector, it is input into the fully connected layer and mapped into a one-dimensional style feature. F S 。

[0013] The specific configuration of the VGG16 network architecture is that convolutional layer 1 uses 64 convolutional kernels with a size of 3×3, the stride is set to 1, the padding is 1, and the activation function is ReLU; pooling layer 1 uses a 2×2 max pooling layer with a stride of 2; Convolutional layer 2 uses 128 convolutional kernels with a size of 3×3, the stride is set to 1, the padding is 1, and the activation function is ReLU; pooling layer 2 uses a 2×2 max pooling layer with a stride of 2; Convolutional layer 3 uses 256 convolutional kernels with a size of 3×3, the stride is set to 1, the padding is 1, and the activation function is ReLU; pooling layer 2 uses a 2×2 max pooling layer with a stride of 2; Convolutional layer 4 uses 512 convolutional kernels with a size of 3×3, the stride is set to 1, the padding is 1, and the activation function is ReLU; pooling layer 2 uses a 2×2 max pooling layer with a stride of 2; Convolutional layer 5 uses 512 convolutional kernels with a size of 3×3, the stride is set to 1, the padding is 1, and the activation function is ReLU; pooling layer 2 uses a 2×2 max pooling layer with a stride of 2.

[0014] The precise feature distribution matching module adopts the empirical cumulative distribution function, and the specific fusion steps are as follows: Process the content feature and the style feature into images with the same batch size, number of channels, width, and height; Flatten the content feature into a one-dimensional vector, then sort the pixel values of each channel in ascending order to obtain the sorted pixel values V c and the corresponding sorting indexes I c , calculate I c the inverse sorting index of, to obtain I cr ; Sort the pixel values of each channel of the style feature in ascending order to obtain the sorted pixel values V s ; According to I cr the sorting order of, map V s to the corresponding position of V c , calculate the difference between the two, and add the difference to V c to obtain the adjusted content feature. The adjusted content feature is aligned with the style feature map in terms of distribution, and the aligned fusion feature map F cn is obtained.

[0015] The restoration steps are as follows: Use the residual block to extract features from the fused feature map and output the feature map O1; Perform upsampling and convolution processing on O1 in sequence to obtain the feature map O3; Fuse O3 with skip2, and then use deformable convolution processing to output the feature map O4; Perform upsampling on O4 again, and extract features through the convolutional layer to obtain the feature map O5; Concatenate O5 and skip1 in the channel dimension and then perform deformable convolution operation to generate the feature map O6; After performing convolution operation on O6, adjust the number of channels to the number of channels of the input image. The output image retains the content features of the content image and incorporates the style features of the style image.

[0016] The beneficial effects of the present invention are: 1. Enhance the ability to model local features By constructing a deformable convolutional network, the present invention significantly enhances the network's ability to model local features, especially for features that exhibit irregular and non-linear distributions in space. This enhancement enables the network to more accurately capture the details of each local region when processing font generation tasks, thereby improving the fineness and accuracy of the generated images.

[0017] 2. Precise feature alignment ability The present invention achieves feature alignment by precisely matching the empirical cumulative distribution function (eCDF) of image features. The eCDF describes the overall distribution of features by statistically sorting feature values, rather than simply matching the mean and variance. This method can precisely align the feature distributions of two images, ensuring that the style of the style image (such as texture, hue) can be accurately transferred to the generated image, rather than simply modifying the color or texture level of the image.

[0018] 3. Avoiding style distortion The present invention effectively avoids the problem of style distortion by achieving precise alignment and effective fusion of content and style features through eCDF. Compared with traditional style transfer methods, the present invention can prevent excessive injection or loss of style, thereby maintaining the authenticity of the image. The generated image not only has the characteristics of the target style but also can maintain a natural and realistic visual effect, significantly improving the quality and practicality of image generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a schematic flow diagram of the present invention; Figure 2 is a schematic diagram of the deformable convolutional neural network in Embodiment 1 of the present invention; Figure 3 is a schematic diagram of the style feature extraction module in Embodiment 1 of the present invention; Figure 4 is a style conversion network diagram in Embodiment 1 of the present invention; Figure 5 is a schematic diagram of the result of image style conversion in Embodiment 2 of the present invention.

[0020] Figure 6 is a schematic diagram of the result of image style conversion in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0021] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0022] Embodiment 1 An image style conversion method based on deformable convolution and precise feature matching according to the present invention, as Figure 1 shown, specifically includes the following steps: Step 1: Obtain Chinese character images of multiple styles and construct a style font dataset In this embodiment, a total of 600 font files are collected from multiple public font libraries, covering a variety of font styles (such as regular script, running script, cursive script, etc.). Each font file contains 990 Chinese character glyphs, including common Chinese characters, Chinese character radicals, and some rare Chinese characters, to ensure the language diversity of the training data and the generalization ability of the model. The resolution of each character image is uniformly 80×80 pixels to meet the unity requirements of the model input.

[0023] For model training and evaluation, the dataset is divided into a training set and a test set according to a ratio of 8:2. Among them, the training set is used for optimizing the parameters of the model, and the test set is used for evaluating the performance of the model. When training and evaluating the model with the training set and the test set, the Chinese character images of different styles inside are randomly divided into style images and content images as the input of the model. The total number of samples is 594,000 images, ensuring the sufficiency of the training process and the reliability of the evaluation results. The above data collection provides a high-quality and diverse training data foundation for the font style conversion model of this embodiment.

[0024] Step 2: Build a font image style conversion model The font image style conversion model includes five parts: a dataset generation module, a content feature extraction network, a style feature extraction network, an exact feature distribution matching module, and an image decoder; Step 2.1: The dataset generation module is used to uniformly normalize the input images to ensure data consistency and the stability of model training. Specifically, this module adjusts the input content images and style images to a unified resolution, generating standard images of size 80×80 pixels. Through this module, a series of font images with consistent resolution and unified format can be generated, providing high-quality input data for subsequent model training and font style conversion.

[0025] Step 2.2: Construct a content feature extraction network. The content feature extraction network is used to extract key content features from the input content images, including feature information such as stroke positions, angles, connection relationships between strokes, and geometric shapes. The specific steps are as follows: Step 2.2.1: Construct residual blocks. Residual blocks are used to extract the content features of the input content images. They are constructed by stacking two convolutional layers, normalization layers, and activation functions in a loop, and can effectively extract the deep information of the input images.

[0026] Specifically, in the residual block, two standard convolutional layers are used to extract image features. The convolutional kernel size of each convolutional layer is 3×3, the stride is 1, and the padding is 1, ensuring that the size of the output feature map is the same as the input. At the same time, the number of output channels is the same as the number of input channels to maintain the number of channels of the feature map unchanged. After each convolutional layer, an instance normalization layer is introduced. This instance normalization layer calculates the mean and variance independently for each channel to reduce the differences between samples and ensure the stability of network training. After the instance normalization layer, the ReLU (Rectified Linear Unit) activation function is applied to introduce non-linear transformation into the network and enhance its expressive ability. Finally, through the residual connection, the input image is added to the feature map after two convolutions, thereby realizing the retention and flow of information and effectively avoiding the problem of gradient disappearance.

[0027] Step 2.2.2, construct the deformable convolution module: The deformable convolution module includes three deformable convolution modules (dcn1, dcn2, dcn3) for solving the problem of fixed position in traditional convolution operations. The three deformable convolution modules are respectively used to process feature extraction tasks at different levels. The structure and parameters of each deformable convolution module are as follows: dcn1 module: The input channel is 3, the output channel is 64, the convolutional kernel size is 7×7, the stride is 1, and the padding is 3. The dcn1 module is used to process the initial input image and extract shallow features of the image, such as basic information like edges and textures; dcn2 module: The input channel is 64, the output channel is 128, the convolutional kernel size is 4×4, the stride is 2, and the padding is 1. On the basis of dcn1, the dcn2 module further extracts the middle-level features of the image, such as the basic shapes of strokes and local structure information; dcn3 module: The input channel is 128, the output channel is 256, the convolutional kernel size is 4×4, the stride is 2, and the padding is 1. The dcn3 module is used to extract the deep features of the image and capture more complex stroke connection relationships, geometric shapes, and global structure information.

[0028] The key to variable convolutional networks (such as deformable convolutional networks) lies in their ability to dynamically adjust the sampling positions of convolutional kernels to adapt to the complex and varying geometric structures in the input images. For Chinese character images, this is particularly important because Chinese characters are composed of numerous strokes, and there are specific layouts, angles, and connection relationships between the strokes. For Chinese characters, such fixed sampling is difficult to capture the curvature, inclination, or complex connections of the strokes, and may overlook key structural details such as the endpoints, turning points, or overlapping regions of the strokes. In deformable convolution, an offset learned by the network is attached to each fixed sampling position. These offsets are predicted from the input features through an additional convolutional layer. These offsets can cause the sampling points of the convolutional kernel to move from the original fixed positions to more informative regions. For example, if a certain stroke is inclined, the network will learn an offset consistent with the inclination direction, causing the sampling point to fall on the stroke rather than deviating from it. By learning the offsets, the network can dynamically adjust the sampling points to accurately capture the center and edge positions of the strokes. In this way, during the convolution operation, the angle and inclination information of the strokes can be more sensitively captured. Since the offsets are adaptively adjusted for local regions, the network can not only extract the information of individual strokes but also integrate the information of each local region to form an understanding of the overall geometric shape of the entire character. This is particularly crucial for Chinese characters with complex structures.

[0029] Generally speaking, by introducing learnable offsets during the convolution process, variable convolutional networks enable the sampling positions to be automatically adjusted according to the input data, thereby taking into account both local details and global structures during feature extraction. This mechanism enables the computer to more accurately identify the positions, angles, connection relationships, and overall geometric shapes of each stroke in Chinese characters, thus greatly enhancing the model's understanding and representation capabilities for complex Chinese character structures.

[0030] Step 2.2.3, Instance Normalization Module To reduce the impact of the batch size on the training process, this embodiment introduces an Instance Normalization (IN) module after each deformable convolution module. Instance normalization reduces the differences between samples by performing independent normalization processing on the feature maps of each sample, thereby helping the model better learn features and improving the stability of the training process. At the same time, the ReLU (Rectified Linear Unit) activation function is used to introduce a non-linear transformation into the network to further enhance the model's expressive ability. The specific implementation is as follows: The input image x enters the dcn1 module. After deformable convolution, the feature map x1 is output, as Figure 2As shown, x1 is normalized by the instance normalization layer (IN1) to obtain the feature map x2. Then, x2 is input into the ReLU activation function. After non-linear transformation, the positive values are retained and the negative values are set to zero to obtain x3. At the same time, x3 is copied as skip1 for skip connection to help the decoder fuse the low-level and high-level feature maps; x3 is input into the dcn2 module and undergoes a convolution operation to output the feature map x4. x4 then passes through the instance normalization layer (IN2) to obtain the second-layer normalized feature map x5. x5 undergoes non-linear transformation through the ReLU activation function. After non-linear transformation, the positive values are retained and the negative values are set to zero to further enhance the expression ability of the model, obtaining x6. At the same time, x is copied as skip2 for skip connection to help the decoder fuse the low-level and high-level feature maps; Similarly, x6 is input into the dcn3 module and undergoes a convolution operation to output the feature map x7. x7 then passes through the instance normalization layer (IN3) to obtain the third-layer normalized feature map x8. x8 undergoes non-linear transformation through the ReLU activation function. After non-linear transformation, the positive values are retained and the negative values are set to zero, enabling the network to capture more complex patterns and finally output the content features F C .

[0031] 2.3. Construction of the style feature extraction network The core of constructing the style feature extraction network is to extract the image style features through the VGG16 network architecture. The extraction process of the style feature extraction network is as follows: The style image with a size of (3, 80, 80) is input into VGG16. In the first-layer convolution operation, 64 convolution kernels with a size of 3×3 are used, the stride is set to 1, and the padding is 1. The purpose of this setting is to keep the spatial size of the feature map unchanged during the convolution operation. After this layer of convolution operation, the size of the output feature map becomes (64, 80, 80), that is, the number of channels of the image increases from 3 to 64, while the width and height remain at 80 unchanged. Subsequently, the convolved feature map is input into the ReLU activation function for processing to obtain the activated feature map. Then, a maximum pooling operation is performed using a 2×2 pooling window with a stride of 2. This operation reduces the width and height of the feature map by half, and after pooling, the size of the feature map becomes (64, 40, 40).

[0032] The second convolutional operation uses 128 convolutional kernels of size 3×3, with a stride of 1 and a padding of 1. After this convolutional operation, the output feature map has a size of (128, 40, 40). Then, this feature map is passed through the ReLU activation function to obtain the activated feature map. Again, a max pooling operation is performed with a 2×2 pooling window and a stride of 2, further reducing the width and height of the feature map by half, and the size of the pooled feature map becomes (128, 20, 20).

[0033] The third convolutional operation uses 256 convolutional kernels of size 3×3, with a stride of 1 and a padding of 1. After the convolutional operation, the output feature map has a size of (256, 20, 20). After passing through the ReLU activation function to obtain the activated feature map, a max pooling operation is performed with a 2×2 pooling window and a stride of 2, and the width and height of the feature map are reduced by half again, becoming (256, 10, 10).

[0034] The fourth convolutional operation uses 512 convolutional kernels of size 3×3, with a stride of 1 and a padding of 1. The size of the output feature map after convolution is (512, 10, 10). After passing through the ReLU activation function to obtain the activated feature map, a max pooling operation is performed with a 2×2 pooling window and a stride of 2, reducing the width and height of the feature map by half, becoming (512, 5, 5).

[0035] The fifth convolutional operation still uses 512 convolutional kernels of size 3×3, with a stride of 1 and a padding of 1. The size of the output feature map after convolution is (512, 5, 5). After passing through the ReLU activation function to obtain the activated feature map, finally, a max pooling operation is performed with a 2×2 pooling window and a stride of 2, and the width and height of the feature map are reduced by half again, finally obtaining a feature map with a size of (512, 2, 2).

[0036] After a series of operations on the convolutional and pooling layers of the above VGG16, a feature map with a size of (512, 2, 2) is obtained. In order to be able to input these features into the subsequent fully connected layer for further processing, a flattening operation needs to be performed on this feature map, that is, flattening the feature map with a size of (512, 2, 2) into a one-dimensional vector with a length of 512×2×2 = 2048.

[0037] The flattened feature vector is then input into the fully connected layer, where the feature vector containing 2048 input features is mapped into a one-dimensional style feature with a length of 4096 F S , as Figure 3 shown, this style feature F SIt can more comprehensively and effectively represent the style information of the input style image, providing a key data basis for subsequent related tasks based on style features.

[0038] Step 2.4: Construct an exact feature distribution matching module The exact feature distribution matching module is used to achieve feature distribution alignment between the content image and the style image, thereby ensuring the effective fusion of content features F C and style features F S as shown in Figure 4 . The specific implementation steps are as follows: Step 2.4.1: Perform size verification and preprocessing on the input content features F C and style features F S . First, verify the input content features F C and style features F S to ensure that their shapes are consistent. Specifically, the content features F C and style features F S must have the same batch size, number of channels, width, and height to ensure that subsequent feature distribution alignment operations are performed in the same feature space, avoiding calculation errors or performance degradation caused by size mismatches.

[0039] Step 2.4.2: Perform exact feature matching according to the empirical cumulative distribution function (eCDF). Specifically: First, sort the content features F C . Flatten F C into a one-dimensional vector, and then sort the pixel values of each channel in ascending order. After sorting, obtain the sorted pixel values V c and the corresponding sorting indices I c , I c indicating the relative position of each pixel value in the entire feature map. Then, by calculating the inverse sorting index of I c , obtain I cr , which is used to rearrange the pixel values of the style features F S so that they are consistent with the content features F CAlign in order. Through the sorting operation, the pixel values of the content features are arranged from smallest to largest, providing a reference basis for the subsequent distribution adjustment of the style features F S of the distribution adjustment.

[0040] Step 2.4.3: Sort the style features F S Perform a sorting process. For the style features F S Ascendingly sort the pixel values of each channel to obtain the sorted pixel values V s .

[0041] Step 2.4.4: Feature matching and adjustment. Adjust the content features according to the sorting result of the style features F S . First, according to I cr the sorting order, map the pixel values of the style features V s to the corresponding positions of the content feature pixel values V c . Then, calculate the difference between the content feature pixel values V c and the style feature pixel values V s , and add this difference to the content feature pixel values V c to obtain the adjusted content features. The adjusted content feature map has been aligned with the style feature map in terms of distribution, obtaining the aligned fused feature map F cn . This process realizes the precise matching of feature distributions, making the distribution of the content features closer to that of the style features.

[0042] Step 2.5, Construct an image decoder. The decoder module is used to gradually restore the extracted feature map to a high-resolution output image, while fusing the content features and style features. The specific implementation steps are as follows: Step 2.5.1, Feature extraction of the residual block. Use the same residual block as in Step 2.2.1 for feature extraction. Pass the feature map F cn through the convolutional layer for feature extraction, and through skip connections (add the input and output feature maps to maintain the flow of information and avoid the problem of gradient disappearance). The output of the residual block is the new deep feature map O1. This step helps to capture the deep semantic information of the input image and lays the foundation for subsequent decoding operations.

[0043] Step 2.5.2, Upsampling and Convolution Processing Layer by Layer. Next, the feature map O1 restores its spatial resolution step by step through an upsampling operation. First, through an upsampling operation, the size of O1 is doubled to obtain O2. Then, the convolutional layer is used to process O2. The convolutional kernel size is 5*5, the stride is 1, and the padding is 2 for further feature extraction. Through this operation, the local information of the image is enhanced, and the feature map O3 is generated. At this time, the size of the image has been expanded and contains more spatial information.

[0044] Step 2.5.3, Deformable Convolution and Feature Fusion. The feature map O3 will be fused with the skip connection skip2. Specifically, by concatenating O3 and skip2 in the channel dimension, a new feature map is formed. Then, the deformable convolution is used to process the concatenated feature map. The convolutional kernel size of the deformable convolution is 5×5, the stride is 1, and the padding is 2. In this process, the feature map is further enhanced through the deformable convolution operation, and the output feature map O4 is obtained. This operation helps to strengthen the detailed expression of the image.

[0045] Step 2.5.4, Continue Upsampling and Convolution Processing. The feature map O4 is upsampled again, and feature extraction is performed through the convolutional layer. The convolutional kernel size of the convolutional layer is 7×7, the stride is 1, and the padding is 3. This convolutional operation helps to further improve the resolution of the image and enhance the feature information of the image. In this step, the size of the feature map is expanded again, and the new deep feature map O5 is obtained.

[0046] Step 2.5.5, Second Deformable Convolution and Fusion Processing. The feature map O5 is concatenated with the skip connection skip1 in the channel dimension to form a new feature map.

[0047] Perform a deformable convolution operation on the concatenated feature map. The convolutional kernel size of the deformable convolution is 3×3, the stride is 1, and the padding is 1. By adaptively adjusting the shape of the convolutional kernel, the image features are optimized, and the feature map O6 is generated.

[0048] Step 2.5.6, Generate the Final Output Image. Finally, perform a convolution operation on O6. The convolution parameters use a convolutional kernel size of 3×3, a stride of 1, and a padding of 1 to adjust the number of channels of the feature map to the number of channels of the input image. At this time, after multiple upsampling and convolution operations, the final output image is generated. The size of this image is the same as that of the input image, and while retaining the content features of the input content image, it incorporates the style features of the style image.

[0049] Step 3, Input the style font dataset into the font image style conversion model for training. Update the network parameters through the training set and optimize the parameter model through the validation set. Then, use the test set to test the optimized model to obtain the Chinese character image style conversion model; Step 4: Input the target content image and the style reference image into the trained network model to obtain the converted image.

[0050] Specifically, input the target content image and the style reference image into the trained Chinese character image style conversion model to obtain the Chinese character image with the target style.

[0051] Embodiment 2 Take Figure 5 the source image in Figure 5 as the target content image, convert its font to the target image, and perform style conversion using the Chinese character image style conversion model obtained in Embodiment 1. The model outputs

[0052] the experimental result image in Take Figure 6 the source image in Figure 6 as the target content image, convert its font to the target image, and perform style conversion using the Chinese character image style conversion model obtained in Embodiment 1. The model outputs

[0053] The style transfer method of the present invention can efficiently achieve the conversion between different font styles and maintain the consistency of content and the accuracy of style.

Claims

1. Image style transfer method based on deformable convolution and precise feature matching, characterized in that: The following steps are involved: Step 1, obtain Chinese character images of various styles, construct a style font dataset, and divide the style font dataset into a training set and a test set; Step 2: Build a font image style transfer model; Step 3, using the training set to train the font image style transfer model to optimize the model, and then using the test set to test the optimized model to obtain the Chinese character image style transfer model; Step 4: Input the target content image and the style reference image into the Chinese character image style transfer model to obtain the converted target style Chinese character image.

2. The image style transfer method based on deformable convolution and precise feature matching according to claim 1, characterized in that: The data ratio of the training set and the test set in step 1 is 8:

2.

3. The image style transfer method based on deformable convolution and precise feature matching according to claim 1, characterized in that: The font image style conversion model includes a data set generation module, a content feature extraction network, a style feature extraction network, an accurate feature distribution matching module and an image decoder; The input content image and style image are adjusted to a uniform resolution by the dataset generation module; the content feature extraction network extracts the content features of the content image, including stroke positions, angles, connection relationships between strokes, and geometric shapes; the style feature extraction network extracts the style features of the style image; The precise feature distribution matching module fuses the content features and style features, and the image decoder restores the fused features to generate the final output image.

4. The image style transfer method based on deformable convolution and precise feature matching according to claim 3, characterized in that: The content feature extraction network includes a residual block, a deformation convolution module and an instance normalization module. The deformation convolution module includes three deformation convolution modules, namely dcn1, dcn2, and dcn3. An instance normalization layer is introduced after each deformation convolution module. The content feature extraction network performs the following steps for content extraction: The input target image is subjected to content feature extraction through the residual block to obtain a content feature map; The content feature map is passed through the dcn1 module to extract the shallow features of the image. The shallow features are normalized by the instance normalization layer and then input into the ReLU activation function. After nonlinear transformation, the positive values ​​are retained and the negative values ​​are set to zero to obtain the output image 1. Image 1 is copied as skip1. Image 1 is then passed through the dcn2 module to extract the middle-level features of the image. The middle-level features are normalized by the instance normalization layer and then input into the ReLU activation function. After nonlinear transformation, the positive values ​​are retained and the negative values ​​are set to zero to obtain the output image 2. Image 2 is copied as skip2. Image 2 is finally passed through the dcn3 module to extract the deep features of the image. The deep features are normalized by the instance normalization layer and then input into the ReLU activation function. After nonlinear transformation, the positive values ​​are retained and the negative values ​​are set to zero, and finally the content features are output. F C .

5. The image style transfer method based on deformable convolution and precise feature matching according to claim 4, characterized in that: The residual block includes two standard convolutional layers, each of which has a convolution kernel size of 3×3, a stride of 1, a padding of 1, and the number of output channels is the same as the number of input channels. An instance normalization layer is introduced after each convolutional layer, and a ReLU activation function is applied after the instance normalization layer.

6. The image style transfer method based on deformable convolution and precise feature matching according to claim 4, characterized in that: The dcn1 module uses the following parameters: 3 input channels, 64 output channels, 7×7 convolution kernel size, 1 step size, and 3 padding; The dcn2 module uses the following parameters: 64 input channels, 128 output channels, 4×4 kernel size, 2 strides, and 1 padding; The dcn3 module uses the following parameters: input channels are 128, output channels are 256, convolution kernel size is 4×4, stride is 2, and padding is 1.

7. The image style transfer method based on deformable convolution and precise feature matching according to claim 3, characterized in that: The style feature extraction network adopts the VGG16 network architecture. The style image passes through the convolution layer and pooling layer of the VGG16 network architecture to obtain a style feature map. The style feature map is flattened into a one-dimensional vector and then input into the fully connected layer to be mapped into a one-dimensional style feature. F S .

8. The image style transfer method based on deformable convolution and precise feature matching according to claim 7, characterized in that: The specific configuration of the VGG16 network architecture is that the convolution layer 1 uses 64 convolution kernels of size 3×3, sets the step size to 1, the padding to 1, and the activation function to ReLU; the pooling layer 1 uses a 2×2 maximum pooling layer with a step size of 2; Convolutional layer 2 uses 128 convolution kernels of size 3×3, sets the stride to 1, padding to 1, and activation function to ReLU; pooling layer 2 uses a 2×2 maximum pooling layer with a stride of 2; Convolution layer 3 uses 256 convolution kernels of size 3×3, sets the step size to 1, padding to 1, and activation function to ReLU; pooling layer 2 uses a 2×2 maximum pooling layer with a step size of 2; Convolutional layer 4 uses 512 convolution kernels of size 3×3, with a step size of 1, padding of 1, and activation function of ReLU; pooling layer 2 uses a 2×2 maximum pooling layer with a step size of 2; Convolution layer 5 uses 512 convolution kernels of size 3×3, sets the step size to 1, padding to 1, and activation function to ReLU; pooling layer 2 uses a 2×2 maximum pooling layer with a step size of 2.

9. The image style transfer method based on deformable convolution and precise feature matching according to claim 3, characterized in that: The precise feature distribution matching module adopts the empirical cumulative distribution function, and the specific fusion steps are: Process content features and style features into images with the same batch size, number of channels, width, and height; After flattening the content features into a one-dimensional vector, the pixel values ​​of each channel are sorted in ascending order to obtain the sorted pixel values V c And the corresponding sort index I c ,calculate I c The reverse sort index of I cr ; Sort the pixel values ​​of each channel of the style feature in ascending order to obtain the sorted pixel values V s ; according to I cr The sort order will be V s Map to V c , calculate the difference between the two, and add the difference to V c The adjusted content features are aligned with the style feature map in terms of distribution, and the aligned fusion feature map is obtained. F cn .

10. The image style transfer method based on deformable convolution and precise feature matching according to claim 4, characterized in that: The recovery steps are: Use the residual block to extract features from the fused feature map and output feature map O1; Upsample and convolve O1 in sequence to obtain feature map O3; O3 is fused with skip2, and then deformable convolution is used to output feature map O4; O4 is upsampled again, and feature extraction is performed through the convolution layer to obtain feature map O5; Concatenate O5 and skip1 in the channel dimension and perform a deformable convolution operation to generate feature map O6; After the convolution operation on O6, the number of channels is adjusted to the number of channels of the input image. The output image retains the content features of the content image and incorporates the style features of the style image.