Chinese traditional calligraphy and painting fusion method and device based on enhanced generative network
Through the method based on the enhanced generation network, the generator network of Chinese character images and landscape painting images is pre-trained and style transfer training is carried out, which solves the problems of character distortion and structural damage during calligraphy and painting generation in the existing technology, and realizes the precise integration of Chinese characters and landscape painting and the essential combination of calligraphy and painting art forms.
Patent Information
- Application Number
- CN202510172718.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
When the existing diffusion model generates traditional Chinese calligraphy and painting, it lacks special data for Chinese calligraphy and traditional Chinese painting, resulting in distortion of characters or damage to the structure, and it is difficult to accurately grasp the geometric shape and stroke characteristics of Chinese characters, and fail to achieve the essential fusion of calligraphy and painting art forms.
Using an enhanced generation network method, the generator network is pre-trained by obtaining the Chinese character image data set and landscape painting image data set, and the style transfer training is used to transfer the structural features of the Chinese character image to the target landscape painting image features to generate a fusion image.
The precise fusion of Chinese character images and landscape painting images is achieved, and the original characters and structure of Chinese characters are maintained. The generated fusion image of calligraphy and painting has the precise combination of Chinese character image characteristics and landscape painting images, achieving the essential fusion of calligraphy and painting art forms.
Smart Images

Figure CN120107080A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a method and device for fusing traditional Chinese calligraphy and painting based on an enhanced generation network. Background Art
[0002] Chinese traditional art has long emphasized the concept of "calligraphy and painting have the same origin", believing that calligraphy and painting are inseparable in terms of brushwork, expression of artistic conception and structural layout, and together they constitute an important foundation of Chinese traditional art. As a unique visual symbol system, Chinese characters not only carry the function of language communication, but also contain rich aesthetic values and cultural connotations. Its stroke structure shows significant similarities and complementarity with traditional Chinese paintings in terms of spatial layout and line rhythm. For example, the stroke trend of Chinese characters echoes the contour lines of landscape paintings, and the two are naturally integrated in artistic expression, forming a unique visual beauty with oriental characteristics.
[0003] In recent years, generative artificial intelligence (AIGC) technology has made significant breakthroughs in the field of artistic creation. In particular, diffusion models represented by Stable Diffusion and large-scale generative models such as MidJourney and DALL-E have demonstrated powerful image generation capabilities through learning from massive multimodal data. However, the current diffusion models are mainly trained on general images, lacking specialized data for Chinese calligraphy and traditional Chinese painting. In addition, they focus too much on global features during the generation process, often ignoring the unique structural details of Chinese characters, resulting in distortion of glyphs or structural damage. In addition, when migrating the style of Chinese painting to Chinese character images, it is difficult to accurately grasp the geometric form and stroke characteristics of Chinese characters, and the generated results often remain at the surface level of color imitation, failing to achieve the essential fusion of calligraphy and painting art forms. Summary of the invention
[0004] The present invention provides a method and device for fusing traditional Chinese painting and calligraphy based on an enhanced generation network, which can ensure the accuracy of the fusion of traditional Chinese painting and calligraphy.
[0005] To achieve the above purpose, the present invention provides a method for integrating traditional Chinese calligraphy and painting based on an enhanced generative network, comprising:
[0006] Obtain Chinese character image datasets and landscape painting image datasets;
[0007] The generator network is pre-trained using a Chinese character image dataset and a landscape painting image dataset to obtain a pre-trained generator, wherein the generator network includes an encoder, a feature transformation network, a residual connection module and a decoder; the output end of the encoder is connected to the input end of the feature transformation network and the input end of the residual connection module respectively; the residual connection module performs residual processing on the output features of the encoder and the output features of the feature transformation network, and inputs the residual processing results into the decoder;
[0008] The pre-trained generator is trained for style transfer to transfer the structural features of the Chinese character image to the target landscape painting image features and generate a fused image.
[0009] Optionally, the acquiring of a Chinese character image dataset and a landscape painting image dataset includes:
[0010] Acquire Chinese character images and landscape painting images;
[0011] Define the pixel threshold in the Chinese character image, and perform binarization processing on the Chinese character image according to the pixel threshold;
[0012] Expand the size of the binarized Chinese character image to obtain a Chinese character image dataset;
[0013] Identify the border in the landscape painting image, and crop the landscape painting image according to the border to obtain a cropped landscape painting image;
[0014] The landscape painting cropped images are divided into squares, the resolution is improved, and the saturation is adjusted to obtain a landscape painting image dataset.
[0015] Optionally, the step of dividing the cropped landscape painting image into squares, improving the resolution, and adjusting the saturation to obtain the landscape painting image dataset includes:
[0016] Identify the image area of the cropped image of the landscape painting, and divide the image area into squares according to a preset size to obtain a landscape painting image;
[0017] The landscape painting image is enhanced in resolution by using super-resolution technology to obtain a clear landscape painting image;
[0018] The clear landscape painting images are linearly enhanced in saturation in color space to obtain high-contrast landscape painting images, and the high-contrast landscape painting images are summarized to obtain a landscape painting image dataset.
[0019] Optionally, the feature transformation network comprises a plurality of cascaded ConvNeXt blocks.
[0020] Optionally, the residual connection module performs residual processing on the output features of the encoder and the output features of the feature transformation network according to the following formula: y=x+(γ·γ scale )·F(x).
[0021] Optionally, the residual connection strength of the residual connection module is controlled during style transfer training.
[0022] Optionally, the style transfer loss function is used to perform style transfer training on the pre-trained generator, and the style transfer loss function for:
[0023]
[0024] Among them, L gan To combat the loss, L cyc is the cycle consistency loss, L identity is the identity loss, L quality is the bidirectional quality constraint loss, where Among them, w(t) is the piecewise weight function, q is an adjustable parameter, t is the current number of training rounds, and T is the total number of training rounds. a is the Chinese character image, G A (a) is the preprocessed Chinese character image, b is the landscape painting image, G B (b) is a landscape painting image, and the MS-SSIM(·) function is the multi-scale structural similarity index.
[0025] In order to solve the above problems, the present invention also provides a Chinese traditional calligraphy and painting fusion device based on an enhanced generation network, the device comprising:
[0026] A data acquisition module, used to acquire a Chinese character image dataset and a landscape painting image dataset;
[0027] A model training module is used to pre-train the generator network using a Chinese character image dataset and a landscape painting image dataset to obtain a pre-trained generator, wherein the generator network includes an encoder, a feature transformation network, a residual connection module and a decoder; the output end of the encoder is respectively connected to the input end of the feature transformation network and the input end of the residual connection module; the residual connection module performs residual processing on the output features of the encoder and the output features of the feature transformation network, and inputs the residual processing results into the decoder;
[0028] The image generation module is used to perform style transfer training on the pre-trained generator to transfer the structural features of the Chinese character image to the target landscape painting image features and generate a fused image.
[0029] In order to solve the above problem, the present invention further provides an electronic device, the electronic device comprising:
[0030] at least one processor; and,
[0031] a memory communicatively connected to the at least one processor; wherein,
[0032] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned traditional Chinese calligraphy and painting fusion method based on enhanced generation network.
[0033] In order to solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is executed by a processor in an electronic device to implement the above-mentioned Chinese traditional calligraphy and painting fusion method based on an enhanced generation network.
[0034] The present invention uses a Chinese character image dataset and a landscape painting image dataset to pre-train a generator network, so that the generator network can accurately obtain the structural detail features of the Chinese character image and maintain the original shape and structure of the Chinese characters in the Chinese character image. In addition, the pre-trained generator is trained in style transfer to transfer the structural features of the Chinese character image to the target landscape painting image features, which can achieve accurate fusion of the features in the Chinese character image and the features of the landscape painting image, and realize the generation of a fusion image of calligraphy and painting. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 A schematic diagram of a flow chart of a method for integrating traditional Chinese calligraphy and painting based on an enhanced generative network provided by an embodiment of the present invention;
[0036] Figure 2 A structural flow chart of an example of a method for integrating traditional Chinese calligraphy and painting based on an enhanced generative network provided by an embodiment of the present invention;
[0037] Figure 3 A schematic diagram of a generator network structure of a method for integrating traditional Chinese calligraphy and painting based on an enhanced generative network provided by an embodiment of the present invention;
[0038] Figure 4 A schematic diagram of a generator network structure of a method for integrating traditional Chinese calligraphy and painting based on an enhanced generative network provided by an embodiment of the present invention;
[0039] Figure 5 A functional module diagram of a Chinese traditional calligraphy and painting fusion device based on an enhanced generation network provided by an embodiment of the present invention;
[0040] Figure 6 A schematic diagram of the structure of an electronic device for implementing the traditional Chinese painting and calligraphy fusion method based on an enhanced generative network provided in one embodiment of the present invention.
[0041] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0042] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0043] The embodiment of the present application provides a method for integrating traditional Chinese calligraphy and painting based on an enhanced generation network. The execution subject of the method for integrating traditional Chinese calligraphy and painting based on an enhanced generation network includes but is not limited to at least one of the electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiment of the present application. In other words, the method for integrating traditional Chinese calligraphy and painting based on an enhanced generation network can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.
[0044] Reference Figure 1 FIG. 1 is a flow chart of a method for integrating traditional Chinese painting and calligraphy based on an enhanced generative network according to an embodiment of the present invention. In this embodiment, the method for integrating traditional Chinese painting and calligraphy based on an enhanced generative network includes:
[0045] S1. Obtain a Chinese character image dataset and a landscape painting image dataset.
[0046] In the embodiment of the present invention, the Chinese character image refers to a Chinese character presented in a visual form, which generally includes the shape, structure and writing style of the Chinese character.
[0047] In the embodiment of the present invention, the landscape painting image refers to a painting or calligraphy work with landscape natural scenery as the theme.
[0048] As an embodiment of the present invention, the step of acquiring a Chinese character image dataset and a landscape painting image dataset includes:
[0049] Acquire Chinese character images and landscape painting images;
[0050] Define the pixel threshold in the Chinese character image, and perform binarization processing on the Chinese character image according to the pixel threshold;
[0051] Expand the size of the binarized Chinese character image to obtain a Chinese character image dataset;
[0052] Identify the border in the landscape painting image, and crop the landscape painting image according to the border to obtain a cropped landscape painting image;
[0053] The landscape painting cropped images are divided into squares, the resolution is improved, and the saturation is adjusted to obtain a landscape painting image dataset.
[0054] Furthermore, the step of dividing the cropped landscape painting image into squares, improving the resolution, and adjusting the saturation to obtain the landscape painting image dataset includes:
[0055] Identify the image area of the cropped image of the landscape painting, and divide the image area into squares according to a preset size to obtain a landscape painting image;
[0056] The landscape painting image is enhanced in resolution by using super-resolution technology to obtain a clear landscape painting image;
[0057] The clear landscape painting images are linearly enhanced in saturation in color space to obtain high-contrast landscape painting images, and the high-contrast landscape painting images are summarized to obtain a landscape painting image dataset.
[0058] In the embodiment of the present invention, the image region of the cropped image of the landscape painting may be identified by identifying the structured region in the image through Hough transform.
[0059] In the embodiment of the present invention, the color space refers to the HSV color space.
[0060] The embodiment of the present invention pre-processes Chinese character images and landscape painting images to ensure model training efficiency and accuracy, and can effectively improve image quality, reduce noise interference and improve recognition accuracy.
[0061] S2. Pre-train the generator network using a Chinese character image dataset and a landscape painting image dataset to obtain a pre-trained generator, wherein the generator network includes an encoder, a feature transformation network, a residual connection module and a decoder; the output end of the encoder is respectively connected to the input end of the feature transformation network and the input end of the residual connection module; the residual connection module performs residual processing on the output features of the encoder and the output features of the feature transformation network, and inputs the residual processing results into the decoder.
[0062] As an embodiment of the present invention, the feature transformation network includes multiple cascaded ConvNeXt blocks.
[0063] In an embodiment of the present invention, the generator network is replaced by a ConvNeXt block from a traditional ResNet block, while retaining the original downsampling and upsampling structure to maintain the overall topology of the network. There are significant differences between the ConvNeXt block and the ResNet block in design concepts and structural features. First, in terms of receptive field, the ConvNeXt block uses a large-size 7×7 convolution kernel to replace the traditional 3×3 convolution, which significantly expands the receptive field range. This design enables the model to capture a wider range of spatial dependencies in a single convolution operation, effectively enhancing the ability to model global style features.
[0064] In an embodiment of the present invention, it is further preferred that the ConvNeXt block also includes normalization processing using LayerNorm (LN); and gated gradient smoothing processing of the input using the GELU activation function. In terms of feature normalization strategy, the ConvNeXt block introduces LayerNorm (LN) to replace BatchNorm (BN). Compared with the characteristic of BN being more sensitive to batch size, LN provides more stable training dynamics and more robust feature expression by independently normalizing the features of each sample. This improvement is particularly important in style transfer tasks because it can better maintain the style consistency between samples. In addition, in terms of activation function selection, ConvNeXt uses the GELU activation function to replace the traditional ReLU activation function. The GELU activation function gates the input by introducing the idea of probability theory, while maintaining nonlinear characteristics, it provides a smoother gradient flow, which helps the model capture more delicate style feature changes. It is worth emphasizing that our solution innovatively maintains the overall architecture of the CycleGAN generator and only accurately replaces the feature extraction module. The traditional generator network is a data sample generation network composed of downsampling and upsampling structures and traditional ResNet blocks.
[0065] In the embodiment of the present invention, the GELU activation function is used to introduce nonlinearity, control signal transmission, reduce the risk of gradient vanishing by smoothing negative areas, and dynamically adjust the activation level. However, the traditional ReLU activation function has a zero gradient in negative areas, resulting in asymmetric output that affects model performance.
[0066] In an embodiment of the present invention, the generator network is pre-trained using a Chinese character image dataset and a landscape painting image dataset, and a ConvNeXt block can be used to replace a ResNet block in a traditional generator network to obtain a generator network; the Chinese character image dataset and the landscape painting image dataset are used as training set data to complete inter-domain conversion in the generator network, and a discriminator is used to determine whether the input data is a real dataset.
[0067] Further preferably, the residual connection module performs residual processing on the output features of the encoder and the output features of the feature transformation network according to the following formula: y=x+(γ·γ scale )·F(x).
[0068] The generator network in the embodiment of the present invention further includes improving the processing process of the residual connection module, and the specific improvement is as follows: using the residual connection function y=x+(γ·γ scale )·F(x), where F(x) is the feature transformation function, γ is the basic residual connection strength parameter, and γ scaleis a dynamic scaling factor, where the γ parameter is optimized and updated by gradient descent, as shown in the following formula: Among them, γ scale is the current value of γ, η is the learning rate, is the partial derivative of the loss function L with respect to γ. When the parameter γ is set to a non-trainable state (i.e. requires_grad = False), the parameter does not participate in the gradient calculation of back propagation, so the optimizer does not update the value of γ, thereby keeping its initial setting unchanged. In addition, γ scale As an additional tuning parameter, it allows to manually adjust the degree of style transfer during inference.
[0069] S3. Perform style transfer training on the pre-trained generator to transfer the structural features of the Chinese character image to the target landscape painting image features, and generate a fused image.
[0070] In the embodiment of the present invention, residual connection refers to a technology used to alleviate the problem of gradient disappearance or explosion in a deep learning model.
[0071] The style transfer training performed in the embodiment of the present invention is to use a group of Chinese character images and designated landscape painting images as style transfer training data for training.
[0072] As an embodiment of the present invention, the residual connection strength of the residual connection module is controlled in the style transfer training.
[0073] Furthermore, the style transfer loss function is used to perform style transfer training on the pre-trained generator, and the style transfer loss function for:
[0074]
[0075] Among them, L gan To combat the loss, L cyc is the cycle consistency loss, L identity is the identity loss, L quality is the bidirectional quality constraint loss, where Among them, w(t) is the piecewise weight function, q is an adjustable parameter, t is the current number of training rounds, and T is the total number of training rounds. a is the Chinese character image, G A (a) is the preprocessed Chinese character image, b is the landscape painting image, G B (b) is a landscape painting image, and the MS-SSIM(·) function is the multi-scale structural similarity index.
[0076] The embodiment of the present invention transfers the structural features of the Chinese character image to the target landscape painting image features, and generates a calligraphy and painting fusion image with the Chinese character image features.
[0077] The present invention uses a Chinese character image dataset and a landscape painting image dataset to pre-train a generator network, so that the generator network can accurately obtain the structural detail features of the Chinese character image and maintain the original shape and structure of the Chinese characters in the Chinese character image. In addition, the pre-trained generator is trained in style transfer to transfer the structural features of the Chinese character image to the target landscape painting image features, which can achieve accurate fusion of the features in the Chinese character image and the features of the landscape painting image, and realize the generation of a fusion image of calligraphy and painting.
[0078] Reference Figure 2 As shown, it is a structural flow chart of an example of a method for integrating traditional Chinese calligraphy and painting based on an enhanced generative network provided by an embodiment of the present invention.
[0079] Reference Figure 3 As shown, it is a schematic diagram of the generator network structure of the Chinese traditional calligraphy and painting fusion method based on the enhanced generation network provided by an embodiment of the present invention.
[0080] Reference Figure 4 As shown, it is a schematic diagram of the generator network structure of the Chinese traditional calligraphy and painting fusion method based on the enhanced generation network provided by an embodiment of the present invention.
[0081] like Figure 5 As shown, it is a functional module diagram of a Chinese traditional calligraphy and painting fusion device based on an enhanced generation network provided by an embodiment of the present invention.
[0082] The Chinese traditional painting and calligraphy fusion device 100 based on the enhanced generation network of the present invention can be installed in an electronic device. According to the functions to be implemented, the Chinese traditional painting and calligraphy fusion device 100 based on the enhanced generation network can include a low data processing module 101, a model training module 102 and an image generation module 103.
[0083] The module described in the present invention may also be referred to as a unit, which refers to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and is stored in a memory of the electronic device.
[0084] In this embodiment, the functions of each module / unit are as follows:
[0085] The data acquisition module 101 is used to acquire a Chinese character image dataset and a landscape painting image dataset.
[0086] In the embodiment of the present invention, the Chinese character image refers to a Chinese character presented in a visual form, which generally includes the shape, structure and writing style of the Chinese character.
[0087] In the embodiment of the present invention, the landscape painting image refers to a painting or calligraphy work with landscape natural scenery as the theme.
[0088] As an embodiment of the present invention, the step of acquiring a Chinese character image dataset and a landscape painting image dataset includes:
[0089] Acquire Chinese character images and landscape painting images;
[0090] Define the pixel threshold in the Chinese character image, and perform binarization processing on the Chinese character image according to the pixel threshold;
[0091] Expand the size of the binarized Chinese character image to obtain a Chinese character image dataset;
[0092] Identify the border in the landscape painting image, and crop the landscape painting image according to the border to obtain a cropped landscape painting image;
[0093] The landscape painting cropped images are divided into squares, the resolution is improved, and the saturation is adjusted to obtain a landscape painting image dataset.
[0094] Furthermore, the step of dividing the cropped landscape painting image into squares, improving the resolution, and adjusting the saturation to obtain the landscape painting image dataset includes:
[0095] Identify the image area of the cropped image of the landscape painting, and divide the image area into squares according to a preset size to obtain a landscape painting image;
[0096] The landscape painting image is enhanced in resolution by using super-resolution technology to obtain a clear landscape painting image;
[0097] The clear landscape painting images are linearly enhanced in saturation in color space to obtain high-contrast landscape painting images, and the high-contrast landscape painting images are summarized to obtain a landscape painting image dataset.
[0098] In the embodiment of the present invention, the image region of the cropped image of the landscape painting may be identified by identifying the structured region in the image through Hough transform.
[0099] In the embodiment of the present invention, the color space refers to the HSV color space.
[0100] The embodiment of the present invention pre-processes Chinese character images and landscape painting images to ensure model training efficiency and accuracy, and can effectively improve image quality, reduce noise interference and improve recognition accuracy.
[0101] The model training module 102 is used to pre-train the generator network using a Chinese character image data set and a landscape painting image data set to obtain a pre-trained generator, wherein the generator network includes an encoder, a feature transformation network, a residual connection module and a decoder; the output end of the encoder is respectively connected to the input end of the feature transformation network and the input end of the residual connection module; the residual connection module performs residual processing on the output features of the encoder and the output features of the feature transformation network, and inputs the residual processing results into the decoder.
[0102] In an embodiment of the present invention, the generator network is replaced by a ConvNeXt block from a traditional ResNet block, while retaining the original downsampling and upsampling structure to maintain the overall topology of the network. There are significant differences between the ConvNeXt block and the ResNet block in design concepts and structural features. First, in terms of receptive field, the ConvNeXt block uses a large-size 7×7 convolution kernel to replace the traditional 3×3 convolution, which significantly expands the receptive field range. This design enables the model to capture a wider range of spatial dependencies in a single convolution operation, effectively enhancing the ability to model global style features.
[0103] In an embodiment of the present invention, it is further preferred that the ConvNeXt block also includes normalization processing using LayerNorm (LN); and gated gradient smoothing processing is performed on the input using the GELU activation function. In terms of feature normalization strategy, the ConvNeXt block introduces LayerNorm (LN) to replace BatchNorm (BN). Compared with the characteristic of BN being more sensitive to batch size, LN provides more stable training dynamics and more robust feature expression by independently normalizing the features of each sample. This improvement is particularly important in style transfer tasks because it can better maintain the style consistency between samples. In addition, in terms of activation function selection, ConvNeXt uses the GELU activation function to replace the traditional ReLU. GELU gates the input by introducing the idea of probability theory, while maintaining nonlinear characteristics, it provides a smoother gradient flow, which helps the model capture more delicate style feature changes. It is worth emphasizing that our solution innovatively maintains the overall architecture of the CycleGAN generator and only accurately replaces the feature extraction module. The traditional generator network is a data sample generation network composed of downsampling and upsampling structures and traditional ResNet blocks.
[0104] As an embodiment of the present invention, the feature transformation network includes multiple cascaded ConvNeXt blocks.
[0105] Furthermore, the residual connection module performs residual processing on the output features of the encoder and the output features of the feature transformation network according to the following formula: y=x+(γ·γscale )·F(x).
[0106] In the embodiment of the present invention, the GELU activation function is used to introduce nonlinearity, control signal transmission, reduce the risk of gradient vanishing by smoothing negative areas, and dynamically adjust the activation level. However, the traditional ReLU activation function has a zero gradient in negative areas, resulting in asymmetric output that affects model performance.
[0107] In an embodiment of the present invention, the generator network is pre-trained using a Chinese character image dataset and a landscape painting image dataset, and a ConvNeXt block can be used to replace a ResNet block in a traditional generator network to obtain a generator network; the Chinese character image dataset and the landscape painting image dataset are used as training set data to complete inter-domain conversion in the generator network, and a discriminator is used to determine whether the input data is a real dataset.
[0108] In the embodiment of the present invention, the generator network further includes improving the residual connection in the ConvNeXt block, and the specific improvement is as follows: using the residual connection function y=x+(γ·γ scale )·F(x), where F(x) is the feature transformation function, γ is the basic residual connection strength parameter, and γ scale is a dynamic scaling factor, where the γ parameter is optimized and updated by gradient descent, as shown in the following formula: Among them, γ scale is the current value of γ, η is the learning rate, is the partial derivative of the loss function L with respect to γ. When the parameter γ is set to a non-trainable state (i.e. requires_grad = False), the parameter does not participate in the gradient calculation of back propagation, so the optimizer does not update the value of γ, thereby keeping its initial setting unchanged. In addition, γ scale As an additional tuning parameter, it allows to manually adjust the degree of style transfer during inference.
[0109] The image generation module 103 is used to perform style transfer training on the pre-trained generator to transfer the structural features of the Chinese character image to the target landscape painting image features and generate a fused image.
[0110] In the embodiment of the present invention, residual connection refers to a technology used to alleviate the problem of gradient disappearance or explosion in a deep learning model.
[0111] The style transfer training performed in the embodiment of the present invention is to use a group of Chinese character images and designated landscape painting images as style transfer training data for training.
[0112] As an embodiment of the present invention, the residual connection strength of the residual connection module is controlled in the style transfer training.
[0113] Furthermore, the style transfer loss function is used to perform style transfer training on the pre-trained generator, and the style transfer loss function for:
[0114]
[0115] Among them, L gan To combat the loss, L cyc is the cycle consistency loss, L identity is the identity loss, L quality is the bidirectional quality constraint loss, where Among them, w(t) is the piecewise weight function, q is an adjustable parameter, t is the current number of training rounds, and T is the total number of training rounds. a is the Chinese character image, G A (a) is the preprocessed Chinese character image, b is the landscape painting image, G B (b) is a landscape painting image, and the MS-SSIM(·) function is the multi-scale structural similarity index.
[0116] The embodiment of the present invention transfers the structural features of the Chinese character image to the target landscape painting image features, and generates a calligraphy and painting fusion image with the Chinese character image features.
[0117] like Figure 6 , which is a schematic diagram of the structure of an electronic device for implementing a method for integrating traditional Chinese calligraphy and painting based on an enhanced generation network provided by an embodiment of the present invention.
[0118] The electronic device may include a processor 10, a memory 11, a communication bus 12 and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a program for a method for integrating traditional Chinese calligraphy and painting based on an enhanced generative network.
[0119] In some embodiments, the processor 10 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor 10 is the control core (ControlUnit) of the electronic device, and uses various interfaces and lines to connect the various components of the entire electronic device, and executes or executes the program or module stored in the memory 11 (for example, executing a Chinese traditional calligraphy and painting fusion method program based on an enhanced generation network, etc.), and calls the data stored in the memory 11 to execute various functions of the electronic device and process data.
[0120] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (for example: SD or DX memory, etc.), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of an electronic device, such as a mobile hard disk of the electronic device. In other embodiments, the memory 11 can also be an external storage device of an electronic device, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Further, the memory 11 can also include both an internal storage unit of the electronic device and an external storage device. The memory 11 can not only be used to store application software and various types of data installed in the electronic device, such as a code of a Chinese traditional calligraphy and painting fusion method program based on an enhanced generation network, but can also be used to temporarily store data that has been output or is to be output.
[0121] The communication bus 12 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize connection and communication between the memory 11 and at least one processor 10, etc.
[0122] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device and other electronic devices. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)), and optionally, the user interface may also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode, organic light-emitting diode) touch device, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device and to display a visual user interface.
[0123] Figure 6 Only an electronic device with components is shown, and those skilled in the art will understand that Figure 6 The structure shown does not constitute a limitation on the electronic device, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0124] For example, although not shown, the electronic device may also include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.
[0125] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0126] The program of a method for integrating traditional Chinese calligraphy and painting based on an enhanced generative network stored in the memory 11 of the electronic device is a combination of multiple instructions. When running in the processor 10, the following can be achieved:
[0127] Obtain Chinese character image datasets and landscape painting image datasets;
[0128] The generator network is pre-trained using a Chinese character image dataset and a landscape painting image dataset to obtain a pre-trained generator, wherein the generator network includes an encoder, a feature transformation network, a residual connection module and a decoder; the output end of the encoder is connected to the input end of the feature transformation network and the input end of the residual connection module respectively; the residual connection module performs residual processing on the output features of the encoder and the output features of the feature transformation network, and inputs the residual processing results into the decoder;
[0129] The pre-trained generator is trained for style transfer to transfer the structural features of the Chinese character image to the target landscape painting image features and generate a fused image.
[0130] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to the description of the relevant steps in the corresponding embodiment of the accompanying drawings, which will not be repeated here.
[0131] Furthermore, if the module / unit integrated in the electronic device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).
[0132] The present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, the computer program can implement:
[0133] Obtain Chinese character image datasets and landscape painting image datasets;
[0134] The generator network is pre-trained using a Chinese character image dataset and a landscape painting image dataset to obtain a pre-trained generator, wherein the generator network includes an encoder, a feature transformation network, a residual connection module and a decoder; the output end of the encoder is connected to the input end of the feature transformation network and the input end of the residual connection module respectively; the residual connection module performs residual processing on the output features of the encoder and the output features of the feature transformation network, and inputs the residual processing results into the decoder;
[0135] The pre-trained generator is trained for style transfer to transfer the structural features of the Chinese character image to the target landscape painting image features and generate a fused image.
[0136] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0137] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0138] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0139] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0140] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.
[0141] The blockchain referred to in this invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.
[0142] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0143] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the system claim can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for integrating traditional Chinese calligraphy and painting based on an enhanced generative network, characterized in that: The method comprises: Obtain Chinese character image datasets and landscape painting image datasets; The generator network is pre-trained using a Chinese character image dataset and a landscape painting image dataset to obtain a pre-trained generator, wherein the generator network includes an encoder, a feature transformation network, a residual connection module and a decoder; the output end of the encoder is connected to the input end of the feature transformation network and the input end of the residual connection module respectively; the residual connection module performs residual processing on the output features of the encoder and the output features of the feature transformation network, and inputs the residual processing results into the decoder; The pre-trained generator is trained for style transfer to transfer the structural features of the Chinese character image to the target landscape painting image features and generate a fused image.
2. The Chinese traditional calligraphy and painting fusion method based on enhanced generative network as claimed in claim 1 is characterized in that: The step of acquiring a Chinese character image dataset and a landscape painting image dataset comprises: Acquire Chinese character images and landscape painting images; Define the pixel threshold in the Chinese character image, and perform binarization processing on the Chinese character image according to the pixel threshold; Expand the size of the binarized Chinese character image to obtain a Chinese character image dataset; Identify the border in the landscape painting image, and crop the landscape painting image according to the border to obtain a cropped landscape painting image; The landscape painting cropped images are divided into squares, the resolution is improved, and the saturation is adjusted to obtain a landscape painting image dataset.
3. The Chinese traditional calligraphy and painting fusion method based on enhanced generative network as claimed in claim 2 is characterized in that: The step of dividing the cropped landscape painting image into squares, improving the resolution, and adjusting the saturation to obtain a landscape painting image dataset includes: Identify the image area of the cropped image of the landscape painting, and divide the image area into squares according to a preset size to obtain a landscape painting image; The landscape painting image is enhanced in resolution by using super-resolution technology to obtain a clear landscape painting image; The clear landscape painting images are linearly enhanced in saturation in color space to obtain high-contrast landscape painting images, and the high-contrast landscape painting images are summarized to obtain a landscape painting image dataset.
4. The Chinese traditional calligraphy and painting fusion method based on enhanced generative network as claimed in claim 1 is characterized in that: The feature transformation network includes multiple cascaded ConvNeXt blocks.
5. The Chinese traditional calligraphy and painting fusion method based on enhanced generative network as claimed in claim 1, characterized in that: The residual connection module performs residual processing on the output features of the encoder and the output features of the feature transformation network according to the following formula: y = x + (γ·γ scale )·F(x).
6. The Chinese traditional calligraphy and painting fusion method based on enhanced generative network as claimed in claim 1, characterized in that: Control the residual connection strength of the residual connection module during style transfer training.
7. The Chinese traditional calligraphy and painting fusion method based on enhanced generative network as claimed in claim 6 is characterized in that: The style transfer loss function is used to perform style transfer training on the pre-trained generator, and the style transfer loss function for: Among them, L gan To combat the loss, L cyc is the cycle consistency loss, L identity is the identity loss, L quality is the bidirectional quality constraint loss, where Among them, w(t) is the piecewise weight function, q is an adjustable parameter, t is the current number of training rounds, and T is the total number of training rounds. a is the Chinese character image, G A (a) is the preprocessed Chinese character image, b is the landscape painting image, G B (b) is a landscape painting image, and the MS-SSIM(·) function is the multi-scale structural similarity index.
8. A Chinese traditional calligraphy and painting fusion device based on enhanced generative network, characterized in that: The device is used to implement the Chinese traditional calligraphy and painting fusion method based on enhanced generative network as described in any one of claims 1 to 7, and the device includes: A data acquisition module, used to acquire a Chinese character image dataset and a landscape painting image dataset; A model training module is used to pre-train the generator network using a Chinese character image dataset and a landscape painting image dataset to obtain a pre-trained generator, wherein the generator network includes an encoder, a feature transformation network, a residual connection module and a decoder; the output end of the encoder is respectively connected to the input end of the feature transformation network and the input end of the residual connection module; the residual connection module performs residual processing on the output features of the encoder and the output features of the feature transformation network, and inputs the residual processing results into the decoder; The image generation module is used to perform style transfer training on the pre-trained generator to transfer the structural features of the Chinese character image to the target landscape painting image features and generate a fused image.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the Chinese traditional calligraphy and painting fusion method based on the enhanced generation network as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for integrating traditional Chinese calligraphy and painting based on an enhanced generative network as described in any one of claims 1 to 7 is implemented.