A Calligraphy and Painting Character Recognition Method Based on the Fusion of Spatial and Frequency Domains
Through the deep learning method of air-frequency domain fusion, the VanillaNet-6 network is used to extract the air-space texture and frequency-domain shape characteristics of calligraphy and painting text, which solves the problem of relying on experience and damage to cultural relics in the existing technology, and achieves high-precision calligraphy and painting text recognition.
Patent Information
- Application Number
- CN202311196354.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-09-18
AI Technical Summary
The existing technology relies on professional experience in the identification of calligraphy and painting works, lacks popularity and may damage cultural relics. Computer vision methods are not accurate in identifying calligraphy and painting, and it is difficult to effectively identify the unique characteristics of calligraphy and painting.
The air-frequency domain fusion method based on deep learning is adopted to extract the air-space texture features and frequency-domain shape features of calligraphy and painting texts through VanillaNet-6 network, and the recognition accuracy is improved through integrated learning and fusion.
It improves the recognition accuracy of calligraphy, painting and writing works, reduces damage to cultural relics, reduces dependence on professional experience, and enhances the popularity and stability of the identification system.
Smart Images

Figure CN117315278B_ABST
Abstract
Description
Technical Field:
[0001] The present invention relates to an identification algorithm for calligraphy and painting text images. Prior Art:
[0002] The identification of calligraphy and painting works is recognized as one of the most complex disciplines in cultural relics identification. It requires not only professional skills, academic cultivation, but also perception and wisdom, and more importantly, a large amount of relevant experience. The main basis for the identification of calligraphy and painting works is the works themselves, namely the era style of calligraphy and painting texts, the personal style of calligraphers and painters, and the texture of the works, etc. Traditional calligraphy and painting works identification faces many problems. It mainly relies on the professional technical cultivation of the appraisers themselves and the accumulation of long-term relevant experience. However, each appraisal expert will be restricted by their own knowledge field and vision. This identification method is mainly for some famous works. The appraiser may make inferences about the works according to the style of the same author and in combination with historical periods and material usage, or compare the works to be detected with the works that have been determined to be genuine, and finally draw a conclusion. The disadvantage of this method is that it relies too much on experience and has high requirements for the quality of the appraisers, that is, it requires the appraisers to have seen a large number of works and have physical objects for comparison in order to draw a credible conclusion.
[0003] For traditional identification methods such as using physical and chemical methods and various high-frequency spectral methods for auxiliary identification, but such methods will bring many problems. First of all, physical and chemical methods will obviously cause inevitable damage to precious cultural relics. Even the method of high-frequency spectral irradiation may cause problems such as the discoloration of the pigments of calligraphy and painting works. Secondly, the supporting equipment for this method is relatively professional and expensive, lacking popularity. Like the traditional identification method, the channels for identification are very narrow. Computer vision technology is a simulation of biological vision using a computer and related imaging equipment, imitating the process of humans obtaining, processing, analyzing, and understanding images. By imitating the identification method of experts, computer vision technology can be used to identify calligraphy and painting works. Due to the large variety and quantity of modern and contemporary calligraphy and painting works, and the sophisticated and diverse forgery methods, these factors have added many obstacles and difficulties to the research of calligraphy and painting identification technology based on computer vision. Therefore, it has gradually become more and more important to use computer vision technology for non-contact identification of calligraphy and painting works. How to use the existing vision technology to construct an identification system for calligraphy and painting works is an urgent problem to be solved at present.
[0004] A discrimination method based on computer vision takes digital works as the research object and uses technologies such as image processing, computer vision, and image recognition to provide people with qualitative or quantitative indicators to determine the authenticity of the works. Literature 1 "Hong Y, Kim J. Art painting identification using convolutional neural network[J]. International Journal of Applied Engineering Research, 2017, 12(4): 532-539." proposed a method of applying CNN to art painting identification; Literature 2 "Zhai S, Liao L, Lin Y, et al. Inscription detection and style identification in Chinese painting[C] / / 2020 Chinese Automation Congress (CAC). IEEE, 2020: 7434-7438." improved by extracting inscriptions in paintings through a deep neural network and combining painting features, providing technical support for identification; Literature 3 "Tang X, Zhang P, Du J, et al. Calligraphy and painting identification method based on hyperspectral imaging and convolution neural network[J]. Spectroscopy Letters, 2021, 54(9): 645-664." adopted the idea of combining hyperspectral imaging and Atlas intelligent learning and proposed a method for extracting atlas features for calligraphy and painting recognition; Literature 4 "Buchana P, Cazan I, Diaz-Granados M, et al. Simultaneous forgery identification and localization in paintings using advanced correlation filters[C] / / 2016 IEEE International Conference on Image Processing (ICIP). IEEE, 2016: 146-150."By training an optimal trade-off synthetic discriminant function (OTSDF) filter on each part of the coarsely segmented image of the original painting, distinguishing between low-quality digital representations of the painting and its forgeries, and highlighting the locations where differences occur and where the replicas are particularly faithful to the original; Document 5 "Wang Z, Lu D, Zhang D, et al. Fake modern Chinese painting identification based on spectral–spatial feature fusion on hyperspectral image [J]. Multidimensional Systems and Signal Processing, 2016, 27: 1031-1044." proposed a feature fusion method based on hyperspectral images for identification of those made using modern high-resolution scanning and printing techniques. These methods have improved the analysis performance of calligraphy and painting works, but they focus on image analysis methods and lack research in the identification and precise analysis of calligraphy and painting works. They only identify the overall content of calligraphy and painting digital images through hyperspectral images and traditional features, and the accuracy rate is not high."
[0005] Object of the invention:
[0006] This project completes the task of identifying calligraphy and painting works through computer vision technology, and proposes a visual recognition algorithm based on time-frequency domain fusion to accurately identify the authors of calligraphy and painting works. In the spatial domain, the feature extraction network often obtains the texture features of the image. However, the identification of Chinese calligraphy and painting works is not only the texture features, but more importantly, the shape features composed of their unique radicals, strokes, and structures. Therefore, the present invention extracts the features of calligraphy and painting characters in the frequency domain to obtain the overall shape features of calligraphy and painting characters. Finally, the texture features in the spatial domain are fused with the shape features in the frequency domain to improve the recognition accuracy of the algorithm for Chinese calligraphy and painting works." Content of the invention:
[0007] The present invention is based on a deep learning model in the field of computer vision, extracts deeper features of Chinese calligraphy and painting character images in the spatial and frequency domains, and the designed network structure is as shown in the appendix Figure 1As shown in the figure. The feature extraction network in the spatial domain uses a minimalist network VanillaNet-6 with only 6 layers as the Backbone. In the network structure, residual blocks, channel attention, self-attention mechanisms, and multi-branch structures are not used, but a very simple straight-through structure containing only convolutional calculations is adopted. In the frequency domain, through the Discrete Cosine Transform (DCT), after concatenating the dimensions of the transformed features with the original feature dimensions, and then through Ensemble Learning, the features extracted in the spatial domain and the frequency domain are fused to improve the stability of the overall model.
[0008] Step 1: Construct the spatial domain network
[0009] 1) Merge convolutional layers
[0010] During network training, two convolutional layers with activation functions are used to increase the nonlinear representation ability of the model, and at the end of training, they are merged into a simple convolutional layer to reduce the inference time. Among them, the activation function can be expressed as:
[0011] A′(x)=(1 - λ)A(x)+λx (1)
[0012] Among them, A(x) represents the activation function of the convolutional layer. Using the hyperparameter λ = epoch / Epochs (epoch represents the current training round, and Epochs represents the total number of rounds), at the beginning of training, A′(x) = A(x), and at the end of training, A′(x) = x, so as to merge it into a convolutional layer with an identity transformation to improve the inference speed.
[0013] 2) Parallel stacking of activation functions
[0014] Adopting parallel stacking of activation functions can provide more nonlinear representations and feature combination capabilities to enhance the performance and adaptability of the neural network. This method can increase the flexibility of the network and is suitable for some tasks that require more nonlinear operations. The parallel stacking of activation functions can be expressed as:
[0015]
[0016] Among them, a i and b i are the scale and bias of each activation function respectively, and n represents the number of activation functions used for stacking.
[0017] To further enrich the approximation ability of the model, global information is learned by changing the input of neighboring values. Given the input where H, W, and C represent the height, width, and number of channels of x respectively, so the activation function can be expressed as:
[0018]
[0019] Among them, h ∈ {1, 2,..., H}, w ∈ {1, 2,..., W}, c ∈ {1, 2,..., C}, and the activation function A s (x h,w,c ) obtains the activation function of neighboring values with different n values, thereby obtaining global information.
[0020] Step 2: Construct a frequency-domain network
[0021] 1) Calculate the DCT coefficient matrix
[0022] Given an N×N image F(u, v), where u and v represent discrete frequency variables in two directions respectively, the DCT coefficient matrix G(u, v) can be expressed as:
[0023]
[0024] Among them, C(u) and C(v) are the scaling factors at positions u and v in the DCT coefficient matrix, taking values of 1 / √2 or 1 according to different positions.
[0025] 2) Extract frequency-domain features
[0026] At the input end of the network, for the given input Perform a two-dimensional discrete cosine transform on the two-dimensional matrices of its three RGB channels independently, convert the image from the spatial domain to the frequency domain, and splice the results of the DCT transform to obtain the corresponding input
[0027] For the obtained input x, after passing through the convolutional layer of the Backbone network, the corresponding frequency-domain feature f spatial , and x DCT After passing through the convolutional layer of the Backbone network with the number of convolutional kernels corrected, the obtained spatial-domain feature f freq And the frequency-domain feature f spatial After concatenating in dimensions, continue to perform feature extraction in the corrected Backbone network to obtain the final Out output, which can be specifically expressed by the following formula:
[0028]
[0029]
[0030]
[0031] Among them, Conv0 represents the first convolutional layer of the Backbone network, which is used to convert the input image from 3 channels to multiple channels in a downsampling manner to learn useful information; mConv represents the Backbone network with the number of convolutional kernels corrected, which is used to match the input after dimension concatenation; the spatial domain output and frequency domain output of the (n - 1)-th convolutional layer are respectively represented by and represent.
[0032] Step 3: Spatial-frequency domain ensemble learning
[0033] Ensemble learning is a machine learning method that combines the prediction results of multiple classifiers or models to obtain more accurate and stable prediction results. By combining the output Out of the spatial domain Backbone network and the output mOut of the corrected frequency domain Backbone network according to a certain ratio γ, the specific prediction result can be expressed by the following formula:
[0034] Prediction = Softmax(γ * Out + (1 - γ) * mOut) (8)
[0035] Beneficial effects:
[0036] The present invention uses the method of deep learning in the field of computer vision to extract the spatial domain and frequency domain features of calligraphy and painting characters, fuses the spatial domain features and frequency domain features, and further improves the generalization ability and stability of the model through the method of ensemble learning. Description of the drawings:
[0037] Figure 1 The network framework diagram designed by the present invention Specific implementation manners:
[0038] As Figure 1 shown, in the present invention, first, the calligraphy and painting character image is preprocessed, and its size is adjusted to a fixed size of 224×224; then, the deep learning model (VanillaNet-6 network) is trained using data, and feature extraction is performed using the deep network; finally, the trained deep learning model is used to make predictions on the test images. For each image, the model will output a vector representing the probabilities of the image belonging to different categories, and the category with the highest probability is selected as the recognition result of the work image, that is, the calligrapher or painter to which the calligraphy and painting characters belong. The specific process of the algorithm implementation is as follows:
[0039] Step 1: Input
[0040] The present invention collected 3,197 Chinese calligraphy and painting art works, and intercepted 68,979 text images therefrom. The size of each image was preprocessed and adjusted to 224×224 by the deep learning framework PyTorch 1.10.0. The calligraphy and painting text images of size 224×224×3 were used as the input of the spatial domain feature network. The text images after discrete cosine transform (DCT) were concatenated with the input of the spatial domain feature network in dimension through Concat, and the concatenated matrix of size 224×224×6 was used as the input of the frequency domain feature network.
[0041] Step 2: Construct the backbone network
[0042] For the optimization of the minimalist network with insufficient non-linear capabilities, during network training, two convolutional layers with activation functions were used to increase the non-linear performance of the model, and they were merged into a simple convolutional layer at the end of training to reduce the inference time. The activation function can be expressed as:
[0043] A′(x)=(1-λ)A(x)+λx (9)
[0044] Where A(x) represents the activation function of the convolutional layer, and the hyperparameter λ = epoch / Epochs, where Epochs represents the total number of 50 rounds. At the beginning of training, A′(x) = A(x), and at the end of training, A′(x) = x, so as to merge it into an identity transformation convolutional layer to improve the inference speed.
[0045] Using parallel stacked activation functions can provide more non-linear representations and feature combination capabilities to enhance the performance and adaptability of the neural network. This method can increase the flexibility of the network and is suitable for some tasks that require more non-linear operations. The parallel stacked activation function can be expressed as:
[0046]
[0047] Where a i and b i are the scale and bias of each activation function respectively, and n represents the number of activation functions used for stacking.
[0048] To further enrich the approximation ability, global information is learned by changing the input of neighboring values. Given an input Where H, W, and C represent the height, width, and number of channels of x respectively. Therefore, the activation function can be expressed as:
[0049]
[0050] where h ∈ {1, 2, ..., H}, w ∈ {1, 2, ..., W}, c ∈ {1, 2, ..., C}, and the activation function A s (x h,w,c ) obtains the activation function of neighboring values with different n values, thereby obtaining global information.
[0051] To fuse two convolutional layers, first convert each batch normalization layer and the convolutional layer in front of it into a single convolution. Denote it as the weight and bias matrices of a convolutional kernel with C in input channels, C out output channels, and kernel size k. The scale, displacement, mean, and variance of batch normalization are denoted as The combined weight and bias matrices are:
[0052]
[0053] where the subscript i ∈ {1, 2, ..., C out} represents the value of the i-th output channel.
[0054] After the input image matrix passes through the convolutional layer, the feature map (also called convolutional feature) of this layer is obtained. The size of the feature map depends on the size of the input image and the settings of the convolutional layer. Each pixel value on each feature map represents the result of the convolution operation between the convolutional kernel and the corresponding area in the input image. Through the convolution operation, local features in the image, such as edges and textures, can be extracted. These feature maps will be passed as input to the next layer of the network for subsequent processing and analysis.
[0055] Step 3: Extract spatial domain features
[0056] The spatial domain feature network consists of 5 merged convolutional layers. For a given input First, through a merged convolutional layer with a stride of 4 and a size of 4×4×3×512, the image with 3 channels is mapped to features with 512 channels. In the next 3 merged convolutional layers, a max pooling layer with a stride of 2 is used to reduce the size of the feature map and double the number of channels of the feature map, where the kernel size of each convolutional layer is 1×1. In the last merged convolutional layer, to improve the representational ability of the model, control the complexity of the model and the number of parameters, the number of channels is not continued to increase, and the feature matrix is scaled to a size of 1×1×4096 Out through a 7×7 average pooling layer. To use the minimum computational cost at each layer while maintaining the information of the feature map, an activation function is applied after each 1×1 convolutional layer. And for the convenience of the network training process, batch normalization is performed after each layer.
[0057] Step 4: Extract frequency domain features
[0058] The frequency-domain feature network has a similar structure to the spatial-domain feature network. Only the input dimension of each layer is doubled compared to the spatial-domain feature network. To keep the size of the final feature output mOut the same as the feature output Out of the spatial-domain feature network, the number of convolutional kernels in the last merging convolutional layer is kept as 4096.
[0059] Given an image matrix F(u, v) of 224×224, where u and v represent discrete frequency variables in two directions respectively, the DCT coefficient matrix G(u, v) can be expressed as:
[0060]
[0061] Among them, C(u) and C(v) are the scaling factors at positions u and v in the DCT coefficient matrix, and their values are or 1 according to different positions.
[0062] At the input end of the network, for the given input Perform a two-dimensional discrete cosine transform on the two-dimensional matrices of its three RGB channels independently, converting the image from the spatial domain to the frequency domain. After splicing the results of the DCT, obtain the corresponding input
[0063] For the obtained input x, after passing through the convolutional layers of the Backbone network, obtain the corresponding frequency-domain feature f spatial , and for x DCT After passing through the convolutional layers of the Backbone network with the number of convolutional kernels corrected, obtain the spatial-domain feature f freq And the frequency-domain feature f spatial After performing dimensional splicing, continue to perform feature extraction in the corrected Backbone network to obtain the final Out output, which can be specifically expressed by the following formula:
[0064]
[0065]
[0066]
[0067] Among them, Conv0 represents the first convolutional layer of the Backbone network, which is used to convert the input image from 3 channels to multiple channels in a downsampling manner to learn useful information; mConv represents the Backbone network with the number of convolutional kernels corrected, which is used to match the input after dimensional splicing; the spatial-domain output and frequency-domain output of the (n - 1)-th convolutional layer are represented by and respectively.
[0068] To match the spatio-temporal frequency features concatenated with the input dimension, the input dimension in the mConv convolutional kernel of the first 5 layers is twice that of the Conv convolutional kernel, and the spatial domain output features and the frequency domain output features are continuously dimensionally fused after each subsequent convolutional layer.
[0069] Step 5: Spatio-temporal frequency domain ensemble learning
[0070] Ensemble learning is a machine learning method that combines the prediction results of multiple classifiers or models to obtain more accurate and stable prediction results. By combining the output Out of the spatial domain Backbone network and the output mOut of the corrected frequency domain Backbone network in a certain proportion γ, the specific prediction result can be expressed by the following formula:
[0071] Prediction = Softmax(FC(γ * Out + (1 - γ) * mOut)) (17)
[0072] Among them, FC represents a fully connected layer with a kernel of 1×1×4096×N to output the classification result; N represents the number of categories; γ represents the hyperparameter for fusing the spatial domain and frequency domain feature outputs in proportion, and better model performance can be obtained by fine-tuning the value of γ. The value of γ in this model is 0.6.
[0073] The Softmax function performs exponential operations on each element in the input vector, and then normalizes the exponential results to obtain the probability of each category. The prediction result output is a vector composed of probabilities, where each element represents the prediction probability of the corresponding category.
Claims
1. A calligraphy and painting text work recognition method based on spatio-frequency domain fusion, including three parts: the design of the spatial domain feature extraction network, the design of the frequency domain feature extraction network, and spatio-frequency domain ensemble learning; (1) Design of the spatial domain feature extraction network During network training, two convolutional layers with activation functions are used to increase the nonlinear performance of the model, and they are merged into a simple convolutional layer at the end of training to reduce the inference time. The activation function can be expressed as: A′(x)=(1-λ)A(x)+λx (1) Among them, A(x) represents the activation function of the convolutional layer. Using the hyperparameter λ = epoch / Epochs, at the beginning of training, A′(x) = A(x), and at the end of training, A′(x) = x. In this way, it is merged into an identity transformation convolutional layer to improve the inference speed; Using parallel stacked activation functions can provide more nonlinear representations and feature combination capabilities to enhance the performance and adaptability of the neural network. This method can increase the flexibility of the network and is applicable to some tasks that require more nonlinear operations. The parallel stacked activation function can be expressed as: where a i and b i are the scale and bias of each activation function respectively, and n represents the number of activation functions used for stacking; To further enrich the approximation ability and learn global information by changing the input of neighboring values, given an input \(x\in\) where \(H\), \(W\), and \(C\) represent the height, width, and number of channels of \(x\), respectively. Therefore, the activation function can be expressed as: where h ∈ {1, 2,..., H}, w ∈ {1, 2,..., W}, c ∈ {1, 2,..., C}, and the activation function A s (x h,w,c ) obtains the activation function of neighboring values with different n values, thereby obtaining global information; (2) Design of the frequency domain feature extraction network The frequency domain feature network has a similar structure to the spatial domain feature network, except that the input dimension of each layer is doubled compared to the spatial domain feature network. To keep the size of the final feature output mOut the same as the feature output Out of the spatial domain feature network, the number of convolutional kernels in the last merged convolutional layer is kept as 4096; Given a 224×224 image matrix F(u,v), where u and v represent the discrete frequency variables in two directions respectively, the DCT coefficient matrix G(u,v) can be expressed as: Among them, C(u) and C(v) are the scaling factors at positions u and v in the DCT coefficient matrix, and take values of or 1; At the input end of the network, for a given input Perform a two-dimensional discrete cosine transform on the two-dimensional matrices of its three RGB channels independently, convert the image from the spatial domain to the frequency domain, and splice the results of the DCT to obtain the corresponding input The obtained input x is passed through the convolutional layer of the Backbone network to obtain the corresponding frequency-domain feature f spatial , and x DCT is passed through the convolutional layer of the Backbone network with the number of convolutional kernels corrected to obtain the spatial-domain feature f freq is concatenated with the frequency-domain feature f spatial in terms of dimensions, and then feature extraction is continued in the corrected Backbone network to obtain the final Out output, which can be specifically expressed by the following formula: Among them, Conv0 represents the first convolutional layer of the Backbone network, which is used to convert the input image from 3 channels to multiple channels in a downsampling manner to learn useful information; mConv represents the Backbone network that corrects the number of convolutional kernels and is used to match the input after dimension concatenation; the spatial domain output and frequency domain output of the (n-1)th convolutional layer are represented by and respectively. To match the spatio-frequency features of the input dimension splicing, the input dimension of the mConv convolutional kernel in the first 5 layers is 2 times that of the Conv convolutional kernel, and the spatial domain output feature and the frequency domain output feature dimensions are continuously fused after each subsequent convolutional layer; (3) Spatio-frequency domain ensemble learning Ensemble learning is a machine learning method that obtains more accurate and stable prediction results by combining the prediction results of multiple classifiers or models. By combining the output Out of the spatial domain Backbone network and the output mOut of the corrected frequency domain Backbone network according to a certain ratio γ, the specific prediction result can be expressed by formula (8); Prediction=Softmax(γ*Out+(1-γ)*mOut) (8).
Citation Information
Patent Citations
SAR image recognition method based on fusion frequency domain and spatial domain network model
CN112926457A
Crop disease identification method based on FCSA-OfficientNetV2
CN114863278A