A method for identifying four-corner coordinates in a painting based on TensorFlow

By using the TensorFlow framework and an improved VGG16 neural network model, combined with Dropout layers and the RMSprop algorithm, the automatic recognition of the four corner coordinates of paintings was achieved, solving the problems of tedious manual operation and low accuracy, and improving recognition efficiency and accuracy.

CN117058356BActive Publication Date: 2025-12-05JIANGSU UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311014561.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-11
Publication Date
2025-12-05
Estimated Expiration
2043-08-11

AI Technical Summary

Technical Problem

In existing technologies, extracting the coordinates of the four corners of a painting requires tedious manual operation, is difficult to automate, and has low accuracy, resulting in low efficiency.

Method used

An improved VGG16 neural network model based on TensorFlow is adopted. Through OpenCV image synthesis and preprocessing, combined with Dropout layer to prevent overfitting, and RMSprop algorithm to optimize the learning rate, the automatic recognition of four corner coordinates is achieved.

Benefits of technology

It improves the accuracy and efficiency of recognizing the four corner coordinates of paintings, reduces manual operation, and enhances the robustness and recognition effect of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058356B_ABST
    Figure CN117058356B_ABST
Patent Text Reader

Abstract

The application discloses a method for identifying four-corner coordinates in a painting work based on TensorFlow, which comprises the following steps: synthesizing a network art picture randomly selected from an art picture set with a network indoor scene picture randomly selected from an indoor scene picture set through OpenCV, and forming a to-be-trained pattern set; preprocessing the to-be-trained pattern set, and making a data set together with a plurality of photos with real art pictures, then dividing the data set into a training set, a verification set and a test set; training and learning on the training set by using an improved VGG16 network model, obtaining the four-corner coordinates of the network art picture or the real art picture after a certain number of training rounds, and then randomly selecting the pictures of the test set for testing to monitor the training effect of the current round. The application adopts the TensorFlow framework, combines the improved VGG16 neural network, solves the problem that the four-corner coordinates of an art work with a complex background in a captured picture can only rely on manual work, and greatly improves the efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method based on TensorFlow for recognizing the four corner coordinates of a painting. Background Technology

[0002] As people continue to pursue a better life, they are more inclined to hang artworks at home. Artworks generally need to be framed. However, if Photoshop is used to render the effect during the preview of the framing, it often wastes a lot of time due to a large amount of manual work when dealing with very opinionated customers.

[0003] There are already relatively mature programs for previewing framing effects, but extracting the four corners of the artwork in the image has always been a problem that needs to be solved. Although customers can manually select them, the operation is cumbersome, and manual selection can also result in inaccurate selection. Summary of the Invention

[0004] This invention provides a method for recognizing the four corner coordinates of a painting based on TensorFlow to solve the problems existing in the prior art.

[0005] The technical solutions adopted in this invention are as follows:

[0006] A method based on TensorFlow for identifying the four corner coordinates of a painting includes the following steps:

[0007] 1) Using OpenCV, a random online artwork image from the artwork image set and a random online indoor scene image from the indoor scene image set are combined to form a training image set.

[0008] 2) Preprocess the training image set and create a dataset together with several photos of real artworks. Then, divide this dataset into training, validation and test sets.

[0009] 3) The improved VGG16 network model is trained on the training set. After a certain number of training rounds, the four corner coordinates of the network artwork image or the real artwork image are obtained. Then, images from the test set are randomly selected for testing to monitor the training effect of the current round.

[0010] Furthermore, in step 1), the art image set is selected from the Art Images dataset of the Open Database;

[0011] The indoor scene image set was selected from the RESIDE-Standard Indoor Training Set dataset.

[0012] Further, in step 2), the training image set is preprocessed, which includes data augmentation operations such as blurring, transposing, normalizing, scaling, adding noise, randomizing brightness, and increasing contrast on the images in the training image set.

[0013] Furthermore, in step 3), the four corner coordinates of the online artwork image or the real artwork image are obtained by using the TensorFlow deep learning framework and an improved VGG16 network model.

[0014] Furthermore, the improvement of the VGG16 network model is as follows: the top layer of the VGG16 network model is improved to become four fully connected layers, and a Dropout layer is added sequentially after the top fully connected layer, for a total of four dropout layers.

[0015] The input to the improved VGG16 network model is an image data of 200*200*3, and the output is a vector of length 8, which corresponds to the x and y coordinates of the four corner points of the network or real artwork image.

[0016] Furthermore, the steps for obtaining the four corner coordinates of the network or real artwork image are as follows:

[0017] Step 3-1) Load the pre-trained weights of the improved VGG16 network model using the TensorFlow deep learning framework;

[0018] Step 3-2) Adjust the top layer of the improved VGG16 network model to four fully connected layers, and add a Dropout layer after the top fully connected layer to prevent overfitting.

[0019] Step 3-3) Input the image data into the adjusted improved VGG16 network model for training.

[0020] Furthermore, in step 3-1), the TensorFlow deep learning framework loads the pre-trained weights of the improved VGG16 network model by calling the tf.keras.applications.vgg16 function;

[0021] In step 3-2), the output values ​​of the four added Dropout layers are all set to 0;

[0022] In step 3-3, the input layer of the improved VGG16 network model receives a set of 200*200*3 image data. At this time, the number of channels of the image data is 3. The image data is converted into a matrix, and then the four corner coordinates are extracted by the Harris corner extraction algorithm.

[0023] Furthermore, in step 3), after obtaining the four corner coordinates of the online artwork image or the real artwork image, the mean absolute error is used to determine the four corner coordinates recognized by the improved VGG16 network model.

[0024] Furthermore, the learning rate of the improved VGG16 network model is optimized using the RMSprop algorithm.

[0025] The present invention has the following beneficial effects:

[0026] This invention uses the TensorFlow framework and a modified VGG16 neural network to solve the problem that capturing the four corner coordinates of artworks with complex backgrounds in images can only be done manually, thus greatly improving efficiency.

[0027] The TensorFlow framework is a high-level API for building machine learning models and supports low-level frameworks, allowing for flexible algorithm design and implementation. TensorFlow can be used for various machine learning tasks, such as classification, regression, clustering, dimensionality reduction, and natural language processing. This invention utilizes TensorFlow's graph computation model, where computations are represented as a data flow graph, where each node represents an operation and each edge represents data flow. TensorFlow stores and computes tensors in the graph, implementing the input and output of the data model. Based on this framework, this invention uses deep learning methods to accurately extract the four corner coordinates of artworks contained in images. Attached Figure Description

[0028] Figure 1 To improve the visualization structure of the VGG16 network model.

[0029] Figure 2 The process for inputting image data into the improved VGG16 network model.

[0030] Figure 3 To improve the convergence curve of the VGG16 network model.

[0031] Figure 4 A schematic diagram of the improved VGG16 network model.

[0032] Figure 5 This is a schematic diagram showing the composite of online artwork images and online indoor scene images using OpenCV.

[0033] Figure 6 A photograph of a real artwork.

[0034] Figure 7 In the image, a, b, and c represent the recognition results after 100, 5000, and 10000 training iterations, respectively. Detailed Implementation

[0035] The invention will now be further described with reference to the accompanying drawings.

[0036] like Figure 2 As shown, the neural network model based on TensorFlow for recognizing the four corner coordinates of paintings in images according to the present invention mainly includes the following steps:

[0037] 1) Using OpenCV, composite an online artwork image randomly selected from an art image set with an online indoor scene image randomly selected from an indoor scene image set, such as... Figure 5 And form a set of patterns to be trained;

[0038] In step 1, the training data includes the Indoor TrainingSet dataset from RESIDE-Standard and the Art Images dataset from an open database.

[0039] This invention uses 1399 clear indoor photos from the Indoor Training Set as the background for synthesizing training data. Using OpenCV, it randomly selects one indoor scene image and one artwork image from the Indoor Training Set dataset and the Art Images dataset, respectively, and combines them to simulate real photos containing artworks.

[0040] When randomly selecting and combining scene images and artwork images from the Indoor Training Set and Art Images datasets using OpenCV, image scaling and channel number adjustment are necessary to ensure the two images are the right size and have the same number of channels. Simultaneously, an affine transformation needs to be performed on the randomly selected artwork images within a certain range to simulate perspective effects in the real world. Finally, the affine-transformed artwork images are drawn onto the selected indoor scene images to simulate the display effect of framed objects in a real-world setting.

[0041] 2) Preprocess the training image set and compare it with several photographs of real artworks (such as...). Figure 6 Together, they are made into a dataset, and then this dataset is divided into a training set, a validation set, and a test set;

[0042] In step 2, 5000 images were synthesized using OpenCV for model training. To ensure the model's recognition performance in the real world, 108 real photographs of artworks were also taken and labeled. A total of 5108 images were generated. This invention used 4500 of the 5000 synthesized images as the dataset, with 70% used as the training set, 30% as the validation set, and the remaining 608 images as the test set.

[0043] In step 2, preprocessing involves performing data augmentation operations on the generated image, such as blurring, transposing, normalizing, scaling, adding noise, and randomizing brightness and contrast, to enhance the adaptability of the data, improve the robustness and accuracy of the model, and reduce the risk of overfitting.

[0044] The dataset designed in this invention randomly crops and optically distorts the generated images within a set range to simulate camera deviations during actual shooting, thereby enhancing the model's perception range.

[0045] 3) such as Figure 7 The network model is trained on the training set. After a certain number of training rounds, the four corner coordinates of the network artwork image or the real artwork image are obtained. Then, images from the test set are randomly selected for testing to monitor the training effect of the current round.

[0046] By using the TensorFlow deep learning framework and employing an improved VGG16 network model, the four corner coordinates of online or real artwork images are obtained.

[0047] The improvement to the VGG16 network model is as follows: the top layer of the VGG16 network model is improved to four fully connected layers, and a Dropout layer is added after the top fully connected layer, for a total of four dropout layers.

[0048] The improved VGG16 network model is set to input 200*200*3 image data, and output as a vector of length 8, which represents the x and y coordinates of the four corner points of the network or real artwork image. The specific steps are as follows:

[0049] Step 3-1: You need to use TensorFlow to load the pre-trained weights of the improved VGG16 model, which can be done by calling the tf.keras.applications.vgg16 function.

[0050] Step 3-2: Adjust the top layer of the improved VGG16 model to four fully connected layers, and add a Dropout layer after the top fully connected layer to prevent overfitting.

[0051] In step 3-2, the top layer of the original VGG16 model consists of three fully connected layers, which follow the convolutional layers and record high-level abstract information of the feature maps.

[0052] This invention modifies the original VGG16 model's top layer from three fully connected layers to four, and adds four additional Dropout layers. These layers randomly "deactivate" input neurons, setting the output value of selected neurons to 0. This ensures that each neuron may be "turned off" in certain training samples, forcing the model to learn a variety of different feature combinations. This reduces the model's over-reliance on certain input features, alleviates overfitting, and improves the network's robustness.

[0053] Step 3-3: Input the image data into the adjusted and improved VGG16 network model for training.

[0054] In step 3-3, after the image data is input into the improved VGG16 network model, the input layer receives a set of 200*200*3 image data. At this point, the image data has 3 channels and is converted into a tensor as input for subsequent calculations. When using the TensorFlow framework, it is necessary to define an input layer not included in the improved VGG16 network model to facilitate the processing of the input data.

[0055] Figure 4 To improve the VGG16 network model, the top layer consists of four fully connected layers (layers 1, 2, 3, and 4), with a dropout layer added after each of these layers. The improved VGG16 network model has five large convolutional layers (layers 1, 2, 3, 4, and 5), each followed by a pooling layer (layers 1, 2, 3, 4, and 5), and each pooling layer is followed by a dropout layer.

[0056] After the tensor is input into convolutional layer 1-1 of the first convolutional layer, a 2D convolution operation is performed on the input tensor, and the ReLU activation function is used to activate it, extracting low-level features of the image. The output tensor of convolutional layer 1-1 is 200*200*64. The output tensor of convolutional layer 1-1 is input into convolutional layer 1-2 of the first convolutional layer, and the ReLU activation function is also used to activate it, further extracting low-level features of the image. The output tensor of convolutional layer 1-2 is also 200*200*64.

[0057] Next, the tensor is input into pooling layer 1. Pooling layer 1 performs max pooling on the output of convolutional layers 1-2, reducing the size of the feature map, decreasing computational cost, and preserving the main information of the features. The output tensor of pooling layer 1 has a shape of 100*100*64. The output of pooling layer 1 applies a dropout operation, randomly discarding a certain proportion of neurons to prevent overfitting.

[0058] The 100*100*64 tensor is then fed into the second, third, fourth, and fifth convolutional layers, where convolution and pooling are performed respectively to extract intermediate, mid-to-high-level, and high-level features, ultimately transforming it into a 12*12*512 tensor. This tensor is then fed into the Flatten layer between the fully connected layer and the fourth convolutional layer, thus unfolding the output of the fifth convolutional layer into a one-dimensional vector, which is then input into the fully connected layer 1. The unfolded vector undergoes a fully connected operation and is activated using the ReLU activation function, combining the low-level, mid-level, and high-level features.

[0059] During model training, the network model used in this invention defines the edges and textures of online or real art images as low-level features; the partial shapes and local structures of online or real art images as mid-level features; and the overall shape and structure of online or real art images as high-level features.

[0060] Next, the tensor is input into the Dropout layer after the fully connected layer 1. That is, the Dropout operation is applied to the output of the fully connected layer 1 to randomly discard a certain proportion of neurons in order to prevent overfitting.

[0061] The output of fully connected layer 1 needs to be fed into fully connected layer 2, fully connected layer 3, and fully connected layer 4, and then fully connected and dropout operations are performed in sequence to finally achieve an output tensor with a length of 8.

[0062] The above steps are attached. Figure 1 and 4 As shown.

[0063] This embodiment describes an example input image data of 200*200*3. In actual operation, the size of the image data can be changed, and it is not limited to the initial data size provided in this paragraph.

[0064] In step 3, this invention uses the mean absolute error (MAE) to improve the recognition performance of the VGG16 network model (i.e., the four-corner coordinates).

[0065] Mean absolute error (MAE) is a regression loss function that is the mean of the sum of the absolute values ​​of the differences between the target value and the predicted value. It represents the average error magnitude of the predicted value without considering the direction of the error.

[0066] In step 3, this invention uses the RMSprop algorithm to optimize the model, with a learning rate of 1e-5. RMSProp addresses this issue by maintaining a moving average of the squared gradient and adjusting the weight updates accordingly.

[0067] In step 4, the input image data is processed by the improved VGG16 network model. The image data is progressively abstracted into increasingly higher-level features, ultimately mapped to a vector space of length 8, which outputs the x and y coordinates of the four corner points. The training convergence curve is attached. Figure 3 As shown.

[0068] The above description is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements without departing from the principle of the present invention, and these improvements should also be considered within the scope of protection of the present invention.

Claims

1. A method for identifying four-corner coordinates in a painting based on TensorFlow, characterized in that: Comprising the following steps: 1) Synthesizing a network art picture randomly selected from an art picture set with a network indoor scene picture randomly selected from an indoor scene picture set by OpenCV, and forming a training pattern set; 2) Preprocessing the training pattern set, and making a data set together with several photos with real art pictures, and then dividing the data set into a training set, a validation set and a test set; 3) Training and learning the network model on the training set, and after a certain number of training, obtaining the four corner coordinates of the network art picture or the real art picture, and then randomly selecting the pictures of the test set for testing to monitor the training effect of the current round; In step 3), the four corner coordinates of the network art picture or the real art picture are obtained by using the TensorFlow deep learning framework and the improved VGG16 network model; The improvement point of the improved VGG16 network model is that the top layer of the VGG16 network model is improved into four fully connected layers, and one Dropout layer is added after the fully connected layer of the top layer, and a total of four Dropout layers are added; The input of the improved VGG16 network model is set to 200*200*3 image data, and the output is a vector with a length of 8, which corresponds to the horizontal and vertical coordinates of the four corner points of the network or real art picture; The four corner point acquisition step of the network or real art picture is: Step 3-1) Load the pre-training weight of the improved VGG16 network model using the TensorFlow deep learning framework; Step 3-2) Adjust the top layer of the improved VGG16 network model to four fully connected layers, and add a Dropout layer after the fully connected layer of the top layer to prevent overfitting; Step 3-3) Input the image data into the adjusted improved VGG16 network model for training.

2. The method of identifying the four corner coordinates in a painting based on TensorFlow as claimed in claim 1, wherein: In step 1), the art picture set selects the Art Images data set of the open database; The indoor scene picture set selects the Indoor Training Set data set from RESIDE-Standard. 3.The method of identifying the four-corner coordinates in the painting work based on TensorFlow according to claim 1, wherein: In step 2), the preprocessing of the training pattern set includes blurring, transposing, normalizing, scaling, adding noise, random brightness and contrast data enhancement operations on the pictures in the training pattern set. 4.The method of identifying the four-corner coordinates in a painting based on TensorFlow according to claim 1, wherein: In step 3-1), the TensorFlow deep learning framework loads the pre-training weight of the improved VGG16 network model by calling the tf.keras.applications.vgg16 function; In step 3-2), the output values of the four added Dropout layers are all set to 0; In step 3-3, the input layer of the improved VGG16 network model receives a group of 200*200*3 image data, at this time the channel number of the image data is 3, the image data is converted into a matrix, and then the Harris corner point extraction algorithm is used for four corner coordinate extraction. 5.The method for identifying the four-corner coordinates in the painting work based on TensorFlow according to claim 4, wherein: In step 3), the average absolute error is used to judge the four corner coordinates identified by the improved VGG16 network model after obtaining the four corner coordinates of the network artwork picture or the real artwork picture.

6. The method of identifying the four corner coordinates in a painting based on TensorFlow as claimed in claim 5, wherein: The improved VGG16 network model optimizes the learning rate of the model through the RMSprop algorithm.

Citation Information

Patent Citations

  • Identity card image classification method based on VGG16 network hierarchical optimization

    CN111598157A

  • Improved VGG16 network pig identity recognition method based on transfer learning

    CN113469356A