An Image Layout Attribute Prediction Method Based on Graph Neural Networks

By converting images into graph structures of nodes and edges, and multi-layer feature extraction and prediction are used to perform multi-layer feature extraction and prediction, the problem of low accuracy of image layout prediction in the prior art is solved, and better structured feature extraction and correlation modeling are achieved.

CN116310508BActive Publication Date: 2025-07-22EAST CHINA NORMAL UNIV

Patent Information

Application Number
CN202310091305.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-06
Publication Date
2025-07-22
Estimated Expiration
2043-02-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively model complex structured features in images, especially the correlation of similar elements in images, resulting in low accuracy in image layout prediction.

Method used

Graph neural network is used to convert images into graph structures of nodes and edges, and high-dimensional layout features are extracted through multi-layer graph convolution and feedforward network, and image layout attribute prediction is performed by combining full-connection layers.

Benefits of technology

Improve the accuracy of image layout prediction, and better extract and fuse structured features in the image, especially the association relationships of similar elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310508B_ABST
    Figure CN116310508B_ABST
Patent Text Reader

Abstract

The present invention discloses an image layout attribute prediction method based on a graph neural network, comprising the following steps: a) determining various attributes to be predicted for the image layout; preprocessing the image and converting it into a graph structure including nodes and edges; b) constructing network blocks; which consist of multiple layers. In each layer, first, it passes through a linear layer, then through a multi-head graph convolutional network, and finally through a two-layer multi-layer perceptron to further enhance the diversity of feature representations; c) for the image processed by multiple layers of such network blocks, the number of output features decreases layer by layer, and the network has a pyramid structure. After several layers, the final feature information of each node can be obtained. d) Finally, these features are input into a fully connected network and a classifier to obtain the predicted values of various layout attributes of the image. Compared with the existing methods, the present invention has strong feature extraction and fusion capabilities, and can improve the accuracy of predicting the image layout to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to an image layout attribute prediction method based on a graph neural network. Background Art

[0002] Image classification and regression tasks are challenging and practical tasks in the field of artificial intelligence, belonging to a type of image processing task, whose goal is to predict the feature values of various attributes in an image. Image layout prediction is a type of image classification and regression task, which is a task that confines the predicted feature values within the framework of "layout". Image layout prediction means: given an image, output its attribute values related to layout.

[0003] Early research on image layout prediction mainly used convolutional neural networks (CNNs) to solve such problems. Image features were extracted through convolutional layers and pooling layers, and then fully connected layers or convolutional layers were used to output prediction values. To better describe the image layout, some methods would also use additional structured information, such as edge detection results or object detection results, to assist in prediction. These models based on simple feature combinations often can only model low-order image and text information and contain more redundant information, and the actual model performance is not good.

[0004] In recent years, researchers have also designed some novel algorithms to improve the performance of chart question answering tasks. For example, generative adversarial networks (GANs) combined with convolutional neural networks are used to solve such problems. GANs have good performance in generating images. It can learn very complex data distributions and improve the training effect; at the same time, since GANs are a type of generative model, they can learn on unlabeled datasets. However, the above methods are difficult to model the structured features in charts, especially the complex structures in images cannot be generalized by the local convolution of CNNs. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide an image layout attribute prediction method based on a graph neural network. In order to fully reflect the structure of the layout in the image, the present invention uses a graph neural network to process an image in the form of a graph structure, which can better extract the structured features in the image, especially the correlation of similar elements in the image.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] An image layout attribute prediction method based on a graph neural network, which specifically includes the following:

[0008] Step 1: For an image whose attributes need to be predicted, convert it into a graph structure with nodes and edges; specifically:

[0009] A1: First, determine the image layout structure attributes, download the ImageNet image dataset, and manually annotate the image dataset for each attribute; then, delimit the training set and the validation set, with the division ratio being 8:2. Among them, the manual annotation is the answer given to the question.

[0010] A2: For a picture with a height of H, a width of W, and a size of H * W * 3, divide it into N D-dimensional vectors, satisfying H * W * 3 = N * D, and regard each vector as a node of a graph; for each node, find its top K nearest neighbors and add an edge between them; thus, a graph G=(V, E) is obtained, where V represents the set of nodes and E represents the set of edges.

[0011] Step 2: Construct the deep learning network block ViG, that is, enter a linear layer, a graph convolutional layer, a linear layer, and two layers of FFN networks in sequence in a block; specifically:

[0012] B1: The graph convolutional part Grapher: Process it in the way of maximum relative graph convolution, including aggregation and update operations; among them, the aggregation operation adopts the multi-head method, that is, apply the attention mechanism, each head has its own different update weights, perform parallel updates and finally aggregate them together.

[0013] B2: The feed-forward network FFN: It is a two-layer multi-layer perceptron, including a hidden layer.

[0014] B3: The ViG network block: Before the feature vector enters the Grapher, first apply a linear layer to convert the features of the nodes into another set of features, then pass through the graph convolutional part, and finally enter the feed-forward network FFN, that is, a ViG block is formed.

[0015] Step 3: Make the graph obtained in Step 1 enter several ViGs with gradually decreasing scales to obtain a high-dimensional layout feature vector; specifically:

[0016] C1: Arrange the ViG blocks in a pyramid structure, set parameters so that the number of output features in each layer of ViG is half of the number of input features. After the initial features pass through several ViGs, a high-dimensional layout feature vector is finally obtained.

[0017] Step 4: Fully connect the high-dimensional layout feature vector with several nodes containing semantic and image layout structure information, and comprehensively obtain the image layout attribute prediction from the output of the nodes; specifically:

[0018] D1: According to the image layout structure attributes determined in A1, establish corresponding attribute nodes one by one, fully connect them with the high-dimensional layout feature vector obtained in Step 3, and add a Softmax function after the attribute nodes.

[0019] D2: The output of the attribute node is the predicted image layout attribute.

[0020] Compared with the prior art, the present invention adopts the above technical solutions and has the following beneficial effects:

[0021] The present invention proposes an image layout attribute prediction method based on a graph neural network, which uses a visual graph neural network to model the dependency relationship between different sub-image blocks in a chart image, and can better extract the structured features in the chart, especially the structural relationship between similar elements in the chart. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flowchart of the present invention;

[0023] FIG. 2 is a schematic flowchart of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The present invention will be further described below in conjunction with specific embodiments and the accompanying drawings.

[0025] Refer to Figure 1 , the present invention specifically includes:

[0026] Step 1: For an image whose attributes need to be predicted, convert it into a graph structure with nodes and edges; specifically:

[0027] (1.1) First, determine the image layout structure attributes according to requirements. In this embodiment, specifically object recognition (predicting object labels, belonging to a classification problem), object position (predicting image coordinates), and object relationship (predicting relationship labels, belonging to a classification problem). Download the ImageNet image dataset, and perform manual annotation on the image dataset for each attribute; then, delimit the training set and the validation set, and the division ratio is 8:2. Among them, the manual annotation is the answer to the question given.

[0028] (1.2) For a picture with a height of H, a width of W, and a size of H * W * 3, divide it into N D-dimensional vectors, satisfying H * W * 3 = N * D, and regard each vector as a node of a graph; for each node, find its top K nearest neighbors and add an edge between them; thus, a graph G=(V, E) is obtained, where V represents the set of nodes and E represents the set of edges;

[0029] Step 2: Construct a deep learning network block ViG, that is, enter a linear layer, a graph convolutional layer, a linear layer, and two FFN networks in sequence in a block; specifically:

[0030] (2.1) Grapher in the graph convolution part: It is processed by using the maximum relative graph convolution, including aggregation and update operations; among them, the aggregation operation adopts a multi-head method, that is, the attention mechanism is applied, and each head has its own different update weights, and parallel updates are performed and finally aggregated together;

[0031] (2.2) Feed-forward network FFN: It is a two-layer multi-layer perceptron, including a hidden layer;

[0032] (2.3) ViG network block: Before the feature vector enters the Grapher, first apply a linear layer to convert the features of the nodes into another set of features, then pass through the graph convolution part, and finally enter the feed-forward network FFN, which is a ViG block;

[0033] Step 3: Make the graph obtained in Step 1 enter several ViGs with gradually decreasing scales to obtain a high-dimensional layout feature vector; specifically:

[0034] (3.1) Arrange the ViG blocks in a pyramid structure, set parameters so that the number of output features in each layer of ViG is half of the number of input features, and finally obtain a high-dimensional layout feature vector after the initial features pass through several ViGs;

[0035] Step 4: Fully connect the high-dimensional layout feature vector with several nodes containing semantic and image layout structure information, and synthesize the outputs of these nodes to obtain the prediction of the image layout attributes; specifically:

[0036] (4.1) According to the image layout structure attributes determined in 1.1, establish corresponding attribute nodes one by one, fully connect them with the high-order feature vectors obtained in 3, and add a Softmax function after the attribute nodes;

[0037] (4.2) The output of the attribute nodes is the category, position of all objects in the image, and the relationships between them.

[0038] Embodiment

[0039] Refer to Figure 2 , after determining the required layout attributes, in this embodiment, the scientific charts collected from the ImageNet dataset and the correct answers manually annotated are preprocessed and then input into the graph neural network constructed by the present invention. First, it is converted into a graph, then after passing through multiple ViGs, a high-order feature vector is obtained, and finally, it is input into the fully connected layer connected to the layout attributes to obtain the category, position of all objects in the graph, and the relationships between them.

[0040] The above is only the preferred embodiment of the present invention. Within the scope defined by the claims of the present invention, certain modifications can be made to it, but all will fall within the protection scope of the present invention.

Claims

1. An image layout attribute prediction method based on a graph neural network, characterized in that, The method includes the following specific steps: Step 1: For the image of the attribute to be predicted, convert it into a graph structure with nodes and edges; specifically: A1: First, determine the image layout structure attribute, download the ImageNet image dataset, and perform manual annotation on the image dataset for each attribute; then, delimit the training set and the validation set, and the division ratio is 8:

2. Among them, the manual annotation is the answer given to the question; A2: For a picture with a height of H, a width of W, and a size of H * W * 3, divide it into N D-dimensional vectors, satisfying H * W * 3 = N * D, and regard each vector as a node of a graph; for each node, find the top K nearest neighbors to it and add an edge between them; thus, a graph G=(V, E) is obtained, where V represents the node set and E represents the edge set; Step 2: Construct a deep learning network block ViG, that is, enter a linear layer, a graph convolutional layer, a linear layer, and two layers of FFN networks in sequence in a block; specifically: B1: The graph convolutional part Grapher: Process it in the way of maximum relative graph convolution, including aggregation and update operations; among them, the aggregation operation adopts a multi-head method, that is, apply the attention mechanism, each head has its own different update weights, perform parallel updates and finally aggregate them together; B2: The feed-forward network FFN: It is a two-layer multi-layer perceptron, including a hidden layer; B3: The ViG network block: Before the feature vector enters the Grapher, first apply a linear layer to convert the features of the nodes into another set of features, then pass through the graph convolutional part, and finally enter the feed-forward network FFN, that is, a ViG block is formed; Step 3: Make the graph obtained in Step 1 enter several ViGs with gradually decreasing scales to obtain a high-dimensional layout feature vector; specifically: C1: Arrange the ViG blocks in a pyramid structure, set parameters so that the number of output features in each layer of ViG is half of the number of input features. After the initial features pass through several ViGs, a high-dimensional layout feature vector is finally obtained; Step 4: Fully connect the high-dimensional layout feature vector with several nodes containing semantic and image layout structure information, and comprehensively obtain the prediction of the image layout attribute from the output of the nodes; specifically: D1: According to the image layout structure attribute determined in A1, establish corresponding attribute nodes one by one, fully connect them with the high-dimensional layout feature vector obtained in Step 3, and add a Softmax function after the attribute nodes; D2: The output of the attribute node is the predicted image layout attribute.

Citation Information

Patent Citations

  • Scene text visual question and answer method based on modal inference graph neural network

    CN113360621A

  • Training graph neural networks using a de-noising objective

    WO2022248735A1

Cited By

  • Text annotation layout generation method based on graph neural network and diffusion model

    CN122655691A