An image processing method, apparatus, device, and storage medium

By introducing noise data and composition information into image aesthetic evaluation, and using target image convolution network for single prediction, the problem of high computational burden in the prior art is solved, and efficient and accurate image aesthetic evaluation is achieved.

CN113822291BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110661548.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-15
Publication Date
2025-06-27
Estimated Expiration
2041-06-15

AI Technical Summary

Technical Problem

The prior art in image aesthetic evaluation requires multiple sampling and blocking processing, resulting in increased computational burden and reduced efficiency.

Method used

By introducing noise data and composition information, a target graph convolution network is used to perform single prediction, avoid blocking processing and improve efficiency.

Benefits of technology

It realizes more efficient image aesthetic evaluation, reduces the computational burden, ensures the accuracy of the evaluation, and is suitable for real-time image evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113822291B_ABST
    Figure CN113822291B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method, apparatus, device, and storage medium, which relate to fields such as cloud technology and blockchain. The image to be evaluated is scaled, and a first feature map is determined according to the scaled image to be evaluated. The first feature map and the obtained size information are input into a target graph convolutional network. Through the first layer layout perception graph convolutional module of the target graph convolutional network, a first adjacency matrix in Euclidean space is obtained, and graph convolution of the first layer layout perception is performed according to the first adjacency matrix to obtain a second feature map, so as to perform prediction according to the second feature map to obtain an evaluation result of the image to be evaluated. This method retains the original size information, that is, retains the composition information of the image to be evaluated, avoids affecting the aesthetic degree of the image due to the change in size, and ensures the accuracy of image aesthetic evaluation. This method realizes more efficient single-pass prediction through the network, greatly reduces the computational burden, and is more suitable for real-time image evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to an image processing method, apparatus, device, and storage medium. Background Art

[0002] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as the field of image aesthetics evaluation. Image aesthetics evaluation is to use a computer system to simulate the human aesthetic perception of images, and then evaluate the aesthetics of images and quantify the aesthetic degree of images. Image aesthetics evaluation can be specifically applied to fields such as image recommendation, image retrieval, and image retouching.

[0003] Currently, binary classification image aesthetics evaluation is mainly performed through a convolutional neural network model. When using a convolutional neural network for image aesthetics evaluation, multiple image patches are mined from the image and different weights are assigned to them through an attention mechanism, so as to fuse the rectangular box features and image features for image aesthetics evaluation.

[0004] However, this method of fusing multiple image patches through multiple samplings requires multiple predictions through a convolutional neural network, bringing more computational burdens and reducing the efficiency of image aesthetics evaluation. Summary of the Invention

[0005] To solve the above technical problems, this application provides an image processing method, apparatus, device, and storage medium, which avoids affecting the aesthetic degree of the image due to size changes, introduces noise data, and ensures the accuracy of aesthetic evaluation. At the same time, since this method takes into account the composition information in the image to be evaluated, it does not require block processing, realizes more efficient single-pass prediction through the network, greatly reduces the computational burden, improves the efficiency of image aesthetics evaluation, and is more suitable for real-time image evaluation.

[0006] The embodiments of this application disclose the following technical solutions:

[0007] In a first aspect, the embodiments of this application provide an image processing method, and the method includes:

[0008] Obtain the image to be evaluated and the size information of the image to be evaluated, where the size information reflects the composition relationship of the image to be evaluated;

[0009] Perform scaling processing on the image to be evaluated, and determine a first feature map according to the scaled image to be evaluated;

[0010] Input the first feature map and the size information into a target graph convolutional network, and obtain a first adjacency matrix in Euclidean space through the first layer layout perception graph convolutional module of the target graph convolutional network;

[0011] Perform graph convolution with first-layer layout awareness based on the first adjacency matrix to obtain a second feature map;

[0012] Perform prediction based on the second feature map to obtain an evaluation result of the image to be evaluated.

[0013] In a second aspect, an embodiment of the present application provides an image processing apparatus, which includes an acquisition unit, a feature map determination unit, an adjacency matrix determination unit, and an evaluation result determination unit:

[0014] The acquisition unit is configured to acquire an image to be evaluated and size information of the image to be evaluated, where the size information reflects the composition relationship of the image to be evaluated;

[0015] The feature map determination unit is configured to perform scaling processing on the image to be evaluated and determine a first feature map according to the scaled image to be evaluated;

[0016] The adjacency matrix determination unit is configured to input the first feature map and the size information into a target graph convolution network, and obtain a first adjacency matrix in Euclidean space through the first-layer layout awareness graph convolution module of the target graph convolution network;

[0017] The feature map determination unit is further configured to perform graph convolution with first-layer layout awareness based on the first adjacency matrix to obtain a second feature map;

[0018] The evaluation result determination unit is configured to perform prediction based on the second feature map to obtain an evaluation result of the image to be evaluated.

[0019] In a third aspect, an embodiment of the present application provides an electronic device for image processing, where the electronic device includes a processor and a memory:

[0020] The memory is configured to store program code and transmit the program code to the processor;

[0021] The processor is configured to execute the method described in the first aspect according to the instructions in the program code.

[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium is configured to store program code, and the program code is used to execute the method described in the first aspect.

[0023] As can be seen from the above technical solutions, when performing aesthetic evaluation on the image to be evaluated, the image to be evaluated is scaled, and the first feature map is determined based on the scaled image to be evaluated. To avoid changing the original size due to the scaling process, the first feature map and the obtained size information can be input into the target graph convolutional network. The first adjacency matrix in the Euclidean space is obtained through the first layout-aware graph convolutional module of the target graph convolutional network. Graph convolution with the first layout awareness is performed based on the first adjacency matrix to obtain the second feature map, so as to perform prediction based on the second feature map to obtain the evaluation result of the image to be evaluated. Since the original size information is embedded when calculating the first adjacency matrix, the second feature map retains the composition information of the image to be evaluated, avoiding affecting the aesthetic degree of the image due to size changes and introducing noise data, ensuring the accuracy of image aesthetic evaluation. At the same time, since this method takes into account the composition information in the image to be evaluated, it does not require block processing, realizes more efficient single-pass network prediction, greatly reduces the computational burden, improves the efficiency of image aesthetic evaluation, and is more suitable for real-time image evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 Schematic diagram of the system architecture of an image processing method provided by an embodiment of the present application;

[0026] Figure 2 Flowchart of an image processing method provided by an embodiment of the present application;

[0027] Figure 3 Flowchart of an architecture for image processing based on a target graph convolutional network provided by an embodiment of the present application;

[0028] Figure 4 Example diagram of the evaluation result provided by an embodiment of the present application;

[0029] Figure 5 Flowchart of a training method for a target graph convolutional network provided by an embodiment of the present application;

[0030] Figure 6 Flowchart of an architecture for a training method for a target graph convolutional network provided by an embodiment of the present application;

[0031] Figure 7A flowchart for retrieval based on an image processing method provided by an embodiment of the present application;

[0032] Figure 8 A structural diagram of an image processing apparatus provided by an embodiment of the present application;

[0033] Figure 9 A structural diagram of a terminal provided by an embodiment of the present application;

[0034] Figure 10 A structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners

[0035] The embodiments of the present application will be described below with reference to the accompanying drawings.

[0036] Currently, binary classification image aesthetics evaluation is mainly carried out through a convolutional neural network model. When using a convolutional neural network for image aesthetics evaluation, since the convolutional neural network model requires the input image to be of a fixed size, it is necessary to preprocess the image by means of stretching and scaling. This means changes the original size of the image, such as the aspect ratio, and affects the aesthetic degree of the image. Training the model with this means is equivalent to introducing noise data to a certain extent. Therefore, in order to avoid the introduction of noise data while maintaining the original size, an image block division mode can be adopted, that is, multiple image blocks are mined from the image and different weights are assigned to them through an attention mechanism, so as to fuse the rectangular frame features and the image features for image aesthetics evaluation.

[0037] However, this method of fusing multiple image blocks through multiple samplings requires multiple predictions through a convolutional neural network, bringing more computational burdens and reducing the efficiency of image aesthetics evaluation.

[0038] In order to solve the technical problem of how to reduce the computational burden and improve the efficiency of image aesthetics evaluation while retaining the original size, an embodiment of the present application provides an image processing method. In the process of feature extraction (that is, obtaining the second feature map), the original size information is embedded, so that the second feature map retains the composition information of the image to be evaluated, avoiding affecting the aesthetic degree of the image due to the change of size and introducing noise data, and ensuring the accuracy of aesthetics evaluation. At the same time, since this method takes into account the composition information in the image to be evaluated, it does not require block processing, realizes more efficient single-pass prediction through the network, greatly reduces the computational burden, improves the efficiency of image aesthetics evaluation, and is more suitable for real-time image evaluation.

[0039] It should be noted that the method provided by the embodiments of the present application can be applied to various scenarios for aesthetic evaluation of images, such as multiple projects and product applications like photo retouching software, social platforms, webcasts, retrieval platforms, etc. It can perform automated aesthetic evaluation of uploaded images, and then can perform image recommendation, retrieval, retouching, etc. with more semantic and aesthetic quality, improving the user experience. For example, when a user conducts a search on a retrieval platform, it can perform aesthetic evaluation on all retrieved images and return more aesthetically pleasing images to the user according to the evaluation results of the images; another example is in photo retouching software. When a user retouches a photo through the photo retouching software, it can perform aesthetic evaluation on the image uploaded by the user, and thus guide the user to retouch the photo according to the evaluation results of the image to obtain a more aesthetically pleasing image.

[0040] It should be noted that the method provided by the embodiments of the present application may be related to the field of cloud computing. For example, cloud computing refers to the delivery and usage model of IT infrastructure, which means obtaining the required resources in a on-demand and easily scalable manner through the network; in a broad sense, cloud computing refers to the delivery and usage model of services, which means obtaining the required services in a on-demand and easily scalable manner through the network. Such services can be related to IT and software, the Internet, or other services. Cloud computing is the product of the development and integration of traditional computer and network technologies such as Grid Computing, Distributed Computing, Parallel Computing, Utility Computing, Network Storage Technologies, Virtualization, and Load Balance.

[0041] With the development of the Internet, real-time data streams, and the diversification of connected devices, as well as the promotion of demands such as search services, social networks, mobile commerce, and open collaboration, cloud computing has developed rapidly. Different from the previous parallel distributed computing, the emergence of cloud computing will, in concept, drive a revolutionary change in the entire Internet model and enterprise management model.

[0042] The present application is, for example, related to the field of artificial intelligence. Artificial Intelligence (AI) is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also about studying the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making. The present application mainly simulates the human aesthetic perception of images through a computer system and then performs aesthetic evaluation on the images.

[0043] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0044] This application, for example, relates to computer vision technology (Computer Vision, CV) in the field of artificial intelligence. Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes to perform machine vision such as object recognition, tracking, and measurement on targets, and further performing graphics processing to make the computer process images that are more suitable for human eyes to observe or be transmitted to instruments for detection. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, and intelligent transportation, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition. This application mainly obtains the image input to the target graph convolutional network through image processing, and extracts the image semantic features through image semantic understanding, so as to perform image aesthetics evaluation based on the image semantic features.

[0045] This application, for example, relates to machine learning in the field of artificial intelligence. Machine learning (Machine Learning, ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. This application mainly trains the target graph convolutional network through machine learning, so as to use the target graph convolutional network for image aesthetics evaluation.

[0046] It should be noted that the method provided in the examples of this application may also involve blockchain. For example, the data used to train the target graph convolutional network (such as sample evaluation images) can be stored on the blockchain, or in some product applications, the images to be evaluated can be stored on the blockchain.

[0047] See Figure 1 , Figure 1It is a schematic diagram of the system architecture of the image processing method provided by an embodiment of this application. The system architecture includes a terminal 101 and a server 102. Product applications such as photo retouching software, social platforms, webcasts, and retrieval platforms can be installed on the terminal 101. Users can trigger the server 102 to execute the image processing method provided by this application through the product applications installed on the terminal 101. Taking the retrieval platform installed on the terminal 101 as an example, when a user enters a keyword on the retrieval platform to retrieve an image, for example, enters "landscape picture", the server 102 can obtain the image according to the keyword, and thus use the obtained image as the image to be evaluated for image aesthetics evaluation; if taking the photo retouching software installed on the terminal 101 as an example, when a user uploads an image for retouching, the server can use the image uploaded by the user as the image to be evaluated for image aesthetics evaluation, and so on.

[0048] When performing image aesthetics evaluation, the server 102 can obtain the image to be evaluated and the size information of the image to be evaluated. The size information reflects the composition relationship of the image to be evaluated. Among them, the image to be evaluated is the image that needs to be evaluated for image beauty. This image can be uploaded through the terminal 101 when the user triggers this image processing method, or it can be an image uploaded by other users obtained by the server 102. Depending on the application scenario, the image to be evaluated may be different. Taking the retrieval platform as an example, the image to be evaluated is an image uploaded by other users obtained by the server 102 according to the keyword input by the user; if taking the photo retouching software as an example, the image to be evaluated can be the image uploaded by the user.

[0049] The server 102 performs scaling processing on the image to be evaluated and determines a first feature map according to the scaled image to be evaluated, so as to obtain a feature map that meets the requirements of the target convolutional network in terms of size.

[0050] In order to retain the original size of the image to be evaluated, retain the composition information of the image to be evaluated itself, avoid affecting the aesthetic degree of the image due to the change in size, and introduce noise data, the server 102 inputs the first feature map and the size information into the target graph convolutional network, and obtains a first adjacency matrix in the Euclidean space through the first layer layout-aware graph convolutional module of the target graph convolutional network. In this way, by performing graph convolution for the first layer layout perception according to the first adjacency matrix, the obtained second feature map retains the composition information of the image to be evaluated, and according to this second feature map for prediction, the evaluation result of the image to be evaluated is more accurate.

[0051] Among them, the above image processing method can be called image aesthetics evaluation, that is, the aesthetic degree of the image is quantified through the above series of processing processes to obtain an evaluation result. This evaluation result is the quantization result of the aesthetic degree of the image, so as to facilitate the intuitive perception of the aesthetic degree of the image.

[0052] Since this method does not require block processing when considering the composition information in the image to be evaluated, it can achieve more efficient single-pass network prediction, greatly reducing the computational burden, improving the efficiency of image aesthetic evaluation, and being more suitable for real-time image evaluation.

[0053] After obtaining the evaluation result of the image to be evaluated, in different application scenarios, this evaluation result has different functions. For example, in the product application of the retrieval platform, if the server 102 obtains the image to be evaluated according to the keyword and gets the evaluation result of the image to be evaluated, the server 102 returns the retrieval result to the user according to the evaluation result. The retrieval result can be a more aesthetically pleasing image. For example, according to this quantitative evaluation result, the aesthetic degrees of all retrieved images are sorted (in descending order of aesthetic degree), and then the top N images in the ranking are returned to the user.

[0054] It should be noted that Figure 1 The provided application scenario (retrieval platform) is only an example. This method can also be applied to product applications such as photo retouching software, social platforms, and webcasting. The embodiments of the present application do not limit this.

[0055] The image processing method provided by the embodiments of the present application can be executed by the server 102. However, in other embodiments of the present application, the terminal 101 can also have a similar function to the server 102, so as to execute the image processing method provided by the embodiments of the present application, or the image processing method provided by the embodiments of the present application can be jointly executed by the terminal 101 and the server 102. The present embodiment does not limit this.

[0056] It should also be noted that Figure 2 The numbers of the terminal 201 and the server 202 in [[ ]] are only illustrative. According to the implementation requirements, the server 102 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal 101 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, a smart TV, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication means. The present application does not limit this here.

[0057] Subsequently, mainly taking the server as the execution subject as an example, the image processing method provided by the embodiments of the present application will be introduced in detail with reference to the accompanying drawings.

[0058] See Figure 2 , Figure 2 shows a flowchart of an image processing method, and the method includes:

[0059] S201. Obtain the image to be evaluated and the size information of the image to be evaluated, where the size information reflects the composition relationship of the image to be evaluated.

[0060] After the image aesthetic evaluation instruction is triggered, the server can execute the image processing method provided in the embodiments of the present application to perform image aesthetic evaluation. When performing image aesthetic evaluation, the server first obtains the image to be evaluated and the size information of the image to be evaluated. The size information can reflect the composition relationship of the image to be evaluated, so as to retain the original size of the image to be evaluated during the subsequent feature extraction process. In a possible implementation manner, the size information can be aspect ratio information. Assuming that the length of the image to be evaluated is h and the width is w, the aspect ratio information can be expressed as r = h / w. Subsequently, the case where the size information is aspect ratio information will be mainly used as an example for introduction.

[0061] It should be noted that in different application scenarios, the way to trigger the image aesthetic evaluation instruction is different. For example, in the product application of a retrieval platform, if a user inputs a keyword through a terminal and hopes to retrieve the corresponding image, the server executes the image processing method provided in the embodiments of the present application after receiving the keyword, that is, when the server receives the keyword sent by the terminal, it is equivalent to triggering the image aesthetic evaluation instruction. Another example is in the product application of a photography retouching software. When a user uploads an image through a terminal, or after the user completes one retouching of the uploaded image, the user can trigger the server to execute the image processing method provided in the embodiments of the present application through a button on the photography retouching software, so as to perform image aesthetic evaluation on the image uploaded by the user, or on the image after one retouching, and then return the evaluation result to the user to guide the user's retouching.

[0062] S202. Perform scaling processing on the image to be evaluated, and determine a first feature map according to the scaled image to be evaluated.

[0063] In order to perform image aesthetic evaluation on the image to be evaluated, feature extraction can be performed on the image to be evaluated through a network model. The network model can be, for example, a Convolutional Neural Networks (CNN) or a Fully Convolutional Networks (FCN). However, the network model has fixed requirements for the size of the input image. For example, it requires the aspect ratio of the input image to be 1. Therefore, it is necessary to perform scaling processing on the image to be evaluated to obtain an input image that meets the size requirements, and then perform feature extraction through this network model to obtain a first feature map. The size of the first feature map is W×H×C, where W, H, and C respectively represent the length, width, and number of channels of the first feature map.

[0064] See Figure 3As shown in 100, the image to be evaluated is scaled to obtain the scaled image to be evaluated, and then a fully convolutional network is used to extract features from the scaled image to be evaluated to obtain a first feature map.

[0065] S203. Input the first feature map and the size information into the target graph convolutional network, and obtain a first adjacency matrix in the Euclidean space through the first layer layout-aware graph convolutional module of the target graph convolutional network.

[0066] Since the image to be evaluated is scaled in S202, this method changes the size of the image to be evaluated and affects the aesthetic degree of the image to be evaluated. To a certain extent, it is equivalent to introducing noise data. Therefore, the size information can be embedded in the Euclidean space through the graph convolutional network for further feature extraction to update the first feature map extracted in S202. Among them, the target graph convolutional network at least includes a first layer layout-aware graph convolutional module. Graph convolution is a neural network based on unstructured graph data, and its core includes a frequency domain convolutional layer and a pooling layer based on the adjacency matrix for perceiving the image layout.

[0067] According to the characteristics of the target graph convolutional network, in order to update the first feature map, a first adjacency matrix in the Euclidean space can be obtained first, so as to perform graph convolution with the first layer layout awareness according to the first adjacency matrix, thereby obtaining an updated second feature map.

[0068] Compared with extracting features by convolution operation, due to the limited receptive field, convolution operation cannot capture the relationship between distant regions in the image and perceive the image layout. However, the graph convolution operation of the target graph convolutional network can increase the receptive field and solve the problem of difficult-to-capture image layout.

[0069] It should be noted that the first adjacency matrix is used to model the relationship between nodes in the Euclidean space. The Euclidean space refers to the space described by constructing a coordinate system according to the control position of pixels in the image. The first adjacency matrix mainly considers two relationships between nodes, namely content similarity and spatial position relationship. Here, the nodes can be pixel points in the first feature map. Based on this, in a possible implementation manner of this embodiment, the method for calculating the first adjacency matrix can be to divide the first feature map according to the spatial position to obtain the nodes corresponding to the first feature map in the Euclidean space, determine a first content similarity matrix according to the node features corresponding to the nodes, and determine a spatial position correlation matrix according to the position information and size information of the nodes in the first feature map. Among them, the first content similarity matrix is used to reflect the content similarity between nodes in the first feature map, and the spatial position correlation matrix is used to reflect the spatial position relationship between nodes in the first feature map. Furthermore, the sum of the first content similarity matrix and the spatial position correlation matrix is determined as the first adjacency matrix.

[0070] Among them, if the size of the first feature map is W×H×C, then there are W×H spatial positions in each channel. Let L = W×H. Regarding each spatial position as a node, there are a total of L nodes. The node features can be represented by to represent.

[0071] In this embodiment, the superscript "sim (similarity)" is used to represent content similarity, and the superscript "spa (spatial)" is used to represent spatial position relationship. The first content similarity matrix can be expressed as A sim , matrix The spatial position correlation matrix can be expressed as A spa , and the first adjacency matrix constructed in the Euclidean space takes into account the correlations between the above two types of nodes. That is, the first adjacency matrix A c = A sim + A spa , and the subscript "c (coordinate)" is used to represent the Euclidean space.

[0072] Next, the determination methods of the first content similarity matrix and the spatial position correlation matrix will be introduced in detail.

[0073] For the first content similarity matrix, the first content similarity matrix reflects the content similarity between nodes in the first feature map and can be composed of quantization values expressing the content similarity between any two nodes. The quantization value expressing the content similarity between any two nodes can be the similarity calculated between the nodes. Therefore, in this embodiment, the similarity between any two nodes can be calculated according to the node features corresponding to the nodes, and then the first content similarity matrix can be determined according to the similarity between any two nodes. For example, the similarities between any two nodes are arranged in a matrix form in a specific order to obtain the first content similarity matrix.

[0074] Among them, calculating the similarity between any two nodes can be, for example, cosine similarity, Euclidean distance, etc. In this embodiment, the similarity between any two nodes is mainly represented by cosine similarity.

[0075] Taking any two nodes as the i-th node and the j-th node respectively as an example, the cosine similarity calculation formula between the i-th node and the j-th node is as follows:

[0076]

[0077] φ(x i ) = ωx i

[0078] φ′(x j ) = ω′x j

[0079]

[0080] Among them, ∥·∥ represents regularization, <·,·> represents the inner product operation, and X i is the node feature of the i-th node, and X j is the node feature of the j-th node, and φ(x i ) = ωx i and φ′(x j ) = ω′x j are two linear transformation operations, mainly used to improve the generalization ability of the model. Here are the learnable parameters.

[0081] For all nodes, the cosine similarity between any two nodes can be obtained, that is, multiple sim(X i , X j ) are obtained. Multiple sim(X i , X j ) form a matrix. For the obtained matrix, the softmax function is further used to normalize each row. Then the first content similarity matrix is calculated as follows:

[0082]

[0083] For the spatial position correlation matrix, the spatial position correlation matrix is used to reflect the spatial position relationship between nodes in the first feature map and can be composed of quantization values expressing the spatial position relationship between any two nodes. The quantization value expressing the spatial position relationship between any two nodes can be represented by the spatial distance between nodes. Therefore, in this embodiment, the spatial distance between any two nodes can be calculated according to the position information and size information of the nodes in the first feature map, and then the spatial position correlation matrix can be determined according to the spatial distance between any two nodes. For example, the spatial distances between any two nodes are arranged in a matrix form in a specific order, and then the spatial position correlation matrix is obtained.

[0084] Taking any two nodes as the i-th node and the j-th node respectively, their position information in the first feature map is represented as position coordinates and To retain the composition relationship of the image to be evaluated, the size information is embedded in the spatial position relationship between nodes. If the size information is represented by the aspect ratio r, the calculation formula for the spatial distance between the i-th node and the j-th node is as follows:

[0085]

[0086] Among them, are the position coordinates of the i-th node and the j-th node respectively, Similarly, for all nodes, the spatial distance between any two nodes can be obtained, that is, multiple dis(i, j) are obtained, and multiple dis(i, j) form a matrix. For the obtained matrix, after normalizing each row of the matrix through the softmax function, the spatial position correlation matrix can be obtained:

[0087]

[0088] S204. Perform graph convolution of the first-layer layout perception according to the first adjacency matrix to obtain a second feature map.

[0089] According to the characteristics of the first-layer layout perception graph convolution module, graph convolution of the first-layer layout perception can be performed based on the first adjacency matrix to obtain a second feature map.

[0090] Perform graph convolution of the first-layer layout perception based on the node features in the Euclidean space and the first adjacency matrix. The specific formula is:

[0091]

[0092] where is the output of the (m + 1)-th layer of graph convolution, represents the output obtained by passing the first feature map through the m-th layer of graph convolution and is the trainable parameter in the m-th layer of graph convolution. For example, if there are 3 layers of graph convolution layers in this first-layer layout perception graph convolution module, then m = 0, 1, 2. σ(·) represents the activation function, and in this embodiment, σ(·) can use the ReLU function.

[0093] It should be noted that for S203 - S204, reference can be made to Figure 3 as shown in 200, determine the first content similarity matrix according to the node features corresponding to the nodes (as Figure 3 shown in 210), and determine the spatial position correlation matrix according to the position information and size information of the nodes in the first feature map (as Figure 3 shown in 220), and then determine the sum of the first content similarity matrix and the spatial position correlation matrix as the first adjacency matrix (as Figure 3 shown in 230), and then use the first adjacency matrix to perform graph convolution of the first-layer layout perception to obtain a second feature map (as Figure 3 shown in 240).

[0094] S205. Perform prediction according to the second feature map to obtain the evaluation result of the image to be evaluated.

[0095] The second feature map obtained by feature extraction through the above method can more accurately reflect the image semantic information of the image to be evaluated in the Euclidean space. Therefore, the evaluation result of the image to be evaluated can be predicted according to the second feature map.

[0096] Among them, the evaluation results may include one or more combinations of aesthetic distribution, average aesthetic score, and aesthetic classification. The aesthetic score refers to rating the aesthetics of the image to be evaluated. For example, it may include ten levels from 1 to 10 points. The average aesthetic score can be obtained by averaging multiple ratings for the same image to be evaluated; the aesthetic distribution is the probability distribution of the image to be evaluated corresponding to each rating. For example, the probability that the image to be evaluated corresponds to the first level of 1 point is 0.11, the probability that the image to be evaluated corresponds to the second level of 2 points is 0.12,..., the probability that the image to be evaluated corresponds to the ninth level of 9 points is 0.99, and the probability that the image to be evaluated corresponds to the tenth level of 10 points is 0.98; the aesthetic classification is used to identify whether the image to be evaluated is a high-aesthetic-quality image or a low-aesthetic-quality image. Among them, the aesthetic distribution can be represented by the result after histogram normalization. The above three evaluation results can be seen in Figure 4 as shown in Figure 4 in (1), it means that it is determined that the aesthetic classification of the image to be evaluated is a high-aesthetic-quality image, Figure 4 in (2), it means that it is determined that the average aesthetic score of the image to be evaluated is 2.94, Figure 4 in (3), it means that it is determined the aesthetic distribution of the image to be evaluated, and this aesthetic distribution is represented by the histogram marked by the dashed box.

[0097] If the output of the target graph convolutional network is the aesthetic distribution, and actually, the average aesthetic score and the aesthetic classification can also be determined according to the aesthetic distribution. Thus, it can be seen that in this embodiment, the image aesthetic multi-task prediction can be realized through the target graph convolutional network.

[0098] It can be seen from the above technical solutions that when performing aesthetic evaluation on the image to be evaluated, the image to be evaluated is scaled, and the first feature map is determined according to the scaled image to be evaluated. To avoid changing the original size due to the scaling process, the first feature map and the obtained size information can be input into the target graph convolutional network. The first adjacency matrix in the Euclidean space is obtained through the first layer layout-aware graph convolutional module of the target graph convolutional network, and the graph convolution of the first layer layout awareness is performed according to the first adjacency matrix to obtain the second feature map, so as to perform prediction according to the second feature map to obtain the evaluation result of the image to be evaluated. Since the original size information is embedded when calculating the first adjacency matrix, the second feature map retains the composition information of the image to be evaluated, avoiding affecting the aesthetic degree of the image due to the change of size and introducing noise data, ensuring the accuracy of the image aesthetic evaluation. At the same time, since this method takes into account the composition information in the image to be evaluated, it does not require block processing, realizes more efficient single-pass prediction through the network, greatly reduces the calculation burden, improves the efficiency of the image aesthetic evaluation, and is more suitable for real-time image evaluation.

[0099] It should be noted that there are multiple ways to implement S205. In one possible implementation, the evaluation result of the image to be evaluated can be directly predicted from the second feature map through the first layer of layout-aware graph convolution module. Specifically, the evaluation result can be obtained by predicting the second feature map through the fully connected layer of the first layer of layout-aware graph convolution module. Of course, in some cases, the first layer of layout-aware graph convolution module may also include a pooling layer. In this case, the second feature map can be sampled through the pooling layer first, and then the evaluation result can be obtained by prediction through the fully connected layer.

[0100] In another possible implementation, the larger the receptive field, the more capable it is of capturing the relationships between more distant regions in the image to be evaluated, perceiving the image layout, and the image layout affects the aesthetic feeling of the image to be evaluated, which in turn affects the evaluation result. Therefore, in order to capture the relationships between more distant regions in the image to be evaluated as much as possible and perceive the image layout, the node features of multiple nodes can be aggregated onto one node, so as to reflect the node features of all nodes on the second feature map through a small number of nodes, which is equivalent to expanding the receptive field and facilitating the perception of the image layout.

[0101] In this case, the implementation of S205 can be to predict the second feature map through the fully connected layer of the first layer of layout-aware graph convolution module to obtain the first prediction result of the image to be evaluated. The determination method of this first prediction result is the same as the method of directly predicting the evaluation result from the second feature map through the first layer of layout-aware graph convolution module described above, and will not be elaborated here (as Figure 3 shown). Then, in order to further expand the receptive field, capture the relationships between more distant regions in the image to be evaluated, and perceive the image layout, a second layer of layout-aware graph convolution module can be set in the target graph convolution network. Through the second layer of layout-aware graph convolution module of the target graph convolution network, the nodes in the Euclidean space are mapped into the latent variable space. Specifically, the node features corresponding to the nodes in the Euclidean space can be aggregated to obtain the node features corresponding to the nodes in the latent variable space. The number of nodes in the latent variable space is less than the number of nodes in the Euclidean space. Then, the third feature map in the Euclidean space is determined according to the node features corresponding to the nodes in the latent variable space, and the second prediction result of the image to be evaluated is obtained by predicting the third feature map through the fully connected layer of the second layer of layout-aware graph convolution module. Furthermore, the evaluation result is jointly determined according to the first prediction result and the second prediction result.

[0102] Of course, in some cases, the second layer of layout-aware graph convolution module may also include a pooling layer. In this case, the third feature map can be sampled through the pooling layer first, and then the second prediction result can be obtained by prediction through the fully connected layer.

[0103] The target graph convolutional network adopted in the embodiments of the present application can not only capture the relationships between objects at different positions in the image in the Euclidean space, but also capture more advanced semantic relationships in the latent variable space, and more reasonably capture the parts that need to be concerned in the aesthetic evaluation of the image.

[0104] Since the nodes in the Euclidean space are mapped to the latent variable space, all node features are aggregated onto a small number of nodes, and the image layout of the entire image to be evaluated is reflected by the small number of nodes. In this way, the third feature map in the Euclidean space is determined according to the node features corresponding to the nodes in the latent variable space, and the third feature map can better reflect the image layout of the entire image to be evaluated, thereby obtaining a more accurate evaluation result.

[0105] It should be noted that the node features in the latent variable space obtained after aggregation can be expressed as where are network learnable parameters, K is the number of nodes in the latent variable space, L is the number of nodes in the Euclidean space, x j is the node feature in the Euclidean space, and b ij is the mapping coefficient for mapping the nodes in the Euclidean space to the latent variable space. The number of nodes in the latent variable space is less than the number of nodes in the Euclidean space, and the number of nodes K in the latent variable space and the number of nodes L in the Euclidean space can have a certain quantitative relationship. For example, K is one Nth of L. In some cases, through actual verification, when N = 4, that is the evaluation result is more accurate.

[0106] In some possible implementation manners, the method for determining the third feature map in the Euclidean space according to the node features corresponding to the nodes in the latent variable space may be to obtain the second adjacency matrix in the latent variable space according to the node features corresponding to the nodes in the latent variable space, perform graph convolution of the second layer layout perception according to the second adjacency matrix in the latent variable space to obtain the response features. Then, perform inverse mapping on the response features to obtain the third feature map.

[0107] Due to the non-Euclidean geometric property of the latent variable space, only the content similarity between nodes is considered and the spatial position relationship is ignored when calculating the second adjacency matrix. Therefore, the method for calculating the second adjacency matrix may be to determine the second content similarity matrix according to the node features corresponding to the nodes in the latent variable space, and then use the second content similarity matrix as the second adjacency matrix.

[0108] It should be noted that the calculation method of the second content similarity matrix is similar to that of the first content similarity matrix, except that the node features in the latent variable space are used when calculating the second content similarity matrix. In this embodiment, the latent variable space can be represented by the subscript "l (latent)", and the second adjacency matrix can be represented by Al representation

[0109] Perform graph convolution for the second - layer layout perception based on the node features and adjacency matrix in this implicit space. The specific formula is as follows:

[0110]

[0111] where is the output of the (m + 1)-th layer of graph convolution, represents the output of the second feature map after passing through the m-th layer of graph convolution are the trainable parameters in the m-th layer of graph convolution. For example, in this second - layer layout perception graph convolution module, there is 1 layer of graph convolution layer, then m = 0. σ(·) represents the activation function. In this embodiment, σ(·) can use the ReLU function.

[0112] Obtain the response feature After that, inverse - map it back to the Euclidean space, and the third feature map output by the second - layer layout perception graph convolution module can be obtained: where B Τ represents the inverse - mapping operation.

[0113] See Figure 3 As shown, the second - layer layout perception graph convolution module can be seen in Figure 3 shown in 300 in. Map the nodes in the Euclidean space to the latent variable space to obtain the node features in the latent variable space (such as Figure 3 shown in 310 in). Determine the second content similarity matrix according to the node features in the latent variable space, and then obtain the second adjacency matrix according to the second content similarity matrix (such as Figure 3 shown in 320 in). Perform graph convolution for the second - layer layout perception according to the second adjacency matrix to obtain the response feature (such as Figure 3 shown in 330 in). Then perform inverse - mapping on this response feature to obtain the third feature map (such as Figure 3 shown in 340 in). Through the fully - connected layer of the second - layer layout perception graph convolution module, predict the third feature map to obtain the second prediction result of the image to be evaluated (such as Figure 3 shown). Furthermore, determine the evaluation result according to the first prediction result and the second prediction result jointly (such as Figure 3 shown in 400 in).

[0114] This application uses a fully convolutional network combined with an object graph convolutional network to capture long-range relationships in the image to be evaluated. Of course, dilated convolution or other methods can also be used to replace the fully convolutional network for feature extraction to obtain the first feature map. The object graph convolutional network uses more layers of layout-aware graph convolutional modules, or replaces it with various other effective new model structures, such as the combination of a graph convolutional network and a spatial pyramid structure / pooling. It can be simplified or extended according to the limitations of model memory occupancy and efficiency requirements in actual applications for Figure 3 the network structure shown.

[0115] When the output of the object graph convolutional network is an aesthetic distribution, that is, the first prediction result output by the first layer of layout-aware graph convolutional module is the first aesthetic distribution, and the second prediction result output by the second layer of layout-aware graph convolutional module is the second aesthetic distribution, the final evaluation result includes at least the aesthetic distribution, such as the target aesthetic distribution. At this time, in order to obtain the target aesthetic distribution, the average value of the first aesthetic distribution and the second aesthetic distribution can be calculated, and the average value of the first aesthetic distribution and the second aesthetic distribution is used as the target aesthetic distribution of the image to be evaluated.

[0116] For example, the first aesthetic distribution and the second aesthetic distribution are p c , p L , p c , S represents the level of aesthetic score, generally 10, then the target aesthetic distribution can be expressed as

[0117] Then, the average aesthetic score of the image to be evaluated is determined according to the target aesthetic distribution, and the aesthetic classification of the image to be evaluated is determined according to the average aesthetic score.

[0118] Other tasks of predicting the aesthetics of the image according to the target aesthetic distribution. For example, the target aesthetic distribution is expressed as p = [p (1) , p (2) , …, p (S) , p (S) is the probability of the image to be evaluated corresponding to the level S. The average aesthetic score can be calculated according to the target aesthetic distribution For aesthetic category c ∈ {0, 1}, in this embodiment, 0 and 1 are used to represent low-quality aesthetic images and high-quality aesthetic images respectively. The aesthetic category is predicted by the relationship between the average aesthetic score and the threshold. If the average aesthetic score is greater than or equal to the threshold, it can be considered that the image to be evaluated is relatively aesthetically pleasing and can be considered a high-quality aesthetic image. At this time, c = 1, otherwise c = 0. Among them, the threshold can be set according to actual needs. For example, the threshold can be S / 2.

[0119] Since the average aesthetic score and aesthetic classification can also be determined according to the aesthetic distribution, the target graph convolutional network can be used in this embodiment to achieve multi-task prediction of image aesthetics.

[0120] It can be understood that, in order to Figure 2 perform image aesthetics evaluation based on the target graph convolutional network in the corresponding embodiment, the original size of the image to be evaluated is retained to avoid affecting the aesthetic degree of the image due to size changes and ensure the accuracy of image aesthetics evaluation. Before applying the target graph convolutional network, the target graph convolutional network can be trained first. The training method of the target graph convolutional network can be referred to Figure 5 as shown, including:

[0121] S501. Obtain a sample evaluation image and the size information of the sample evaluation image.

[0122] S502. Perform scaling processing on the sample evaluation image, and determine a first sample feature map according to the scaled sample evaluation image.

[0123] S503. Input the first sample feature map and the size information of the sample evaluation image into the target graph convolutional network, and obtain a first sample adjacency matrix in the Euclidean space through the first-layer layout-aware graph convolutional module of the target graph convolutional network.

[0124] S504. Perform graph convolution with first-layer layout awareness according to the first sample adjacency matrix to obtain a second sample feature map.

[0125] S505. Perform prediction according to the second sample feature map to obtain a predicted evaluation result.

[0126] Refer to Figure 6 as shown, performing scaling processing on the sample evaluation image and obtaining a first sample feature map is equivalent to preprocessing the sample evaluation image, and then inputting the first sample feature map obtained after preprocessing into the target graph convolutional network to obtain a predicted evaluation result.

[0127] To retain the original size of the sample evaluation image, avoid affecting the aesthetic degree of the image due to size changes, and ensure the accuracy of image aesthetics evaluation, the size information of the sample evaluation image can be used as an input (refer to Figure 6 the dotted line shown in Figure 1 and input it into the target graph convolutional network together with the first sample feature

[0128] It should be noted that the specific implementation manners of S501 - S505 are similar to those of S201 - S205 in the Figure 2 corresponding embodiment, and will not be elaborated here.

[0129] S506. Construct a target loss function based on the predicted evaluation result and the labeled evaluation result of the sample evaluation image.

[0130] S507. Train the target graph convolutional network according to the target loss function.

[0131] After obtaining the predicted evaluation result, a target loss function can be constructed based on the predicted evaluation result and the labeled evaluation result of the sample evaluation image. Then, the model parameters of the target graph convolutional network can be adjusted according to the target loss function until the target loss function is minimized, thereby completing the training of the target graph convolutional network.

[0132] As can be seen from the above technical solution, when training the target graph convolutional network, the size information of the sample evaluation image is used as input and input into the target graph convolutional network together with the first sample feature for further feature extraction, retaining the composition information of the sample evaluation image, being able to effectively capture the aesthetic attributes of the image, so that the trained target graph convolutional network can perform image aesthetic evaluation more accurately. In addition, since this method takes into account the composition information in the sample evaluation image without the need for block processing, it realizes more efficient single-pass network prediction, greatly reducing the computational burden and improving the efficiency of image aesthetic evaluation, and is more suitable for real-time image evaluation. Figure 1 It should be noted that the evaluation result can include one or more combinations of aesthetic distribution, average aesthetic score, and aesthetic classification. For example

[0133] as shown Figure 6 where the evaluation result includes aesthetic distribution, average aesthetic score, and aesthetic classification. Figure 6

[0134] When the predicted evaluation result and the labeled evaluation result are aesthetic distributions, the target loss function can be the loss function of the Earth Mover's Distance (EMD). This loss function mainly calculates the minimum cost of transforming one of the two probability distributions into the other. That is to say, the target graph convolutional network is trained based on the target loss function of the aesthetic distribution distance. At this time, the calculation formula of the target loss function can be expressed as:

[0135]

[0136] The cumulative aesthetic distribution function obtained by accumulating the (i) probability corresponding to level k of the sample evaluation image, p is the cumulative aesthetic distribution function obtained by accumulating the labeled evaluation result, S is the divided aesthetic score level, and r is a hyperparameter in the target loss function that can be set according to requirements during training. For example, r can be set to 2.

[0137] In this embodiment, the target graph convolutional network is trained with a target loss function based on the aesthetic distribution distance. Therefore, the output of the target graph convolutional network is the aesthetic distribution. In fact, the average aesthetic score and aesthetic classification can also be determined according to the aesthetic distribution. Thus, the target graph convolutional network can be used to implement multi-task prediction of image aesthetics, that is, an evaluation result including the aesthetic distribution, average aesthetic score, and aesthetic classification can be obtained through one model. And it significantly outperforms other existing reference systems in terms of the metrics of the benchmark dataset in the field of image aesthetic evaluation.

[0138] After the target graph convolutional network is trained, online services can be provided according to the Figure 2 application process therein, which will not be elaborated here.

[0139] Next, the image processing method provided by the embodiments of this application will be introduced in combination with actual application scenarios. When a user conducts a search on a retrieval platform, image aesthetic evaluation can be performed on all the retrieved images, and then more aesthetically pleasing images can be returned to the user according to the image evaluation results. The specific implementation process can be referred to Figure 7 , and the method includes:

[0140] S701. The user inputs a keyword on the terminal and triggers the search.

[0141] S702. The server obtains the keyword and the images to be evaluated corresponding to the keyword.

[0142] S703. Perform scaling processing on the images to be evaluated, and determine the first feature map according to the scaled images to be evaluated.

[0143] S704. Determine the first content similarity matrix according to the node features of the nodes in the first feature map.

[0144] S705. Determine the spatial position correlation matrix according to the position information and size information of the nodes in the first feature map.

[0145] S706. Determine the sum of the first content similarity matrix and the spatial position correlation matrix as the first adjacency matrix.

[0146] S707. Perform graph convolution of the first layer of layout perception according to the first adjacency matrix to obtain the second feature map.

[0147] S708. Through the fully connected layer of the first layer of layout perception graph convolution module, predict the second feature map to obtain the first prediction result of the images to be evaluated.

[0148] S709. Map the nodes in the Euclidean space to the latent variable space according to the first feature map.

[0149] S710. Determine the second content similarity matrix according to the node features corresponding to the nodes in the latent variable space.

[0150] S711. Use the second content similarity matrix as the second adjacency matrix.

[0151] S712. Perform graph convolution with second-layer layout awareness based on the second adjacency matrix in the latent variable space to obtain response features.

[0152] S713. Perform inverse mapping on the response features to obtain the third feature map.

[0153] S714. Through the fully connected layer of the second-layer layout awareness graph convolution module, predict the third feature map to obtain the second prediction result of the image to be evaluated.

[0154] S715. Determine the evaluation result according to the first prediction result and the second prediction result.

[0155] S716. Return the retrieval result according to the evaluation result.

[0156] Based on Figure 2 the image processing method provided in the corresponding embodiment, the embodiment of the present application further provides an image processing device. Refer to Figure 8 , the image processing device 800 includes an acquisition unit 801, a feature map determination unit 802, an adjacency matrix determination unit 803, and an evaluation result determination unit 804:

[0157] The acquisition unit 801 is configured to acquire the image to be evaluated and the size information of the image to be evaluated, and the size information reflects the composition relationship of the image to be evaluated;

[0158] The feature map determination unit 802 is configured to perform scaling processing on the image to be evaluated, and determine the first feature map according to the scaled image to be evaluated;

[0159] The adjacency matrix determination unit 803 is configured to input the first feature map and the size information into the target graph convolution network, and obtain the first adjacency matrix in the Euclidean space through the first-layer layout awareness graph convolution module of the target graph convolution network;

[0160] The feature map determination unit 802 is further configured to perform graph convolution with first-layer layout awareness according to the first adjacency matrix to obtain the second feature map;

[0161] The evaluation result determination unit 804 is configured to perform prediction according to the second feature map to obtain the evaluation result of the image to be evaluated.

[0162] In a possible implementation manner, the adjacency matrix determination unit 803 is specifically configured to:

[0163] Divide the first feature map according to spatial positions to obtain nodes corresponding to the first feature map in the Euclidean space;

[0164] Determine a first content similarity matrix according to the node features corresponding to the nodes, and determine a spatial position correlation matrix according to the position information and the size information of the nodes in the first feature map. The first content similarity matrix is used to reflect the content similarity between nodes in the first feature map, and the spatial position correlation matrix is used to reflect the spatial position relationship between nodes in the first feature map;

[0165] Determine the sum of the first content similarity matrix and the spatial position correlation matrix as the first adjacency matrix.

[0166] In a possible implementation manner, the adjacency matrix determination unit 803 is specifically configured to:

[0167] Calculate the similarity between any two of the nodes according to the node features;

[0168] Determine the first content similarity matrix according to the similarity between any two of the nodes.

[0169] In a possible implementation manner, the adjacency matrix determination unit 803 is specifically configured to:

[0170] Calculate the spatial distance between any two of the nodes according to the position information and the size information of the nodes in the first feature map;

[0171] Determine the spatial position correlation matrix according to the spatial distance between any two of the nodes.

[0172] In a possible implementation manner, the evaluation result determination unit 804 is specifically configured to:

[0173] Perform prediction on the second feature map through the fully connected layer of the first layer layout-aware graph convolutional module to obtain the evaluation result of the image to be evaluated.

[0174] In a possible implementation manner, the evaluation result determination unit 804 is specifically configured to:

[0175] Perform prediction on the second feature map through the fully connected layer of the first layer layout-aware graph convolutional module to obtain the first prediction result of the image to be evaluated;

[0176] Through the second-layer layout-aware graph convolution module of the target graph convolutional network, aggregate the node features corresponding to the nodes in the Euclidean space to obtain the node features corresponding to the nodes in the latent variable space, where the number of nodes in the latent variable space is less than the number of nodes in the Euclidean space;

[0177] Determine the third feature map in the Euclidean space according to the node features corresponding to the nodes in the latent variable space;

[0178] Predict the third feature map through the fully connected layer of the second-layer layout-aware graph convolution module to obtain the second prediction result of the image to be evaluated;

[0179] Determine the evaluation result according to the first prediction result and the second prediction result.

[0180] In a possible implementation manner, the evaluation result determination unit 804 is specifically configured to:

[0181] Obtain the second adjacency matrix in the latent variable space according to the node features corresponding to the nodes in the latent variable space;

[0182] Perform graph convolution with second-layer layout awareness according to the second adjacency matrix in the latent variable space to obtain response features;

[0183] Perform inverse mapping on the response features to obtain the third feature map.

[0184] In a possible implementation manner, the evaluation result determination unit 804 is specifically configured to:

[0185] Determine the second content similarity matrix according to the node features corresponding to the nodes in the latent variable space;

[0186] Use the second content similarity matrix as the second adjacency matrix.

[0187] The evaluation result includes one or more combinations of aesthetic distribution, average aesthetic score, and aesthetic classification.

[0188] In a possible implementation manner, if the first prediction result is the first aesthetic distribution and the second prediction result is the second aesthetic distribution, the evaluation result determination unit 804 is specifically configured to:

[0189] Use the average of the first aesthetic distribution and the second aesthetic distribution as the target aesthetic distribution of the image to be evaluated;

[0190] Determine the average aesthetic score of the image to be evaluated according to the target aesthetic distribution;

[0191] Determine the aesthetic classification of the image to be evaluated according to the average aesthetic score.

[0192] In a possible implementation, the device further includes a training unit:

[0193] The training unit is configured to obtain a sample evaluation image and size information of the sample evaluation image; perform a scaling process on the sample evaluation image, and determine a first sample feature map according to the scaled sample evaluation image; input the first sample feature map and the size information of the sample evaluation image into the target graph convolutional network, and obtain a first sample adjacency matrix in the Euclidean space through the first layer layout-aware graph convolutional module of the target graph convolutional network; perform graph convolution with the first layer layout awareness according to the first sample adjacency matrix to obtain a second sample feature map; perform prediction according to the second sample feature map to obtain a predicted evaluation result; construct a target loss function according to the predicted evaluation result and the labeled evaluation result of the sample evaluation image; and train the target graph convolutional network according to the target loss function.

[0194] In a possible implementation, the predicted evaluation result and the labeled evaluation result are aesthetic distributions.

[0195] As can be seen from the above technical solutions, when performing aesthetic evaluation on an image to be evaluated, the image to be evaluated is scaled, and a first feature map is determined according to the scaled image to be evaluated. To avoid changing the original size due to the scaling process, the first feature map and the obtained size information can be input into the target graph convolutional network, and a first adjacency matrix in the Euclidean space is obtained through the first layer layout-aware graph convolutional module of the target graph convolutional network. Graph convolution with the first layer layout awareness is performed according to the first adjacency matrix to obtain a second feature map, so as to perform prediction according to the second feature map to obtain the evaluation result of the image to be evaluated. Since the original size information is embedded when calculating the first adjacency matrix, the second feature map retains the composition information of the image to be evaluated, avoiding affecting the aesthetic degree of the image due to the change of size and introducing noise data, and ensuring the accuracy of image aesthetic evaluation. At the same time, since this method takes into account the composition information in the image to be evaluated, it does not require block processing, realizes more efficient single-pass prediction through the network, greatly reduces the computational burden, improves the efficiency of image aesthetic evaluation, and is more suitable for real-time image evaluation.

[0196] The embodiment of the present application further provides an electronic device for image processing. The electronic device may be a terminal. Taking the terminal as a smart phone as an example:

[0197] Figure 9 Shown is a block diagram of a part of the structure of a smart phone related to the terminal provided by the embodiment of the present application. Refer to Figure 9, the smart phone includes components such as a Radio Frequency (RF) circuit 910, a memory 920, an input unit 930, a display unit 940, a sensor 950, an audio circuit 960, a wireless fidelity (WiFi) module 970, a processor 980, and a power supply 990. The input unit 930 may include a touch panel 931 and other input devices 932. The display unit 940 may include a display panel 941. The audio circuit 960 may include a speaker 961 and a microphone 962. Those skilled in the art can understand that Figure 9 the structure of the smart phone shown in

[0198] does not limit the smart phone, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. The memory 920 can be used to store software programs and modules. The processor 980 executes various functional applications and data processing of the smart phone by running the software programs and modules stored in the memory 920. The memory 920 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the smart phone (such as audio data, a phone book, etc.). In addition, the memory 920 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0199] The processor 980 is the control center of the smart phone, connects various parts of the entire smart phone using various interfaces and lines, and executes various functions of the smart phone and processes data by running or executing the software programs and / or modules stored in the memory 920, and by calling the data stored in the memory 920. Optionally, the processor 980 may include one or more processing units; preferably, the processor 980 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 980.

[0200] In this embodiment, the processor 980 in the terminal may execute the following steps:

[0201] Obtain an image to be evaluated and size information of the image to be evaluated, where the size information reflects the composition relationship of the image to be evaluated;

[0202] Scale the image to be evaluated, and determine the first feature map according to the scaled image to be evaluated;

[0203] Input the first feature map and the size information into the target graph convolutional network, and obtain the first adjacency matrix in Euclidean space through the first layer layout-aware graph convolutional module of the target graph convolutional network;

[0204] Perform graph convolution with the first layer layout awareness according to the first adjacency matrix to obtain a second feature map;

[0205] Make a prediction according to the second feature map to obtain the evaluation result of the image to be evaluated.

[0206] The electronic device may further include a server. Embodiments of the present application also provide a server. Please refer to Figure 10 as shown Figure 10 is a structural diagram of the server 1000 provided by the embodiment of the present application. The server 1000 may vary greatly due to configuration or performance differences, and may include one or more central processing units (Central Processing Units, abbreviated as CPUs) 1022 (for example, one or more processors) and a memory 1032, and one or more storage media 1030 for storing application programs 1042 or data 1044 (for example, one or more mass storage devices). Among them, the memory 1032 and the storage media 1030 may be transient storage or persistent storage. The program stored in the storage media 1030 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processor 1022 may be set to communicate with the storage media 1030 and execute a series of instruction operations in the storage media 1030 on the server 1000.

[0207] The server 1000 may further include one or more power supplies 1026, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1058, and / or one or more operating systems 1041, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0208] In this embodiment, the central processor 1022 in the server 1000 may perform the following steps:

[0209] Obtain the image to be evaluated and the size information of the image to be evaluated, where the size information reflects the composition relationship of the image to be evaluated;

[0210] Perform scaling processing on the image to be evaluated, and determine a first feature map according to the scaled image to be evaluated;

[0211] Input the first feature map and the size information into a target graph convolutional network, and obtain a first adjacency matrix in Euclidean space through the first layer layout-aware graph convolutional module of the target graph convolutional network;

[0212] Perform graph convolution with first layer layout awareness according to the first adjacency matrix to obtain a second feature map;

[0213] Perform prediction according to the second feature map to obtain an evaluation result of the image to be evaluated.

[0214] According to one aspect of the present application, there is provided a computer-readable storage medium for storing program code for executing the image processing method described in each of the foregoing embodiments.

[0215] According to one aspect of the present application, there is provided a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations of the foregoing embodiments.

[0216] The terms "first", "second", "third", "fourth", etc. (if any) in the description of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0217] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.

[0218] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0219] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0220] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0221] As mentioned above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of this application.

Claims

1. An image processing method, characterized in that, The method includes: Obtaining an image to be evaluated and size information of the image to be evaluated, where the size information reflects the composition relationship of the image to be evaluated; Performing a scaling process on the image to be evaluated, and determining a first feature map according to the scaled image to be evaluated; Inputting the first feature map and the size information into a target graph convolutional network, and obtaining a first adjacency matrix in Euclidean space through the first layer layout-aware graph convolutional module of the target graph convolutional network; Performing graph convolution with the first layer layout awareness according to the first adjacency matrix to obtain a second feature map; Performing prediction according to the second feature map to obtain an evaluation result of the image to be evaluated; Inputting the first feature map and the size information into a target graph convolutional network, and obtaining a first adjacency matrix in Euclidean space through the first layer layout-aware graph convolutional module of the target graph convolutional network, including: Dividing the first feature map according to spatial positions to obtain nodes corresponding to the first feature map in Euclidean space; Determining a first content similarity matrix according to the node features corresponding to the nodes, and determining a spatial position correlation matrix according to the position information of the nodes in the first feature map and the size information, where the first content similarity matrix is used to reflect the content similarity between nodes in the first feature map, and the spatial position correlation matrix is used to reflect the spatial position relationship between nodes in the first feature map; Determining the sum of the first content similarity matrix and the spatial position correlation matrix as the first adjacency matrix.

2. The method according to claim 1, wherein The determining the first content similarity matrix according to the node features corresponding to the nodes includes: Calculating the similarity between any two of the nodes according to the node features; Determining the first content similarity matrix according to the similarity between any two of the nodes.

3. The method according to claim 1, characterized in that, The determining the spatial position correlation matrix according to the position information of the nodes in the first feature map and the size information includes: Calculating the spatial distance between any two of the nodes according to the position information of the nodes in the first feature map and the size information; Determining the spatial position correlation matrix according to the spatial distance between any two of the nodes.

4. The method according to claim 1, wherein The performing prediction according to the second feature map to obtain an evaluation result of the image to be evaluated includes: Performing prediction on the second feature map through the fully connected layer of the first layer layout-aware graph convolutional module to obtain an evaluation result of the image to be evaluated.

5. The method according to claim 1, characterized in that, The performing prediction according to the second feature map to obtain an evaluation result of the image to be evaluated includes: Performing prediction on the second feature map through the fully connected layer of the first layer layout-aware graph convolutional module to obtain a first prediction result of the image to be evaluated; Aggregating the node features corresponding to the nodes in the second feature map through the second layer layout-aware graph convolutional module of the target graph convolutional network to obtain node features corresponding to the nodes in the latent variable space, where the number of nodes in the latent variable space is less than the number of nodes in the second feature map; Determining a third feature map in Euclidean space according to the node features corresponding to the nodes in the latent variable space; Predict the third feature map through the fully connected layer of the second-layer layout-aware graph convolution module to obtain a second prediction result of the image to be evaluated; Determine the evaluation result according to the first prediction result and the second prediction result.

6. The method according to claim 5, wherein The determining the third feature map in the Euclidean space according to the node features corresponding to the nodes in the latent variable space includes: Obtain a second adjacency matrix in the latent variable space according to the node features corresponding to the nodes in the latent variable space; Perform graph convolution with second-layer layout awareness according to the second adjacency matrix in the latent variable space to obtain response features; Perform inverse mapping on the response features to obtain the third feature map.

7. The method according to claim 6, wherein The obtaining the second adjacency matrix in the latent variable space according to the node features corresponding to the nodes in the latent variable space includes: Determine a second content similarity matrix according to the node features corresponding to the nodes in the latent variable space; Use the second content similarity matrix as the second adjacency matrix.

8. The method according to any one of claims 1 to 7, characterized in that The evaluation result includes one or more combinations of aesthetic distribution, average aesthetic score, and aesthetic classification.

9. The method according to claim 5, characterized in that, If the first prediction result is a first aesthetic distribution and the second prediction result is a second aesthetic distribution, the determining the evaluation result according to the first prediction result and the second prediction result includes: Use the average of the first aesthetic distribution and the second aesthetic distribution as the target aesthetic distribution of the image to be evaluated; Determine the average aesthetic score of the image to be evaluated according to the target aesthetic distribution; Determine the aesthetic classification of the image to be evaluated according to the average aesthetic score.

10. The method according to claim 1, characterized in that The training method of the target graph convolution network includes: Obtain a sample evaluation image and size information of the sample evaluation image; Perform scaling processing on the sample evaluation image, and determine a first sample feature map according to the scaled sample evaluation image; Input the first sample feature map and the size information of the sample evaluation image into the target graph convolution network, and obtain a first sample adjacency matrix in the Euclidean space through the first-layer layout-aware graph convolution module of the target graph convolution network; Perform graph convolution with first-layer layout awareness according to the first sample adjacency matrix to obtain a second sample feature map; Perform prediction according to the second sample feature map to obtain a predicted evaluation result; Construct a target loss function according to the predicted evaluation result and the labeled evaluation result of the sample evaluation image; Train the target graph convolution network according to the target loss function.

11. The method according to claim 10, characterized in that, The predicted evaluation result and the labeled evaluation result are aesthetic distributions.

12. An image processing apparatus, characterized in that, The device includes an acquisition unit, a feature map determination unit, an adjacency matrix determination unit, and an evaluation result determination unit: The acquisition unit is used to acquire the image to be evaluated and the size information of the image to be evaluated, and the size information reflects the composition relationship of the image to be evaluated; The feature map determination unit is used to perform scaling processing on the image to be evaluated, and determine a first feature map according to the scaled image to be evaluated; The adjacent matrix determination unit is configured to input the first feature map and the size information into a target graph convolutional network, and obtain a first adjacent matrix in the Euclidean space through the first layer layout-aware graph convolutional module of the target graph convolutional network; The feature map determination unit is further configured to perform graph convolution with the first layer layout awareness according to the first adjacent matrix to obtain a second feature map; The evaluation result determination unit is configured to make a prediction according to the second feature map to obtain an evaluation result of the image to be evaluated; The adjacent matrix determination unit is specifically configured to: Divide the first feature map according to the spatial position to obtain nodes corresponding to the first feature map in the Euclidean space; Determine a first content similarity matrix according to the node features corresponding to the nodes, and determine a spatial position correlation matrix according to the position information of the nodes in the first feature map and the size information. The first content similarity matrix is used to reflect the content similarity between nodes in the first feature map, and the spatial position correlation matrix is used to reflect the spatial position relationship between nodes in the first feature map; Determine the sum of the first content similarity matrix and the spatial position correlation matrix as the first adjacent matrix.

13. The device according to claim 12, characterized in that, The adjacent matrix determination unit is specifically configured to: Calculate the similarity between any two of the nodes according to the node features; Determine the first content similarity matrix according to the similarity between any two of the nodes.

14. The device according to claim 12, characterized in that, The adjacent matrix determination unit is specifically configured to: Calculate the spatial distance between any two of the nodes according to the position information of the nodes in the first feature map and the size information; Determine the spatial position correlation matrix according to the spatial distance between any two of the nodes.

15. The device according to claim 12, characterized in that, The evaluation result determination unit is specifically configured to: Make a prediction on the second feature map through the fully connected layer of the first layer layout-aware graph convolutional module to obtain an evaluation result of the image to be evaluated.

16. The device according to claim 12, characterized in that, The evaluation result determination unit is specifically configured to: Make a prediction on the second feature map through the fully connected layer of the first layer layout-aware graph convolutional module to obtain a first prediction result of the image to be evaluated; Aggregate the node features corresponding to the nodes in the second feature map through the second layer layout-aware graph convolutional module of the target graph convolutional network to obtain node features corresponding to the nodes in the latent variable space. The number of nodes in the latent variable space is less than the number of nodes in the second feature map; Determine a third feature map in the Euclidean space according to the node features corresponding to the nodes in the latent variable space; Make a prediction on the third feature map through the fully connected layer of the second layer layout-aware graph convolutional module to obtain a second prediction result of the image to be evaluated; Determine the evaluation result according to the first prediction result and the second prediction result.

17. The device according to claim 16, characterized in that, The evaluation result determination unit is specifically configured to: Obtain a second adjacent matrix in the latent variable space according to the node features corresponding to the nodes in the latent variable space; Perform graph convolution with the second layer layout awareness according to the second adjacent matrix in the latent variable space to obtain response features; Inverse mapping is performed on the response feature to obtain the third feature map.

18. The device according to claim 17, characterized in that, The evaluation result determination unit is specifically configured to: Determine a second content similarity matrix according to the node features corresponding to the nodes in the latent variable space; Use the second content similarity matrix as the second adjacency matrix.

19. The device according to any one of claims 12-18, characterized in that, The evaluation result includes one or more combinations of aesthetic distribution, average aesthetic score, and aesthetic classification.

20. The device according to claim 16, characterized in that If the first prediction result is the first aesthetic distribution and the second prediction result is the second aesthetic distribution, the evaluation result determination unit is specifically configured to: Use the average of the first aesthetic distribution and the second aesthetic distribution as the target aesthetic distribution of the image to be evaluated; Determine the average aesthetic score of the image to be evaluated according to the target aesthetic distribution; Determine the aesthetic classification of the image to be evaluated according to the average aesthetic score.

21. The device according to claim 12, characterized in that, The device further includes a training unit: The training unit is configured to obtain a sample evaluation image and the size information of the sample evaluation image; perform scaling processing on the sample evaluation image, and determine a first sample feature map according to the scaled sample evaluation image; input the first sample feature map and the size information of the sample evaluation image into the target graph convolutional network, and obtain a first sample adjacency matrix in the Euclidean space through the first-layer layout-aware graph convolutional module of the target graph convolutional network; perform first-layer layout-aware graph convolution according to the first sample adjacency matrix to obtain a second sample feature map; perform prediction according to the second sample feature map to obtain a predicted evaluation result; construct a target loss function according to the predicted evaluation result and the labeled evaluation result of the sample evaluation image; and train the target graph convolutional network according to the target loss function.

22. The device according to claim 21, wherein The predicted evaluation result and the labeled evaluation result are aesthetic distributions.

23. An electronic device for image processing, characterized in that, The electronic device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the method according to any one of claims 1-11 according to the instructions in the program code.

24. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program code, and the program code is used to execute the method according to any one of claims 1-11.

25. A computer program product, characterized in that, The computer program product includes instructions, and when the instructions run on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Remote sensing image road extraction method based on graph convolution

    CN112766280A

  • Identifying image aesthetics using region composition graphs

    US20200151546A1