A deep neural network algorithm based on a semantic attribute visual conversion reconstructor
By constructing a deep neural network algorithm with semantic attribute module and target reconstruction, the problem of insufficient semantic and spatial characteristics of image blocks in the visual transformer network model is solved, and more effective image semantic object modeling and spatial representation are achieved.
Patent Information
- Application Number
- CN202211133367.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-17
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-09-17
AI Technical Summary
The existing visual transformer network model fails to effectively utilize the semantic meaning and spatial characteristics of small image blocks in image processing, resulting in insufficient global feature representation.
The semantic attribute module is constructed to extract the semantic attribute feature vectors and key component position coordinates of the image, convert them into position feature vectors through a linear fully connected layer, and build a graph model using a multi-layer semantic attribute converter and semantic target reconstruction device to calculate node similarity coefficients to improve the reconstruction feature representation of the image.
The modeling ability and spatial representation ability of deep neural networks to image semantic objects is improved, and the semantic understanding of global features is enhanced.
Smart Images

Figure CN115482449B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence algorithms, and particularly relates to a deep neural network algorithm based on a semantic attribute visual conversion reconstructor. Background Art
[0002] In recent years, compared with traditional shallow neural networks, deep neural networks use a network structure connected by several neuron layers to fully exploit the information in data and have achieved great success in various fields of artificial intelligence. The convolutional neural network is one of the classic models in deep neural networks. This model can effectively process data with translational invariance such as images by introducing convolutional layers and pooling layers. Since the excellent performance demonstrated by AlexNet (Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. NIP, 2012: 1097–1105.) in the image classification challenge, more and more complex and effective neural networks have been gradually proposed, further promoting the research boom of deep learning technology in the field of computer vision. These widely used general network models include VGG (K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. ICLR, 2015.), GoogleNet (Christian Szegedy, Wei Liu, Yangqing Jia, et al. Going deeper with convolutions. CVPR, 2015), ResNet (Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun. Deep residual learning for image recognition. CVPR, 2016: 770–778), DenseNet (Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. CVPR, 2017: 4700–4708), etc.
[0003] However, convolutional neural networks model local features within an image, and the convolutional filters used can only perceive local regions of the input image, that is, the receptive field of convolutional neural networks is limited. An important measure to improve this shortcoming is to use the self-attention mechanism to increase the global connectivity of the network. This series of network models is the transformer network model originating from the field of machine translation. This type of model is very good at modeling long-range dependencies. Recently, there has been a research boom in introducing the transformer network model into the field of computer vision. Vision transformer networks have achieved competitive performance with convolutional neural networks. For example, Dosovitskiy et al. (Dosovitskiy et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.) proposed dividing an image into a series of 16×16 image patches, and then using several transformer layers to process these patches and finally establish the global features of the image. However, the current vision transformer networks only simply divide the image, and the image patches generated in this way do not have semantic meanings, and the global feature representation established based on the image patches does not consider the spatial characteristics of the image patches. Summary of the Invention
[0004] To solve the above problems, the present invention discloses a deep neural network algorithm based on a semantic attribute vision transformation reconstructor.
[0005] To achieve the above invention purpose, the present invention adopts the following technical solutions:
[0006] A deep neural network algorithm based on a semantic attribute vision transformation reconstructor, comprising the following steps:
[0007] Step 1, construct a semantic attribute module B1 to extract the semantic attribute feature vector of the image and the position coordinates of the key components;
[0008] Step 2, construct a position vector module B2 to convert the position coordinates of the key components into a d-dimensional position feature vector by using a linear fully connected layer;
[0009] Step 3, construct an L-layer semantic attribute transformer B3 to transform the feature obtained by adding the semantic attribute feature vector and the position feature vector to obtain K semantic feature vectors, where each layer of the semantic attribute transformer consists of a semantic attention calculation module and a feed-forward fully connected layer;
[0010] Step 4, construct the semantic object reconstructor B4 of the image. Represent the K semantic feature vectors as the nodes of a graph, calculate the similarity coefficients between pairwise nodes, obtain the reconstruction matrix P of the target image, calculate the C eigenvectors of this matrix and concatenate them as the reconstruction vector of the target image;
[0011] Step 5, calculate the error value between the reconstruction vector of the target image and the true label of the image through the loss function, and use the error value to backtrain and optimize the network parameters to make the algorithm reach the optimal.
[0012] Further, as a preferred technical solution of the present invention, the specific steps of the said Step 1 are as follows:
[0013] S11: Collect several image patches regarding semantic attributes, and use the pixel value vectors within these image patches to train the K-class classifier C(·);
[0014] S12: For any image represented as I(x, y), where (x, y) represents any pixel point within the image, calculate the first-order gradients I x and I y in the horizontal and vertical directions respectively, as well as the second-order gradients and in the horizontal and vertical directions. Establish the gradient correlation matrix, that is:
[0015]
[0016] S13: Calculate the eigenvalues and the trace of the matrix M, where the eigenvalues are represented as λ1, λ2, and the trace of the matrix is represented as ρ. Define the attribute detection candidate region function:
[0017]
[0018] where t is an adjustable parameter; judge the relationship between N and the threshold T. When N is greater than T, then (x, y) is regarded as a semantic attribute candidate region point;
[0019] S14: Convert the pixel values within the image patch with a radius of r centered at (x, y) into vectors and input them into the trained classifier C(·) to output the probability values of K semantic attribute categories, and at the same time obtain the K key components of the image; where the position coordinates of the k-th key component are represented as (x k , y k );
[0020] S15: Convert the pixel values of the image regions of each channel within a radius of r centered at the extracted key component position coordinates into d-dimensional semantic attribute feature vectors, where the semantic attribute feature vector of the k-th key component is represented as
[0021] Furthermore, as a preferred technical solution of the present invention, the specific steps of step 2 are as follows:
[0022] S21: Construct a d-dimensional linear fully connected layer Ψ ω (·); where w is the parameter of the fully connected layer;
[0023] S22: Use the linear fully connected layer to convert the position coordinates of the key components into d-dimensional position feature vectors. Among them, the d-dimensional position feature vector ψ k of the k-th key component after the conversion of the position coordinates is obtained by ψ k = Ψ w (x k , y k ).
[0024] Furthermore, as a preferred technical solution of the present invention, the specific steps of step 3 are as follows:
[0025] S31: For the K key components of the image, superimpose the corresponding position feature vectors and semantic attribute feature vectors to obtain a new semantic attribute feature vector. The superimposed transformation formula of the semantic attribute feature vector of the k-th key component is: z k = z k + l k ;
[0026] S32: Combine the semantic attribute feature vectors processed by S31 into an input matrix and perform layer-by-layer transformation processing on it using the semantic attention calculation module and the feed-forward fully connected layer of the L-layer semantic attribute converter, specifically as follows:
[0027] Among them, the input semantic attribute feature vector matrix of the semantic attention calculation module of the l-th layer is represented as The linear matrices for self-transformation in this layer are respectively represented as Multiply them with Z l respectively to obtain the query matrix Q = [q1, q2,... q K , the key value matrix M = [m1, m2,... m K , and the value matrix V = [v1, v2,... v K , that is:
[0028]
[0029]
[0030]
[0031] Calculate the similarity coefficient between each element of the query matrix Q and the value matrix V using the cosine similarity function to obtain the attention matrix The calculation formula for the ij-th element is as follows:
[0032]
[0033] Based on the above attention matrix, the semantic attribute feature vector matrix is transformed to obtain Its calculation formula is as follows:
[0034]
[0035] Each layer of the semantic attribute converter also includes a feed-forward fully connected layer F W (·), where W is the parameter matrix in the network layer, and the feed-forward fully connected layer is transformed to obtain the discrete semantic attribute feature matrix representation of the l-th layer as:
[0036]
[0037] Furthermore, as a preferred technical solution of the present invention,
[0038] The specific steps of step 4 are as follows:
[0039] S41: The discrete semantic attribute feature matrix after passing through L layers of semantic attribute converters is represented as These K semantic attribute feature vectors are represented as the nodes of a graph, and the calculation formula for the similarity coefficient between any two nodes is:
[0040]
[0041] S42: Establish a reconstruction matrix of the target image Calculate the C eigenvectors of this matrix, where the c-th eigenvector is represented as These C feature vectors are concatenated, and the reconstruction vector of the target feature is represented as
[0042] The deep neural network algorithm based on the semantic attribute visual conversion reconstructor described in the present invention, compared with the prior art by adopting the above technical solutions, has the following technical effects:
[0043] (1) By constructing a semantic feature module, the modeling ability of the current deep neural network for semantic objects can be improved.
[0044] (2) By constructing a reconstructor for establishing the feature representation of the target image, the representation ability of the current deep neural network for the object space in the image can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1It is the network structure diagram of a deep neural network algorithm of a semantic attribute visual conversion reconstructor according to the present invention. Detailed implementation manners
[0046] The following further explains the present invention in detail with reference to the accompanying drawings, so that those skilled in the art can understand the present invention more deeply and be able to implement it. However, the following explanations by reference to examples are only for explaining the present invention and do not limit the present invention.
[0047] As Figure 1 shown, it is the network structure diagram of a deep neural network algorithm of a semantic attribute visual conversion reconstructor according to the present invention. The deep neural network algorithm model consists of a semantic attribute module B1, a position vector module B2, an L-layer semantic attribute converter B3, a semantic target reconstructor B4, etc.
[0048] A deep neural network algorithm based on a semantic attribute visual conversion reconstructor includes the following steps:
[0049] Step 1, construct a semantic attribute module B1 to extract the semantic attribute feature vector of the image and the position coordinates of the key components;
[0050] Step 2, construct a position vector module B2 to convert the position coordinates of the key components into a d-dimensional position feature vector by using a linear fully connected layer;
[0051] Step 3, construct an L-layer semantic attribute converter B3 to convert the features after adding the semantic attribute feature vector and the position feature vector to obtain K semantic feature vectors, where each layer of the semantic attribute converter consists of a semantic attention calculation module and a feed-forward fully connected layer;
[0052] Step 4, construct a semantic target reconstructor B4 of the image, represent the K semantic feature vectors as the nodes of the graph, calculate the similarity coefficient between pairwise nodes to obtain the reconstruction matrix P of the target image, calculate C eigenvectors of this matrix and concatenate them as the reconstruction vector of the target image;
[0053] Step 5, calculate the error value between the reconstruction vector of the target image and the true label of the image through the loss function, and use the error value to backtrain and optimize the network parameters to make the algorithm reach the optimal.
[0054] The specific steps of Step 1 are as follows:
[0055] S11: Collect several image patches about semantic attributes and use the pixel value vectors in these image patches to train a K-class classifier C(·);
[0056] S12: For any image represented as I(x, y), where (x, y) represents any pixel point in the image, calculate the first-order gradients I x and I y in the horizontal and vertical directions respectively, as well as the second-order gradients in the horizontal and vertical directions and Establish a gradient correlation matrix, i.e.:
[0057]
[0058] S13: Calculate the eigenvalues and trace of matrix M, where the eigenvalues are represented as λ1, λ2, and the trace of the matrix is represented as ρ. Define the function for detecting candidate regions of attributes:
[0059]
[0060] where t is an adjustable parameter; judge the relationship between N and the threshold T. When N is greater than T, then (x, y) is regarded as a candidate region point for semantic attributes;
[0061] S14: Convert the pixel values in the image patch with a radius of r centered on (x, y) into a vector and input it into the trained classifier C(·) to output the probability values of K semantic attribute categories, and at the same time obtain K key components of the image; where the position coordinates of the k-th key component are represented as (x k , y k );
[0062] S15: Convert the pixel values of the image region of each channel within a radius of r centered on the extracted position coordinates of the key components into a d-dimensional semantic attribute feature vector, where the semantic attribute feature vector of the k-th key component is represented as
[0063] The specific steps of step 2 are as follows:
[0064] S21: Construct a d-dimensional linear fully connected layer Ψ ω (·); where w is the parameter of the fully connected layer;
[0065] S22: Use the linear fully connected layer to convert the position coordinates of the key components into a d-dimensional position feature vector, where the d-dimensional position feature vector ψ k of the k-th key component after conversion, is calculated by ψ k = Ψ w (x k , y k ).
[0066] The specific steps of step 3 are as follows:
[0067] S31: For K key components of the image, superimpose the corresponding position feature vectors and semantic attribute feature vectors to obtain new semantic attribute feature vectors. The superimposed transformation formula for the semantic attribute feature vector of the k-th key component is: z k = z k + l k ;
[0068] S32: Combine the semantic attribute feature vectors processed by S31 into an input matrix Use the semantic attention calculation module and the feed-forward fully connected layer of the L-layer semantic attribute converter to perform layer-by-layer transformation processing on it, specifically as follows:
[0069] Among them, the input semantic attribute feature vector matrix of the semantic attention calculation module of the l-th layer is expressed as The linear matrices for self-transformation in this layer are respectively expressed as Multiply them with Z l respectively to obtain the query matrix Q = [q1, q2,... q K , the key value matrix M = [m1, m2,... m K , and the value matrix V = [v1, v2,... v K , that is:
[0070]
[0071]
[0072]
[0073] Use the cosine similarity function to calculate the similarity coefficients between each element of the query matrix Q and the value matrix V to obtain the attention matrix The calculation formula for the ij-th element is:
[0074]
[0075] Based on the above attention matrix, transform the semantic attribute feature vector matrix to obtain Its calculation formula is as follows:
[0076]
[0077] Each layer of the semantic attribute converter also includes a feed-forward fully connected layer F W (·), where W is the parameter matrix in the network layer, and the feed-forward fully connected layer performs a transformation on to obtain the discrete semantic attribute feature matrix of the l-th layer, which is expressed as:
[0078]
[0079] The specific steps of Step 4 are as follows:
[0080] S41: The discrete semantic attribute feature matrix after passing through the L-layer semantic attribute converter is denoted as These K semantic attribute feature vectors are represented as the nodes of a graph, and the calculation formula for the similarity coefficient between any two nodes is:
[0081]
[0082] S42: Establish the reconstruction matrix of the target image Calculate the C eigenvectors of this matrix, where the c-th eigenvector is denoted as Concatenate these C eigenvectors, then the reconstruction vector of the target feature is denoted as
[0083] In specific implementation, Step 1: Preprocessing of the classification dataset; for the image classification dataset I, it is divided into a training set, a validation set, and a test set. Randomly extract H images from the training set, and the h-th image and the classification label are denoted as (I h , y h ), and input them into the deep neural network based on the semantic attribute visual conversion reconstructor;
[0084] Step 2: Feature extraction of semantic attributes; the h-th image passes through the semantic attribute module B1, and the attribute detection algorithm detects the positions of K key components in the image. The position of the k-th component is denoted as (x k , y k ). Taking (x k , y k ) as the center, convert the pixel values of each channel in the image region within a small block with a radius of r into a d-dimensional feature vector to obtain the semantic attribute feature vector. The semantic attribute feature vector of the k-th semantic attribute is denoted as
[0085] Step 3: Extraction of position feature vectors; use the fully connected layer L ω (·) to convert (x k , y k ) into the position feature vector l k = L w (x k , y k )
[0086] Step 4: Operations of the semantic attribute converter on the semantic attribute feature vectors; superimpose each semantic attribute feature vector in the image I h with its corresponding position vector, and perform transformations on the K semantic attribute features in the image I h through the L-layer semantic attribute converter to obtain the feature Regarding the K semantic attribute features in the image as the nodes of a graph, calculate the similarity coefficient between any nodes to establish the reconstruction matrix P of the target image, calculate the C eigenvectors of this matrix P and concatenate them, then the final feature representation of the image is obtained as
[0087] Step 5: Optimization of the parameters of the deep neural network based on the semantic attribute visual transformation reconstructor; First, input H images into the deep neural network based on the semantic attribute visual transformation reconstructor in sequence to obtain the feature representation of the images; Then calculate the cross-entropy loss function according to its corresponding manually annotated class labels; Finally, use the gradient descent algorithm to optimize the parameters in the network to obtain the deep neural network model with the optimal performance.
[0088] The present invention first uses the semantic attribute module B1 to obtain the position coordinates of the key components in the image, converts the pixel values in the area where the key component positions are located into semantic attribute feature vectors, and at the same time uses the position vector module B2 to convert the position coordinates into position feature vectors through a linear fully connected layer. Then, use the multi-layer semantic attribute converter B3 to convert the features after adding the semantic attribute feature vectors and the position feature vectors to obtain discrete semantic attribute features. Finally, use the semantic target reconstructor B4 to establish a graph model for the discrete semantic attribute features to obtain the final feature representation of the image. The present invention processes digital images through the deep neural network algorithm of the semantic attribute visual transformation reconstructor composed of modules B1, B2, B3, and B4, which can improve the computer's ability to model semantic objects and spatial representation ability in images.
[0089] The specific implementation methods described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only the specific implementation methods of the present invention and are not intended to limit the scope of the present invention. Any equivalent changes and modifications made by those skilled in the art without departing from the concept and principles of the present invention shall fall within the scope of protection of the present invention.
Claims
1. A deep neural network algorithm based on a semantic attribute visual conversion reconstructor, characterized in that, It includes the following steps: Step 1: Construct a semantic attribute module B1 to extract the semantic attribute feature vector of the image and the position coordinates of the key components; Step 2: Construct a position vector module B2 and use a linear fully connected layer to convert the position coordinates of the key components into a d-dimensional position feature vector; Step 3: Construct an L-layer semantic attribute converter B3 to convert the features after adding the semantic attribute feature vector and the position feature vector to obtain K semantic feature vectors, where each layer of the semantic attribute converter consists of a semantic attention calculation module and a feed-forward fully connected layer; Step 4: Construct a semantic object reconstructor B4 of the image, represent the K semantic feature vectors as the nodes of the graph, calculate the similarity coefficients between pairwise nodes to obtain the reconstruction matrix P of the target image, calculate the C eigenvectors of this matrix and concatenate them as the reconstruction vector of the target image; Step 5: Calculate the error value between the reconstruction vector of the target image and the true label of the image through the loss function, and use the error value to backpropagate and train to optimize the network parameters to make the algorithm reach the optimal; The specific steps of Step 1 are as follows: S11: Collect several image patches about semantic attributes and use the pixel value vectors in these image patches to train a K-class classifier C(·); S12: For any image represented as I(x, y), where (x, y) represents any pixel point within the image, calculate the first-order gradients I x and I y in the horizontal and vertical directions respectively for this point, as well as the second-order gradients in the horizontal and vertical directions and Establish the gradient correlation matrix, that is: S13: Calculate the eigenvalues and the trace of the matrix M, where the eigenvalues are denoted as λ1, λ2, and the trace of the matrix is denoted as ρ, and define the attribute detection candidate region function: where t is an adjustable parameter; judge the relationship between N and the threshold T. When N is greater than T, then (x, y) is regarded as a semantic attribute candidate region point; S14: Take the pixel values within the image patch with radius r centered at (x, y) and convert them into vectors, then input them into the trained classifier C(·) to output the probability values of K semantic attribute categories. At the same time, obtain the K key components of the image. The position coordinates of the k-th key component are represented as (x k , y k ); S15: Taking the coordinates of the extracted key component positions as the center, convert the pixel values of the image regions of each channel within a radius of r into d-dimensional semantic attribute feature vectors, where the semantic attribute feature vector of the k-th key component is expressed as 2. The deep neural network algorithm of a semantic attribute visual conversion reconstructor according to claim 1, characterized in that, The specific steps of Step 2 are as follows: S21: Construct a d-dimensional linear fully-connected layer Ψ w (·); where w is the parameter of the fully-connected layer; S22: Convert the position coordinates of the key components into a d-dimensional position feature vector using a linear fully connected layer, where the d-dimensional position feature vector ψ after the conversion of the position coordinates of the k-th key component k , is obtained from ψ k = Ψ w (x k , y k ).
3. The deep neural network algorithm of a semantic attribute visual conversion reconstructor according to claim 2, wherein The specific steps of Step 3 are as follows: S31: For the K key components of the image, superimpose the corresponding position feature vector and the semantic attribute feature vector to obtain a new semantic attribute feature vector. The superimposition transformation formula for the semantic attribute feature vector of the k-th key component is: z k = z k + l k ; S32: Combine the semantic attribute feature vectors after S31 processing into an input matrix Perform layer-by-layer transformation on it using the semantic attention calculation module and the feed-forward fully connected layer of the L-layer semantic attribute converter, specifically as follows: The input semantic attribute feature vector matrix of the semantic attention calculation module in the l-th layer is represented as The linear matrices for self-transformation in this layer are respectively represented as Multiply with Z l respectively to obtain the query matrix Q = [q1, q2, … q K , the key value matrix M = [m1, m2, … m K , and the value matrix V = [v1, v2, … v K , that is: Calculate the similarity coefficient between each element of the query matrix Q and the value matrix V using the cosine similarity function to obtain the attention matrix The calculation formula for the ij-th element is as follows: Based on the above attention matrix, the semantic attribute feature vector matrix is transformed to obtain Its calculation formula is as follows: Each layer of the semantic attribute converter also includes a feed-forward fully-connected layer F W (·), where W is the parameter matrix in the network layer, and the feed-forward fully-connected layer is transformed to obtain the discrete semantic attribute feature matrix of the l-th layer, which is expressed as:
4. The deep neural network algorithm of a semantic attribute-based visual conversion reconstructor according to claim 3, characterized in that, The specific steps of Step 4 are as follows: S41: The discrete semantic attribute feature matrix after passing through the L-layer semantic attribute converter is expressed as These K semantic attribute feature vectors are represented as the nodes of the graph, and the calculation formula for the similarity coefficient between any two nodes is: S42: Establish the reconstruction matrix of the target image Calculate the C eigenvectors of this matrix, where the c-th eigenvector is denoted as Concatenate these C eigenvectors, then the reconstruction vector of the target feature is denoted as