An Image Recognition Method Based on Adaptive Stacked Capsule Autoencoder
The introduction of adaptive modules AM and OCAE in PCAE decoder through the adaptive stacking capsule autoencoder (ASCAE) to discover objects in PCAE decoder, which solves the shortcomings of existing stacking capsule autoencoder in complex image reconstruction and classification, and achieves more efficient image reconstruction and classification performance.
Patent Information
- Application Number
- CN202411667899.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-11-21
AI Technical Summary
The existing stacked capsule autoencoders are poor in reconstruction and have low unsupervised classification accuracy when processing complex images, especially the CIFAR10 dataset. This is mainly because the fixed template cannot effectively describe complex images and cannot explain how the capsule represents the pose of the object.
Using the adaptive stacked capsule autoencoder (ASCAE), by introducing an adaptive module AM into the PCAE decoder, image features are projected into the image space to form local image tiles, and reconstructed images are generated through geometric transformation. At the same time, objects in the image are discovered using the object capsule autoencoder in OCAE.
Improve the reconstruction quality and unsupervised classification performance of complex images. ASCAE performs better than traditional SCAE on CIFAR10 and FMNIST datasets, allowing better image features and classification.
Smart Images

Figure CN119810613B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to an image recognition method based on an adaptive stacked capsule autoencoder. Background Art
[0002] The capsule network (CapsNet) is an artificial neural network based on vector neurons (capsules). CapsNet can learn the spatial relationships between objects in an image. However, CapsNet has a defect that the spatial geometric information extracted by the capsule network is difficult to interpret. Currently, it can only be experimentally proven that the extracted capsules contain spatial information, but it is impossible to explain how the capsules represent the pose of an object. To solve this problem, Kosiorek et al. proposed the stacked capsule autoencoder (SCAE).
[0003] The SCAE model is divided into two parts: the part capsule autoencoder (PCAE) and the object capsule autoencoder (OCAE). The PCAE encoder is responsible for extracting capsules from the input image. The information in the capsules contains the special features, pose, and probability of existence of a certain part in the image. A template for drawing the image part is defined in the PCAE decoder, and this decoder uses the pose and probability of existence of the part extracted by the encoder to reconstruct the input image, that is, to combine the parts into the original image. Different from other capsule models, these poses are 6D vectors that describe translation, rotation, scaling, and shear. The image reconstruction process is basically the same as the image construction process in computer graphics. In the OCAE, the special features are converted into object capsules. These object capsules represent the objects existing in the image. They also contain the pose of the object, but do not contain appearance information, that is, they are not used to reconstruct the image. Generally speaking, SCAE uses pose, template, and probability of existence to describe an image, which is an intuitive and interpretable method. And SCAE is the first unsupervised capsule model.
[0004] Although SCAE has achieved excellent results in image reconstruction and unsupervised classification tasks, its effectiveness is limited to simple images such as MNIST (a handwritten digit image dataset). On complex datasets, the performance of SCAE is much worse. For example, on the CIFAR10 dataset (a dataset of realistic object images), SCAE cannot reconstruct meaningful images and has a low accuracy in unsupervised classification. This is because MNIST only contains handwritten digits with a pure black background. These images have little noise and are easy to reconstruct and classify. However, CIFAR10 images contain complex appearances and backgrounds. This results in complex structures and high noise in CIFAR10 images. In SCAE, the templates are composed of trainable parameters, but the number is limited. After the model converges, the parameters of the templates also tend to be stable and cannot be changed according to the features of the images. Therefore, Kosiorek et al. call it a fixed template. And the expressiveness of the fixed template is insufficient to model complex images, resulting in SCAE being unable to effectively reconstruct complex images such as CIFAR10 images. Summary of the Invention
[0005] In view of the above-mentioned shortcomings of the prior art, the present invention proposes an image recognition method based on an adaptive stacked capsule autoencoder, which constructs a capsule autoencoder ASCAE based on an adaptive method. In ASCAE, the PCAE encoder part first converts an image into part capsules. The present invention improves the internal data structure of the part capsules to achieve the decoupling of the information used to describe local images and special features.
[0006] The image recognition method based on the adaptive stacked capsule autoencoder proposed by the present invention includes the following steps:
[0007] Step 1: Obtain the image to be analyzed;
[0008] Step 2: Construct an adaptive stacked capsule autoencoder ASCAE;
[0009] Step 3: Input the image to be analyzed into ASCAE;
[0010] Step 4: Use ASCAE to reconstruct the image and use a classifier to achieve image recognition, and output the recognized image.
[0011] Furthermore, the adaptive stacked capsule autoencoder ASCAE in Step 2 includes a part capsule autoencoder PCAE and an object capsule autoencoder OCAE. The capsule autoencoder PCAE includes an encoder and a decoder. An adaptive module AM is set in the decoder of PCAE. The adaptive AM module is a fully connected neural network composed of two layers of fully connected layers.
[0012] Furthermore, a ResCNN deep convolution module and an attention pooling layer connected to its output are provided in the encoder of the capsule autoencoder PCAE. The ResCNN deep convolution module is used to perform convolution calculations on the input original image to obtain a feature map, and output the feature map to the attention pooling layer. The attention pooling layer is used to convert the input feature map into part capsules.
[0013] Furthermore, the part capsules include a special feature component, an image feature component, an existence probability component, and a pose component:
[0014] The existence probability component is used to describe the possibility of the part existing in the image;
[0015] The pose component is used to describe the spatial geometric information of the part;
[0016] The image feature component is used to represent the appearance information of the part;
[0017] The special feature component is used to describe other image features except the appearance of the part, and serves as the input data of the OCAE.
[0018] Furthermore, the adaptive module AM is used to draw the appearance of the part capsules according to the image feature mapping of the part capsules, form tiles of the local image, and these tiles are then geometrically transformed to change the pose, and finally form a reconstructed image.
[0019] Furthermore, the ResCNN convolution module includes 4 sequentially connected residual blocks, and each residual block contains three convolutional layers and a residual connection from the first convolutional layer to the third convolutional layer.
[0020] Furthermore, the specific operation steps of step 4 include:
[0021] Step 4.1: Obtain the input image I to be analyzed, denoted as I(h, w, c), where h is the height, w is the width, and c is the channel;
[0022] Step 4.2: Input the image I into the encoder of the PCAE, and convert I into M part capsules, that is:
[0023]
[0024] where, PE represents the encoder of the PCAE, x represents the pose, d represents the existence probability, z s represents the special feature, z i represents the image feature, and the subscript 1:M represents from 1 to M;
[0025] Step 4.3: The adaptive module AM projects all M image features into M image patches through the following projection formula, each Both contain the appearance information of an encoded component capsule, and the projection formula is:
[0026]
[0027] where T m has a height, width, and number of channels of h t , w t , c; m ∈ (1, M); represents the image features contained in the m-th component capsule;
[0028] Step 4.4: The PCAE decoder learns an alpha channel corresponding to each image patch T m respectively.
[0029] Step 4.5: Based on the pose x m in the component capsule, perform an affine transformation on T m and to obtain the transformed image patches and
[0030]
[0031] where Transform represents the affine transformation;
[0032] Step 4.6: Combine the existence probability d through the following formula to obtain the reconstructed image: m
[0033]
[0034] where represents the obtained reconstructed image, and ⊙ represents the product of the corresponding elements in the image;
[0035] Step 4.7: Use the obtained reconstructed image and the original image I to calculate the loss value between the reconstructed image and the original image using the mean squared error L mse thereby training the PCAE;
[0036] Step 4.8: Input the special features of the component capsule into the OCAE, and concatenate x m , d m and together to obtain a token. Encode the M tokens formed by the M component capsules into K object capsules by the encoder of the OCAE;
[0037] Step 4.9: Use the object decoder of OCAE to decode the corresponding object capsule into 5 components:
[0038] OV k ,a k ,OP k,1:M ,a k,1:M ,λ k,1:M = MLP k (c k )
[0039] where MLP k represents the object decoder implemented by a multi-layer perceptron, OV k represents a 3×3 object-observer relationship matrix, c k represents the k-th object capsule, a k represents the existence probability of c k , OP k,1:M represents a 3×3 object-part relationship matrix, a k,1:M represents the predicted existence probability of the part capsule, λ k,1:M represents the relevant standard deviation;
[0040] Step 4.10: Obtain the predicted pose μ of the k-th object capsule with respect to the m-th part capsule according to the following formula k,m :
[0041] μ k,m = OV k OP k,m
[0042] where OV k represents the affine transformation between the object capsule and the image observer, OP k,m represents the affine transformation between the object capsule and the part capsule;
[0043] Step 4.11: Model the pose as a single Gaussian mixture through the following formula
[0044] p(x m |k,m) = N(x m |μ k,m ,λ k,m )
[0045] where p represents the Gaussian mixture, λ k,m represents the standard deviation of the m-th part capsule;
[0046] Step 4.12: Calculate the pose likelihood L pose , and train OCAE by maximizing L pose . The calculation formula for the pose likelihood L pose is as follows:
[0047]
[0048] Among them, a k represents the existence probability of the k-th object capsule, and all the existence probabilities a k constitute an existence vector with k elements; a k,m represents the prediction result of the existence probability of the k-th object capsule for the m-th part capsule, and a i,j represents the prediction result of the existence probability of the i-th object capsule for the j-th part capsule;
[0049] Step 4.13: By minimizing L mse and maximizing L pose establish the loss function of ASCAE:
[0050] L = L mse -L pose
[0051] Among them, L represents the loss function of ASCAE;
[0052] Step 4.14: Use the existence vector composed of the existence probabilities a k of all object capsules to represent the feature code of the input image I. Based on the constructed loss function L, train the ASCAE model to optimize the feature code, and then the classifier uses the optimized feature code of the image I to implement image classification and output the image recognition result.
[0053] Therefore, the present invention adopts the above-mentioned image recognition method based on an adaptive stacked capsule autoencoder, and has the following beneficial effects:
[0054] First, the present invention proposes a novel capsule autoencoder with an adaptive module. Using the adaptive module, the information in the capsule can be projected into the template, and the template is an image patch. Through this dynamically generated template, complex images of high quality can be efficiently reconstructed. Moreover, there is no need to increase the number of templates.
[0055] Second, the present invention proposes an image feature component in the data structure of the capsule. Through the image feature component, the capsule realizes the decoupling of information between the image appearance and the object, enabling PCAE to pay more attention to reconstructing the image, while OCAE pays more attention to discovering objects.
[0056] Finally, the experimental results prove that ASCAE has better image reconstruction performance than SCAE on CIFAR10 and FMNIST. In addition, the performance of ASCAE in unsupervised and supervised image classification tasks has been improved.
[0057] Next, through the accompanying drawings and embodiments, the technical solution of the present invention will be further described in detail. Description of the Drawings
[0058] Figure 1 This is the architecture of the ASCAE model proposed by the present invention; among them, solid arrows represent information flow, and the gradient flows in the opposite direction, and the dashed arrows represent the information flow where the gradient stops.
[0059] Figure 2 This is the architecture of ResCNN.
[0060] Figure 3 (a)-(c) are the reconstructed images on CIFAR10; among them, Figure 3 (a) is the original image in the dataset, Figure 3 (b) is the ASCAE reconstructed image, Figure 3 (c) is the SCAE reconstructed image.
[0061] Figure 4 (a)-(c) are the reconstructed images on FMNIST; among them, Figure 4 (a) is the original image in the dataset, Figure 4 (b) is the ASCAE reconstructed image, Figure 4 (c) is the SCAE reconstructed image.
[0062] Figure 5 (a)-(c) are the reconstructed images on the transformed CIFAR10; Figure 5 (a) is the original image in the dataset, Figure 5 (b) is the ASCAE reconstructed image, Figure 5 (c) is the SCAE reconstructed image. Specific implementation manners
[0063] In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0064] The present invention proposes a capsule autoencoder based on an adaptive method, which improves the data structure inside the component capsule. The improved component capsule includes four components: image features, existence probability, pose, and special feature components. Among them, the image feature component is a vector describing a local image region. In the PCAE decoder of the ASCAE, an adaptive module AM is introduced, which is implemented by a spatial mapping function. AM projects the image features into the image space to form tiles of the reconstructed image (i.e., local images, which are also the images presented by the components). The existence probability component describes the possibility of the corresponding tile existing. The pose component describes the affine transformation of the tile. All these components except the special feature component are used to reconstruct the image. Based on computer graphics, multiple sub-images in the same image space are combined into a complete image. The special feature component is used to describe some abstract information of the local image and does not participate in image reconstruction. The special features are input into the OCAE of the ASCAE to discover objects.
[0065] 1. Architecture of the ASCAE Model
[0066] As Figure 1 shown, the ASCAE model proposed by the present invention includes two autoencoders: the component capsule autoencoder (PCAE) and the object capsule autoencoder (OCAE). Among them, PCAE is used to reconstruct images, and OCAE is used to discover objects existing in the images.
[0067] (1) PCAE
[0068] Combined Figure 1 , in the encoder of PCAE, first, the image is input into the deep convolutional network ResCNN for convolutional calculation to obtain a feature map. Then, the obtained feature map is input into the attention-based pooling layer, and the feature map is converted into component capsules through the attention-based pooling layer.
[0069] As Figure 2 shown, ResCNN is a deep convolutional network composed of four sequentially connected residual modules. Conv in the figure represents a convolutional layer. For example, Conv(5,1,64) means it has 64 convolutions, a 5×5 kernel, a stride of 1, and ⊕ represents an addition operation. As can be seen from Figure 2 , ResCNN contains four identical residual blocks. Each residual block contains three convolutional layers and a residual connection from the first convolutional layer to the third convolutional layer. In addition, each layer in ResCNN includes a batch normalization process.
[0070] The component capsule contains four components: special features, image features, existence probability, and pose components, where:
[0071] The existence probability component describes the likelihood of the capsule's existence;
[0072] The pose component describes the image patch information represented by the spatial geometric capsule;
[0073] The image feature component is used to represent the appearance information of the patch;
[0074] The special feature component does not participate in the image reconstruction process, is used to describe other attributes of the capsule, and serves as the input data for OCAE.
[0075] Different from SCAE, the independent structures of the image features and special features of the component capsule make PCAE more focused on image reconstruction, while OCAE is more concerned with high-level object description. Through this capsule data structure, the information interference caused by only using special features alone in SCAE can be solved.
[0076] The image features in each component capsule are mapped to the image space through the AM in the decoder of PCAE. In the decoder of PCAE, the projected vectors are reshaped into patches representing local images. The AM module is a fully connected neural network composed of two layers of fully connected layers, that is, a multi-layer perceptron network. Its first layer consists of 56 neurons, and the number of neurons in the second layer is determined by the size of the reconstructed image. If a 32-sized color image is used, then the number of neurons should be 32×32×3 = 3072. Therefore, it can be trained to adaptively generate images. The AM module draws the appearance of the component (i.e., the patches that make up the complete image) based on the input image feature information. These patches change their poses through geometric transformations and finally generate the reconstructed image. And this geometric transformation requires not only the existence probability and pose components but also the alpha channel. Each patch of the local image corresponds to an alpha channel, and the alpha channel represents transparency and is used to compose the image. The alpha channel is independent of the patch and is defined by trainable parameters. Finally, the mean square error (MSE) between the reconstructed image and the input image can be used to train PCAE.
[0077] (2) OCAE
[0078] The OCAE has a similar structure to the OCAE in the SCAE in the literature "Kosiorek, A.R.; Sabour, S.; Teh, Y.W.; Hinton, G.E. Stacked Capsule Autoencoders. In Proceedings of the Proc. Adv. Neural Inf. Process. Syst., 2019, Vol. 32, pp. 15486–15496". The OCAE is used for pose prediction. The predicted poses and the poses in the part capsules are used to train the OCAE.
[0079] 2. Algorithm Implementation of the ASCAE Model
[0080] Define the input image I as an image with the shape of (h, w, c), where h is the height, w is the width, and c is the channel. The image I is converted into M part capsules through the ResCNN convolutional module and the attention-based pooling layer, and this process can be expressed as:
[0081]
[0082] where PE represents the encoder of PCAE, x represents the pose, d represents the existence probability, z s represents the special feature, z i represents the image feature, the subscript 1:M represents from 1 to M, the superscript s represents special, and i represents the image;
[0083] For example, capsule m consists of four parts: pose x m , existence probability d m , special feature and image feature containing the appearance information of capsule m. The AM module projects the image feature in the m-th part capsule into the image patch T m :
[0084]
[0085] where the height, width, and channel of T m are h t , w t , c respectively.
[0086] In addition, the calculation of each part capsule is implemented by the same AM module, so that all part capsules can be projected into the same image space, which is beneficial to image synthesis.
[0087] To more effectively combine the patches into a reconstructed image, the PCAE decoder also learns an alpha channel for each T m
[0088] Based on T m and The process of image reconstruction includes the following two steps:
[0089] First, use x in Equation (3) to m perform an affine transformation on T m to obtain the transformed patch Since the pose of should correspond to that of T m Therefore, perform the same affine transformation on to convert it into
[0090]
[0091] where Transform represents the affine transformation;
[0092] Secondly, according to the following formula, jointly calculate the patch alpha channel existence probability d m to obtain the reconstructed image:
[0093]
[0094] where represents the obtained reconstructed image, d m represents the existence probability, and ⊙ is the product of the corresponding elements in the image;
[0095] Finally, use the mean square error L to calculate the loss value between the obtained reconstructed image and the original image I, so as to train the PCAE and continuously improve the image synthesis ability of the adaptive stacked capsule autoencoder. mse The special features in the part capsules do not participate in the image reconstruction work, and these features are used to discover the possible objects in the image. The OCAE uses the special features as the main input information, combines the existence probability and pose of the part capsules, and outputs the object capsules that describe the possible objects in the image.
[0096]
[0097] Since the OCAE needs the part capsules output by the PCAE to discover the objects existing in the image I. In the OCAE, first, the x m , d m and Connect together to form a token, and the M tokens of the M component capsules are encoded into K object capsules by the encoder Set Transformer of OCAE. Different from the OCAE in the prior art, in the OCAE of the present invention, T m does not need to be input into the Set Transformer to identify the identity of the component capsules, because the identity information of the component capsule m is already included in .
[0098] In OCAE, its target decoder is implemented by a multi-layer perceptron (MLP), and each object capsule corresponds to an MLP k , and through this decoder, the k-th object capsule c k is decoded into 5 components:
[0099] OV k , a k , OP k,1:M , a k,1:M , λ k,1:M = MLP k (c k ) (5)
[0100] wherein, MLP k represents the target decoder implemented by a multi-layer perceptron, OV k represents a 3×3 object-observer relationship matrix, a k represents the existence probability of c k , OP k,1:M represents a 3×3 object-component relationship matrix, a k,1:M represents the predicted existence probability of the component capsule, and λ k,1:M represents the relevant standard deviation.
[0101] Next, the predicted pose μ k,m of the current object capsule k with respect to the component capsule m is calculated by Equation (6):
[0102] μ k,m = OV k OP k,m (6)
[0103] wherein, OV k represents the affine transformation between the object capsule and the image observer, and OP k,m represents the affine transformation between the object capsule and the component capsule;
[0104] In OCAE, the pose x m is modeled as a single Gaussian mixture p, and each pose x m is interpreted as an independent mixture predicted from the object capsule c k :
[0105] p(x m |k,m) = N(x m |μ k,m ,λ k,m ) (7)
[0106] where μ k,m and λ k,m are the center and standard deviation of the mixture components.
[0107] The OCAE is trained by maximizing the pose likelihood L pose , and the pose likelihood L pose = p(x 1:M ,d 1:M ):
[0108]
[0109] where a k represents the existence probability of the k-th object capsule, and all the existence probabilities a k constitute an existence vector with k elements; a k,m represents the prediction result of the existence probability of the k-th object capsule for the m-th part capsule, and a i,j represents the prediction result of the existence probability of the i-th object capsule for the j-th part capsule.
[0110] The ASCAE can construct the loss function L of the ASCAE by minimizing L mse and maximizing L pose , and the formula of the loss function is:
[0111] L = L mse - L pose (9).
[0112] The existence vector can describe which objects may exist in the input image. Therefore, the existence vector is used to represent the feature code of the input image, and the ASCAE model is trained through the loss function L to optimize the feature code. Finally, the classifier uses the optimized feature code of the input image to implement image classification and outputs the image recognition result.
[0113] Specifically, the present invention selects a LIN-MATCH classifier or a LIN-PRED classifier. LIN-MATCH is an unsupervised classifier that first clusters the feature codes and then performs bipartite graph matching to obtain the classification. The LIN-PRED classifier is a classifier of a supervised method, which is internally composed of a single-layer fully connected neural network and needs to be trained with labeled images before use to enable it to have the classification ability. The fully connected neural network can directly convert the input feature code into the category of the image, thereby realizing the recognition of the image.
[0114] In summary, the adaptive stacked capsule auto - encoder proposed by the present invention first decomposes an image into multiple components, then recombines each component, analyzes various possible constituent objects in the image during this process, and uses the existence probabilities of these objects to describe the features of the image, finally realizing image recognition.
[0115] Embodiment
[0116] To verify the method proposed by the present invention, the following experiments are carried out.
[0117] 1. Model selection
[0118] In PCAE, the width h m and height w t of T are both 14. On CIFAR10, the number of channels c is 3. On FMNIST, c is 1. The size of the alpha channel t is the same as that of T . The number M of part capsules is 32. The number K of object capsules is 64. The number of elements of the special feature m and the image feature are both 16. The number of elements in the pose x is 6. The AM module is implemented by a network of two fully - connected layers. The first layer has 256 units, and the second layer has h m ×w t ×c units. The other hyper - parameter configurations of the ASCAE model are the same as those of SCAE. t
[0119] The ASCAE model is trained using the RMSProp optimizer. The initial learning rate is set to 0.00003, and the learning rate strategy is to reduce the learning rate by 4% every 10000 updates of the model parameters. The images in all the above datasets are not augmented with data. Since it is difficult for SCAE to reconstruct the original images in CIFAR10, Kosiorek et al. used the images after Sobel filtering as the reconstruction target to emphasize the importance of shape. However, the ASCAE proposed by the present invention can reconstruct the original images. Therefore, the ASCAE model does not require image pre - processing, such as the Sobel filter, on both the FMNIST dataset and the CIFAR10 dataset. The number of images used for one training is 128, and the number of training iterations is 800. The main loss function is shown in formula (9), and the other auxiliary loss functions are the same as those of SCAE, such as prior sparsity and posterior sparsity.
[0120] 2. Dataset setting
[0121] This embodiment mainly studies the reconstruction and recognition performance of ASCAE on complex images. Therefore, two commonly used and representative datasets are used to evaluate ASCAE, namely the CIFAR10 and FMNIST datasets. Among them, CIFAR10 is a collection of real-world images. It has 10 categories, with 6000 samples in each category. It includes a training set and a test set. The training set has 50000 samples, and the test set has 10000 samples. Each sample is a 3-channel color image with a size of 32×32. The images in CIFAR10 contain objects with complex shapes and complex background noises. Therefore, it is difficult to classify the samples in CIFAR10 using an unsupervised model.
[0122] FMNIST is a collection of grayscale images, and the objects in it are clothes, shoes, etc. The objects in FMNIST are derived from real objects, so its appearance structure is more complex than that of MNIST. FMNIST is also a dataset commonly used to evaluate capsule models. It has 10 categories, with 7000 samples in each category. Its training set has 60000 samples, and the test set has 10000 samples. Each sample is a single-channel grayscale image with a size of 28×28. In the experiment of this embodiment, in order to be consistent with CIFAR10, the images in FMNIST are rescaled to 32×32.
[0123] 3. Image Reconstruction
[0124] To study the performance of ASCAE, the reconstructed images of the ASCAE model are compared with those of the SCAE. The comparison results are as Figure 3 and Figure 4 shown. Figure 3 In (a) in [figure], it shows the original image in the CIFAR10 dataset; Figure 3 In (b) in [figure], it shows the reconstructed image of ASCAE for the CIFAR10 dataset; Figure 3 In (c) in [figure], it shows the reconstructed image of SCAE for the CIFAR10 dataset. Figure 4 In (a) in [figure], it shows the original image in the FMNIST dataset; Figure 4 In (b) in [figure], it shows the reconstructed image of ASCAE for the FMNIST dataset; Figure 4 In (c) in [figure], it shows the reconstructed image of SCAE for the FMNIST dataset. From Figure 3 in (b) and Figure 3 in (c) comparison, it can be obtained that in CIFAR10, ASCAE is significantly better than SCAE. The reconstructed image of SACE is only composed of some color blocks, and in some images, even the outline of the object cannot be distinguished. These results intuitively show the defect of the fixed template of SCAE.
[0125] In SCAE, fixed templates are trained to describe the local parts of images. However, after training is completed, the parameters in the templates are fixed. In other words, the patches that the templates can represent are fixed. Moreover, in SACE, the number of fixed templates is limited. The images in CIFAR10 are complex and cannot be effectively described by a limited number of fixed templates. In addition, fixed templates are independently defined variables and have no tight connection with the input images. These reasons lead to Figure 3 the phenomenon shown in (c). The patches of the ASCAE model are generated through a projection process, which is a mapping from image features to the image space. The trained AM module has sufficient ability to generate complex local images, overcoming the above-mentioned defects of fixed templates. In addition, comparing Figure 3 (a) in Figure 3 with (b) in
[0126] Figure 4 the reconstruction results of ASCAE and SCAE on FMNIST. It can be seen from Figure 4 (c) that although SCAE can reconstruct images, color blocks are still visible. However, the reconstructed images of ASCAE are visually closer to the original images. Although the images in FMNIST are relatively simple and can be represented by fixed templates, ASCAE is still superior to SCAE.
[0127] 4. Image Classification
[0128] PCAE provides special feature information for OCAE, which is the key to realizing image classification. Therefore, the classification performance of ASCAE is evaluated in this embodiment.
[0129] There are two methods to evaluate the classification performance of ASCAE: LIN-MATCH and LIN-PRED. LIN-MATCH is an unsupervised method. LIN-PRED is a supervised image classification method. In LIN-PRED, a linear classifier is trained using labeled samples to achieve classification. Its input is the existence vector output by OCAE, just like LIN-MATCH. However, the classifier does not propagate gradients to ASCAE. This ensures that the parameters of ASCAE are still trained by unsupervised image reconstruction and are not affected by supervised classification.
[0130] Table 1 shows the results of classification using ASCAE. The best results in Table 1 are shown in bold, and "-" indicates that the results of this model's paper are not reported.
[0131] Table 1 Classification Results of ASCAE
[0132]
[0133]
[0134] As can be seen from Table 1, on CIFAR10 and FMNIST, ASCAE, SCAE, and other unsupervised models were compared. The results show that ASCAE has better performance than SCAE. In LIN-MATCH, ASCAE achieved performance improvement on both datasets. For example, on CIFAR10, the accuracy of ASCAE is 4.59% higher than that of SCAE. In LIN-PRED, the results of ASCAE are still better than those of SCAE. This performance improvement proves that ASCAE can extract better existence vectors than SCAE for image description. This is because the performance of PCAE in ASCAE proposed in the present invention has been significantly improved, so that more efficient image features can be learned. In addition, due to the independent structure of the component capsules, more reasonable special feature information can be provided for OCAE, enabling the improved OCAE to output better existence vectors.
[0135] 5. Robustness to Affine Transformations
[0136] Robustness to affine transformations is a major feature of the capsule autoencoder. In SCAE, the transformed MNIST images were used to evaluate affine robustness. However, MNIST only contains easily recognizable digit images. For complex images, we used the transformed CIFAR10 images to verify the affine robustness of ASCAE. The transformed CIFAR10 data was obtained by affine transformation of the image samples in the original CIFAR10 test set. These affine transformation processes include: First, the image is randomly rotated by an angle between -20 and +20 degrees; Second, a multiple between 0.8 and 1.2 is randomly selected to scale the image; Third, the image is randomly moved 3 pixels in both the horizontal and vertical directions, so that while translating the object, the main part of the image will not be moved out of the boundary. In this embodiment, ASCAE was trained on the CIFAR10 training set without any data augmentation. Then, the trained ASCAE was evaluated on the transformed CIFAR10 dataset. In this experiment, the transformed image samples used for testing were not involved in the training process at all.
[0137] Figure 5 (a) shows the original image in the CIFAR10 dataset after affine transformation, Figure 5 (b) shows the reconstructed image of ASCAE, Figure 5 (c) shows the reconstructed image of SCAE. From Figure 5 (a)- Figure 5 (c), it can be seen that ASCAE can still reconstruct the objects in the transformed image. However, the reconstructed image of SCAE is still only composed of color blocks and is worse in appearance thanFigure 3 Medium (c) difference. Figure 5 The results prove that the reconstruction ability of ASCAE has stronger affine robustness than SCAE.
[0138] 6. Ablation experiments
[0139] Table 2 Ablation experiment results of ASCAE on the CIFAR10 dataset
[0140]
[0141] To further study ASCAE, ablation experiments were conducted on CIFAR10. The experimental results are shown in Table 2. Here, the image recognition accuracies obtained by two classifiers, LIN-MATCH and LIN-PRED, were used to conduct ablation experiments on ASCAE. The best results in Table 2 are shown in bold. The "without ResCNN" experiment means using a simple convolutional network instead of ResCNN in ASCAE. In LIN-MATCH, the classification accuracy decreased by 0.73%. This result indicates that the feature maps extracted by ResCNN have better effects. The "without alpha channel" experiment means removing the alpha channel from the image reconstruction process. The results show that the performance of ASCAE will decline without the alpha channel. Therefore, the alpha channel is also very important for image reconstruction. The "alpha channel from mapping" experiment means that the alpha channel is generated by mapping the image features of AM. The experimental results show that this mapped alpha channel is not reasonable. The alpha channel should be defined by trainable parameters independent of the tiles. The "OCAE input with tiles" means using the tiles generated by AM as part of the OCAE input. However, the experiments show that the tiles do not need to be input into OCAE. What OCAE needs is abstract information rather than visual information.
[0142] Experiments on CIFAR10 and FMNIST show that ASCAE can reconstruct images better than SCAE. The adaptive method of ASCAE is more effective than the fixed template. Improving PCAE with the adaptive method can also improve the unsupervised classification performance of ASCAE.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. An image recognition method based on an adaptive stacked capsule autoencoder, characterized in that It includes the following steps: Step 1: Obtain the image to be analyzed; Step 2: Construct an Adaptive Stacked Capsule Autoencoder (ASCAE); Step 3: Input the image to be analyzed into the ASCAE; Step 4: Use the ASCAE to reconstruct the image and use a classifier to achieve image recognition, and output the recognized image; The Adaptive Stacked Capsule Autoencoder (ASCAE) in Step 2 includes a Part Capsule Autoencoder (PCAE) and an Object Capsule Autoencoder (OCAE). The Part Capsule Autoencoder (PCAE) includes an encoder and a decoder. An Adaptive Module (AM) is set in the decoder of the PCAE. The Adaptive AM module is a fully connected neural network composed of two layers of fully connected layers; A ResCNN deep convolutional module and an attention pooling layer connected to its output are set in the encoder of the Part Capsule Autoencoder (PCAE). The ResCNN deep convolutional module is used to perform convolutional calculations on the input original image to obtain a feature map and output the feature map to the attention pooling layer. The attention pooling layer is used to convert the input feature map into part capsules; The part capsules include a special feature component, an image feature component, an existence probability component, and a pose component: The existence probability component is used to describe the possibility of the part existing in the image; The pose component is used to describe the spatial geometric information of the part; The image feature component is used to represent the appearance information of the part; The special feature component is used to describe other image features except the appearance of the part and serves as the input data of the OCAE.
2. The image recognition method based on an adaptive stacked capsule autoencoder according to claim 1, wherein The Adaptive Module (AM) is used to draw the appearance of the part capsule according to the image feature mapping of the part capsule to form patches of the local image. These patches are then transformed geometrically to change the pose and finally form a reconstructed image.
3. The image recognition method based on an adaptive stacked capsule autoencoder according to claim 1, wherein, The ResCNN deep convolutional module includes 4 sequentially connected residual blocks, and each residual block contains three convolutional layers and a residual connection from the first convolutional layer to the third convolutional layer.
4. The image recognition method based on an adaptive stacked capsule autoencoder according to claim 1, characterized in that The specific operation steps of Step 4 include: Step 4.1: Obtain the input image I to be analyzed and denote it as I(h, w, c), where h is the height, w is the width, and c is the channel; Step 4.2: Input the image I into the encoder of the PCAE and convert I into M part capsules, that is: Among them, PE represents the encoder of PCAE, represents the pose, represents the existence probability, represents the special feature, represents the image feature, subscript 1: M represents from 1 to M; Step 4.3: The adaptive module AM projects all M image features into M image patches through the following projection formula, each of which contains the appearance information of an encoded part capsule, and the projection formula is: Among them, The height, width, and channels are h t , w t , c respectively; ; represents the image features contained in the m th component capsule; Step 4.4: The PCAE decoder learns a corresponding alpha channel for each image block ; ; Step 4.5: Based on the attitude in the component capsule , transform and by affine transformation to obtain the transformed The replaced image block and : Among them, represents an affine transformation; Step 4.6: Through the following formula, , , the existence probability are jointly calculated to obtain the reconstructed image: Among them, represents the obtained reconstructed image, represents the product of the corresponding elements in the image; Step 4.7: Using the obtained reconstructed image and the original image Use the mean square error to calculate the loss value between the reconstructed image and the original image, so as to train the PCAE; Step 4.8: Input the special features of the part capsules into the OCAE, and combine , and together to obtain a token. Encode the M tokens formed by the part capsules M by the encoder of the OCAE into K object capsules; Step 4.9: Use the object decoder of the OCAE to decode the corresponding object capsule into 5 components: Among them, MLP k represents an object decoder implemented by a multi-layer perceptron, represents an object-observer relationship matrix of size 3×3, represents the k th object capsule, represents the probability of existence, represents a 3×3 object-part relationship matrix, represents the predicted probability of existence of the part capsule, represents the relevant standard deviation; Step 4.10: Obtain the predicted attitude of the k th object capsule with respect to the m th component capsule according to the following formula : Among them, represents a 3×3 object-observer relationship matrix, represents a 3×3 object-component relationship matrix; Step 4.11: Model the pose as a single Gaussian mixture by the following formula: where p represents a Gaussian mixture, denotes the m standard deviation of the Step 4.12: Calculate the pose likelihood by the following formula , and train the OCAE by maximizing . The calculation formula of the pose likelihood is as follows: Among them, represents the existence probability of the k th object capsule. All the existence probabilities constitute an existence vector with k elements; represents the prediction result of the existence probability of the k th object capsule on the m th part capsule, represents the prediction result of the existence probability of the i th object capsule on the j th part capsule; Step 4.13: By minimizing L mse and maximizing L pose establish the loss function of ASCAE: Among them, L represents the loss function of ASCAE; Step 4.14: Use the existence vector composed of the existence probabilities of all object capsules to represent the feature code of the input image I. Based on the constructed loss function L train the ASCAE model to optimize the feature code, and then the classifier uses the optimized feature code of the image I to implement image classification and output the image recognition result.
Citation Information
Patent Citations
Centrifugal pump fault diagnosis method based on EAS and stacked capsule auto-encoder
CN113530850A
Object discovery in images through categorizing object parts
US20220230425A1