A Brain-Inspired Global-Local Dual-Channel Image Classification Method and System
The brain-inspired global-local dual-channel image classification method integrates CNNs for local details and Transformers for global context, enhancing CNN performance by fusing features through a gating mechanism to improve accuracy and robustness.
Patent Information
- Application Number
- CN202211149221.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-21
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-21
AI Technical Summary
Existing deep convolutional neural networks have problems with texture paranoia and semantic confusion in image classification, and lack the ability to capture global structures or contexts in local semantics, resulting in misclassification results.
The brain-inspired global-local dual-channel image classification method is adopted. By using the CNN model as the local channel basic module for extracting local details, and combining the Transformer model as the global channel basic module for extracting global topology correlation information, the modulator is used to fuse the dual-channel output features to form a global-local dual-channel image classification model.
It improves the classification accuracy and robustness of the model, improves the generalization ability of the model, and can better utilize the local fine-grained information and global spatial topological information of the image.
Smart Images

Figure CN115439696B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a brain-inspired global-local dual-channel image classification method and system. Background Art
[0002] Deep convolutional neural networks have been widely applied in the fields of computer vision, natural language processing, speech recognition, etc. They are one of the best learning algorithms for understanding image content and have shown excellent performance in related tasks such as image classification, semantic segmentation, object detection, and retrieval.
[0003] The powerful learning ability of deep convolutional neural networks lies in the use of multiple feature extraction stages, and the significant improvement in their representation ability is mainly achieved through architectural innovation. With the popularity of the idea of using layer blocks as structural units, architectural innovation in networks mainly focuses on the reorganization of processing units and the design of new blocks. After Google proposed the innovative concept of split, transform, and merge, the concept of intra-layer branches was first introduced, allowing features to be extracted at different spatial scales and making multi-branch topologies an important idea for network architecture innovation. ResNet adds shortcut connections to form a dual-branch residual module, and the shortcut branch helps to propagate information during the training process, greatly reducing the training difficulty. ResNeXt uses group convolution to aggregate multiple residual transformations in each building unit. The input feature map is transformed into several groups in the channel dimension and processed by multiple branches respectively. The Inception network constructs its architecture by stacking Inception modules, and each module aggregates branches of multiple convolutional layers, allowing for flexible combination of various operations to represent features of different patterns. Architectural innovation and hardware support allow for the construction of deeper and deeper networks, and the model can be extended to larger and more complex problems.
[0004] However, when the network reaches a certain depth, the recognition performance will saturate with a significant increase in computational power, and the improvement of the performance of existing deep convolutional neural networks has encountered a bottleneck. Moreover, CNN still faces some unsolved problems. (1) Recent research has found that CNNs trained on large-scale image datasets (such as ImageNet) are more biased towards extracting the texture features of images. More specifically, when CNN is applied to image recognition, local texture features may dominate the global object shape, which may lead to misclassification of objects with complex textures. This texture bias problem may be due to the fact that CNN uses a set of small convolutional filters to densely extract local features, and these filters tend to adapt to local texture patterns rather than global object shape information. (2) Another common problem of CNN is the semantic confusion problem. Different feature map channels learned by the CNN model may focus on different local parts (or local semantics) in the input image, and thus obtain the final image recognition result through the competition of local semantics captured by different feature map channels. Due to the lack of the ability to capture the global structure or context in local semantics, some local semantics may dominate other semantics and may lead to misclassification results. These problems have an impact on the improvement of the robustness and generalization ability of the model. Summary of the Invention
[0005] The purpose of the present invention is to provide a brain-inspired global-local dual-channel image classification method and system to improve the problem in the prior art that the existing deep convolutional neural network classification model lacks the ability to capture the global structure or context in local semantics, and some local semantics may dominate other semantics and may lead to misclassification results.
[0006] In a first aspect, an embodiment of the present application provides a brain-inspired global-local dual-channel image classification method, including the following steps:
[0007] Obtain and divide the image set to be classified into a training set and a test set;
[0008] Select a CNN model, use its building unit as the local channel basic module of the dual-channel model, and extract local detailed information from the input features to obtain a feature representation with local information;
[0009] Select a Transformer model, use its encoding layer component as the global channel basic module of the dual-channel model, and extract global topological correlation information from the input features to obtain a feature representation with global information;
[0010] Use the local channel basic module and the global channel basic module of the dual-channel model as parallel dual channels, and connect them to a modulator respectively to form a dual-channel building unit. Through the modulator, fuse the output features of the dual channels to obtain the output features of the dual-channel module;
[0011] Stack multiple of the said dual-channel building units according to the hierarchical architecture of the CNN model to obtain a global-local dual-channel image classification model;
[0012] Train the global-local dual-channel image classification model using the said training set to obtain a trained global-local dual-channel image classification model;
[0013] Use the trained global-local dual-channel image classification model to classify the said test set to obtain an image classification result.
[0014] Based on the first aspect, in some embodiments of the present invention, extracting global topological correlation information from the input features to obtain a feature representation with global information includes the following steps:
[0015] Map the said input features into N one-dimensional Tokens;
[0016] Generate K groups of global topological representations from the said N one-dimensional Tokens;
[0017] Convert the N output Tokens of the said K groups of global topological representations into a multi-dimensional output feature map to obtain a feature representation with global information.
[0018] Based on the first aspect, in some embodiments of the present invention, fusing the output features of the dual channels through the said modulator to obtain the output features of the dual-channel module includes the following steps:
[0019] The said modulator fuses the feature representation with global information and the feature representation with local information in a learnable manner through a gating mechanism to generate the output features of the dual-channel module.
[0020] Based on the first aspect, in some embodiments of the present invention, fusing the feature representation with global information and the feature representation with local information to generate the output features of the dual-channel module includes the following steps:
[0021] The said modulator modulates the output features of the dual channels through a gating mechanism to obtain the output features of the dual-channel module. The expression of the output features of the dual-channel module is: where σ(·) is the sigmoid activation function, is the feature representation with global information of the l-th layer, represents the feature representation with local information of the l-th layer, Y l is the output feature of the dual-channel module of the l-th layer.
[0022] Based on the first aspect, in some embodiments of the present invention, the modulator fuses the feature representation with global information and the feature representation with local information in a learnable manner to generate the output feature of the dual-channel module, including the following steps:
[0023] The modulator adopts an additive fusion method and dynamically learns the proportional relationship between the features of the two channels by using the learnable parameter λ to obtain the output feature of the dual-channel module. The expression for obtaining the output feature of the dual-channel module is: Where is the ReLU activation function, is the feature representation with global information of the l-th layer, represents the feature representation with local information of the l-th layer, Y l is the output feature of the dual-channel module of the l-th layer, and λ is the proportional relationship dynamically learned between the features of the two channels.
[0024] Based on the first aspect, in some embodiments of the present invention, the local channel basic module in the dual-channel construction unit is composed of N L stacked CNN model construction units, and the global channel basic module in the dual-channel construction unit is composed of N G stacked encoding layer components of the Transformer model.
[0025] Second aspect, an embodiment of the present application provides a brain-inspired global-local dual-channel image classification system, including:
[0026] An image set division module to be classified, configured to obtain and divide the image set to be classified into a training set and a test set;
[0027] A local path module, configured to select a CNN model, use its construction unit as the local channel basic module of the dual-channel model, and extract local detailed information from the input features to obtain a feature representation with local information;
[0028] A global path module, configured to select a Transformer model, use its encoding layer component as the global channel basic module of the dual-channel model, and extract global topological correlation information from the input features to obtain a feature representation with global information;
[0029] A modulation module, configured to use the local channel basic module and the global channel basic module of the dual-channel model as parallel dual channels, and connect them to the modulator respectively to form a dual-channel construction unit, and fuse the output features of the two channels through the modulator to obtain the output feature of the dual-channel module;
[0030] A model construction module, configured to stack a plurality of the dual-channel construction units according to the hierarchical architecture of a CNN model, so as to obtain a global-local dual-channel image classification model;
[0031] A model training module, configured to train the global-local dual-channel image classification model by using the training set, so as to obtain a trained global-local dual-channel image classification model;
[0032] An image classification module, configured to classify the test set by using the trained global-local dual-channel image classification model, so as to obtain an image classification result.
[0033] In a third aspect, an embodiment of the present application provides an electronic device, which includes a memory for storing one or more programs; and a processor. When the one or more programs are executed by the processor, the method described in any one of the first aspects above is implemented.
[0034] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the first aspects above is implemented.
[0035] The embodiment of the present invention has at least the following advantages or beneficial effects:
[0036] An embodiment of the present invention provides a brain-inspired global-local dual-channel image classification method and system. The method includes dividing an image set to be classified into a training set and a test set; then selecting a deep convolutional neural network (CNN) model with a layer structure, using its building units as the basic modules of the local channel of the dual-channel model, and extracting local detailed information from the input features to obtain a feature representation with local information; selecting a Transformer model, using its encoding layer components as the basic modules of the global channel of the dual-channel model, and extracting global topological correlation information from the input features to obtain a feature representation with global information; taking the CNN module of the local channel and the Transformer encoding layer module of the global path as parallel dual-channels, and then fusing the output features of the two channels through a top-down modulator to form a building unit of the dual-channel image classification model, and finally obtaining the output features of a layer of dual-channel modules; stacking the dual-channel building units and constructing the model according to the hierarchical architecture of the CNN model to obtain a global-local dual-channel image classification model; then training using the training set according to the batch-based stochastic gradient descent method and some data augmentation methods. After the model is trained, the test set of the images to be classified is tested to complete the classification prediction. By adding a new global path representation method and combining a top-down feature modulation mechanism, a global-local dual-path classification model is constructed, providing a computational model based on the dual-channel visual recognition mechanism of the human brain, called the Global-Local network (abbreviation: GLNet), which can make full use of the local fine-grained information and global spatial topological information of the image, and proposes a general principle for converting a typical deep convolutional neural network (CNN) structure with a layer structure into a global-local dual-path classification model structure, and applying this conversion to some representative baseline CNN models, thereby improving the classification accuracy of the model and enhancing the robustness and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 It is a flowchart of a brain-inspired global-local dual-channel image classification method provided by an embodiment of the present invention;
[0039] Figure 2 It is an architecture diagram of a brain-inspired global-local dual-path model provided by an embodiment of the present invention;
[0040] Figure 3A novel feature-to-feature (F2F) Transformer module provided by an embodiment of the present invention;
[0041] Figure 4 A block diagram of a brain-inspired global-local dual-channel image classification system structure provided by an embodiment of the present invention;
[0042] Figure 5 A block diagram of an electronic device provided by an embodiment of the present invention.
[0043] Icons: 110 - Image set partitioning module to be classified; 120 - Local path module; 130 - Global path module; 140 - Modulation module; 150 - Model construction module; 160 - Model training module; 170 - Image classification module; 101 - Memory; 102 - Processor; 103 - Communication interface. Detailed implementation manners
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. Components of the embodiments of the present application described and illustrated herein usually can be arranged and designed in various different configurations.
[0045] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0046] It should be noted that the term "including", "comprising", or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitations, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.
[0047] In the description of the present application, it should also be noted that unless otherwise clearly specified and limited, the term "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0048] Embodiment
[0049] The following will, with reference to the accompanying drawings, elaborate on some embodiments of the present application. Without conflict, the following various embodiments and the various features in the embodiments can be combined with each other.
[0050] Since VGG and Google have verified the recognition ability of deep convolutional neural networks (CNNs), a series of research works have focused on improving the reorganization of CNN processing units and the design of new blocks. In the innovation of CNN network architectures, multi-path topologies are often adopted.
[0051] Deep CNNs usually perform well in complex tasks. However, due to the increase in depth, performance degradation, gradient vanishing, or explosion problems may occur. In the face of this challenge, the concepts of multi-path and cross-layer connection are proposed. ResNet adds shortcut connections to form a two-branch residual module, and the shortcut branch helps to propagate information during the training process. DenseNet skips some intermediate layers in a multi-path manner and connects two non-consecutive layers, thereby realizing cross-layer information flow and enabling the gradient to approach lower layers. For this reason, different shortcut connection methods emerge in an endless stream, such as zero-padded connections, projection-based connections, pruning connections, skip connections, 1×1 convolution connections, etc. However, although the problem of gradient vanishing has been solved and the design idea of multi-branches has been proven feasible, when the network reaches a certain depth, the recognition performance will saturate with a significant increase in computing power, and existing CNN models have encountered bottlenecks in improving performance. Therefore, the present invention focuses on discovering the main problems of existing CNN models, specifically solving the defects, and improving the model performance.
[0052] Due to the limitations of the working mechanism of convolutional networks, CNN models often only focus on local information and tend to process the fine-grained local features of images. The classification results are actually obtained through the semantic competition of each local feature, which has an impact on the improvement of the robustness and generalization ability of the model. Recently, the Transformer model has received extensive attention in the field of computer vision. Different from the operation mode of convolutional networks, this deep neural network based on the self-attention mechanism can avoid the local feature bias problem caused by the receptive field constraint of convolutional networks by modeling the long-range dependence relationship between local parts of images. However, compared with the natural inductive bias (such as translational invariance, etc.) and bionic characteristics of CNNs based on prior knowledge in the image coding domain, the Transformer model does not have these advantages in image problems. Therefore, it is more difficult to learn and requires a larger dataset or stronger data augmentation and training techniques to achieve better training results.
[0053] The latest research on the human visual cognitive mechanism shows that human object recognition is divided into the ventral pathway and the subcortical pathway. The ventral pathway processes fine local information (texture, edges) from bottom to top to form a detailed representation of visual stimuli, while the subcortical pathway specifically processes the rough and global (such as shape, topology) information of the input visual stimuli and analogously forms a global representation of the object category. Finally, the rough and overall representation from the subcortical pathway guides and regulates the fine local representation from the ventral pathway in a "top-down" manner, thereby forming a more accurate representation of the object category. This method not only retains the shape and texture related to the category but also suppresses the interference of irrelevant details and noise. Inspired by the research of brain cognitive science, the present invention proposes a global-local dual-channel image classification method based on the dual-channel visual cognitive process of the human brain, using the CNN model building unit and the Transformer model encoding component as the basic modules of the local processing pathway and the global processing pathway respectively.
[0054] Please refer to Figure 1 , Figure 1 which is a flowchart of a brain-inspired global-local dual-channel image classification method provided by an embodiment of the present invention. The brain-inspired global-local dual-channel image classification method includes the following steps:
[0055] Step S110: Obtain and divide the image set to be classified into a training set and a test set; the training set is used for training the model, and the test set is used for testing the trained model to obtain a better-performing model. The above division can be a random division or a division according to a preset ratio, which is not limited in this embodiment.
[0056] Step S120: Select a CNN model, use its building units as the local channel basic modules of the dual-channel model, and extract local detailed information from the input features to obtain a feature representation with local information; in this embodiment, selecting a CNN model can be a layer-structured deep convolutional neural network (CNN) model.
[0057] The obtained feature representation with local information can be: Denote the input feature of the \(l\)th local CNN path as where \(H_1, W_1, C_1\) represent the height, width, and number of channels of the input feature respectively. Use to represent the \(l\)th local information processing channel of the dual-channel image classification model. Then the output feature of the \(l\)th local channel of the global-local dual-path image classification model is
[0058] Step S130: Select a Transformer model, use its encoding layer components as the global channel basic modules of the dual-channel model, and extract global topological correlation information from the input features to obtain a feature representation with global information; in this embodiment, it can be to represent the \(l\)th Transformer encoding layer processing path, as the \(l\)th global information processing channel of the dual-channel image classification model. The input and output features of this layer channel are the same as the local channel input features in Step S120. Therefore, the output feature of the \(l\)th global channel of the global-local dual-path image classification model is
[0059] The \(l\)th global path needs to achieve two goals, namely: (1) Generate a feature representation that captures global context information in the input feature map \(X\); (2) Output which must be aligned with the fine local feature representation generated by the local path so that they can be fused together to form the output feature map \(Y\) l . Transformer is mainly applied in the natural language field. In order to extract the global topological information of the input features and facilitate the fusion with the output of the local path , please refer to Figure 3 , Figure 3 which is a novel feature-to-feature (F2F) Transformer module provided by the embodiment of the present invention. The present invention proposes a new Transformer encoder module, called the Feature-to-Feature (F2F) module.
[0060] The F2F module is an improvement on the standard Transformer encoder. The input of the standard encoder is N one-dimensional tokens obtained by splitting and linearly transforming the original image. In the present invention, both the input and output of the dual-channel building block are feature maps. Therefore, F2T and T2F modules are required to perform the conversion between feature maps and tokens. The global information processing channel of the l-th layer is constructed by stacking F2F modules, and each F2F module consists of the following three modules:
[0061] (1) Feature to Token (F2T) module, which converts the input three-dimensional feature map into N one-dimensional tokens. The core component of the Transformer encoder is the MHSA. The input of the MHSA is N one-dimensional tokens, and then the similarity of these tokens is calculated, and finally an output with global characteristics is obtained. Converting the three-dimensional feature map into N tokens, so that each token represents a part of the feature map, and then inputting it into the MHSA is actually calculating the relationship between different parts of the three-dimensional feature map. Thus, it is convenient for the MHSA module to calculate and extract the global information of the three-dimensional features.
[0062] (2) Multi-Head Self-Attention module (MHSA), which generates K groups of global topological representations from N input tokens; the essence of multi-head is multiple independent attention calculations (K heads calculate K groups of independent topological representations), serving as an integrated role to prevent overfitting; compared with single-head, multi-head can extract more features (K groups of global topological representations), and the information of the expressed features will be richer.
[0063] (3) Token to Feature (T2F) module, which converts the K groups of N output tokens in the MHSA into a three-dimensional output feature map. The output of the l-th layer must be aligned with the fine-grained local feature representation generated by the local path so that they can be fused together to form the output feature map Y l . Since only features of the same scale can be fused, tokens need to be converted into three-dimensional features to achieve feature fusion.
[0064] The extraction of global topological correlation information from the input features through the new Transformer encoder module to obtain a feature representation with global information includes the following steps: First, convert the input feature map into N one-dimensional tokens. Then, generate K groups of global topological representations from the N one-dimensional tokens.
[0065] Finally, convert the N output tokens of the K groups of global topological representations into a multi-dimensional output feature map to obtain a feature representation with global information.
[0066] The new Transformer encoder module constructed through the F2F module can not only extract the global topological information of the input features, but also has the same size of input and output as the local path to facilitate feature fusion. At the same time, it can also prevent the influence of different feature scales on the fusion effect.
[0067] Step S140: Use the local channel basic module and the global channel basic module of the dual-channel model as parallel dual-channels, and connect them to the modulator respectively to form a dual-channel construction unit. The output features of the two dual-channels are fused through the modulator to obtain the output features of the dual-channel module. Specifically, select the CNN module of the local channel in step S120 and the Transformer encoding layer module of the global path in step S130 as parallel dual-channels, and then fuse the output features of the two channels through a top-down modulator to form a construction unit of the dual-channel image classification model, and finally obtain the output features of one layer of the dual-channel module.
[0068] Step S150: Stack multiple such dual-channel construction units according to the hierarchical structure of the CNN model to obtain a global-local dual-channel image classification model; thus, a new global-local dual-path model architecture is constructed, providing a computational model based on the dual-channel visual recognition mechanism of the human brain, called the Global-Local network (abbreviation: GLNet). Please refer to Figure 2 , Figure 2 which is the architecture diagram of the brain-inspired global-local dual-path model provided by the embodiment of the present invention. The single-layer structure of the global-local dual-channel image classification model consists of parallel global and local channels and a top-down feature modulation mechanism. The above global-local dual-channel image classification model may include multiple dual-channel construction units to construct a low-level dual-channel model construction layer, a high-level dual-channel model construction layer, etc. The above stacking according to the hierarchical structure of the CNN model may be that the input image first enters the low-level dual-channel model construction layer and the high-level dual-channel model construction layer constructed by the dual-channel construction unit, then outputs to the fully connected layer, and finally outputs the classification result.
[0069] Among them, the modulator fuses the feature representation with global information and the feature representation with local information in a gated mechanism and in a learnable manner to generate the output features of the dual-channel module. The modulator uses the global topological features from g (M a ) and in a learnable manner (M ) to modulate the local fine feature representation from to generate the final output Y l . Specifically, it can be obtained through the following two modulation mechanisms:
[0070] (1) The modulator modulates the output features of the dual channels through a gating mechanism to obtain the output features of the dual-channel module. The expression of the output features of the dual-channel module is: where σ(·) is the sigmoid activation function, is the feature representation with global information at the l-th layer, represents the feature representation with local information at the l-th layer, Y l is the output feature of the dual-channel module at the l-th layer.
[0071] (2) The modulator adopts an additive fusion method to dynamically learn the proportional relationship between the features of the dual channels by using the learnable parameter λ, and obtains the output features of the dual-channel module. The expression of the obtained output features of the dual-channel module is: where is the ReLU activation function, is the feature representation with global information at the l-th layer, represents the feature representation with local information at the l-th layer, Y l is the output feature of the dual-channel module at the l-th layer, and λ is to dynamically learn the proportional relationship between the features of the dual channels.
[0072] Experiments show that the second modulation mechanism can dynamically learn the correlation between the two channels, is more suitable for the training of the dual-channel model, and the model accuracy is higher.
[0073] Among them, the local channel basic module in the dual-channel construction unit is composed of N L stacked CNN model construction units, and the global channel basic module in the dual-channel construction unit is composed of N G stacked encoding layer components of the Transformer model. The number of single-layer structures, the inter-layer organization method, and the fully connected layer of the global-local dual-channel image classification model are the same as those of the deep convolutional neural network model.
[0074] Step S160: Train the global-local dual-channel image classification model with the training set to obtain a trained global-local dual-channel image classification model; according to the batch-based stochastic gradient descent method and some data augmentation methods, use the training set to train the global-local dual-channel image classification model in step S150. The process of training the above model can be realized by existing technologies and will not be elaborated here.
[0075] Step S170: Classify the test set using the trained global-local dual-channel image classification model to obtain the image classification result. After the model is trained, test the test set of the image to be classified to complete the classification prediction.
[0076] In the above implementation process, the image set to be classified is divided into a training set and a test set; then a deep convolutional neural network (CNN) model with a layer structure is selected, and its building unit is used as the local channel basic module of the dual-channel model, and local detailed information of the input features is extracted to obtain a feature representation with local information; a Transformer model is selected, and its encoding layer component is used as the basic module of the global channel of the dual-channel model, and global topological correlation information of the input features is extracted to obtain a feature representation with global information; the CNN module of the local channel and the Transformer encoding layer module of the global path are used as parallel dual channels, and then the output features of the two channels are fused through a top-down modulator to form a building unit of the dual-channel image classification model, and finally the output features of a layer of dual-channel modules are obtained; stack the dual-channel building units, and build the model according to the hierarchical architecture of the CNN model to obtain the global-local dual-channel image classification model; then, according to the batch-based stochastic gradient descent method and some data augmentation methods, use the training set for training. After the model is trained, test the test set of the image to be classified to complete the classification prediction. By adding a new global path representation method and combining a top-down feature modulation mechanism, a global-local dual-path classification model is constructed, providing a computational model based on the dual-channel visual recognition mechanism of the human brain, called the Global-Local network (abbreviation: GLNet), which can make full use of the local fine-grained information and global spatial topological information of the image, and propose a general principle for converting the typical layer structure deep convolutional neural network (CNN) structure into the global-local dual-path classification model structure, and apply this conversion to some representative baseline CNN models, thereby improving the classification accuracy of the model and enhancing the robustness and generalization ability of the model.
[0077] The present invention has been tested on multiple image classification data sets (CIFAR10, CIFAR100, ImageNet, fine-grained image classification data set CUB Bird-200), and the experimental results show that the present invention significantly improves the accuracy of the model.
[0078] As shown in Table 1: Comparison Table of Test Accuracy on Thumbnail (CIFAR10 / 100) Image Classification Datasets. The present invention compares the test accuracy with other benchmark methods on the thumbnail (CIFAR10 / 100) image classification dataset. The benchmark methods for comparison are: ResNet / ResNeXt (ResNet20, ResNet56, ResNeXt29), SENet method based on the "squeeze and excitation" mechanism combined with existing CNN blocks, GCNet method based on the attention mechanism to capture the global context information of the input image, and brain-inspired global-local dual-path network GLNet method. It can be seen from the data results in Table 1 that the accuracy rate and improvement rate of the brain-inspired global-local dual-path network GLNet method of the present invention are higher than those of other methods.
[0079] Table 1: Comparison Table of Test Accuracy on Thumbnail (CIFAR10 / 100) Image Classification Datasets
[0080]
[0081]
[0082] As shown in Table 2: Comparison Table of Test Accuracy on ImageNet Image Classification Dataset and CUB Bird-200 Fine-Grained Image Classification Dataset. The present invention compares the test accuracy with other benchmark methods on the ImageNet image classification dataset and CUB Bird-200 fine-grained image classification dataset. The benchmark methods for comparison are: ResNet / ResNeXt (ResNet18, ResNet50, ResNeXt50), SENet method based on the "squeeze and excitation" mechanism combined with existing CNN blocks, GCNet method based on the attention mechanism to capture the global context information of the input image, and brain-inspired global-local dual-path network GLNet method. It can be seen from the data results in Table 2 that the accuracy rate and improvement rate of the brain-inspired global-local dual-path network GLNet method of the present invention are higher than those of other methods.
[0083] Table 2: Comparison Table of Test Accuracy on ImageNet Image Classification Dataset and CUB Bird-200 Fine-Grained Image Classification Dataset
[0084]
[0085]
[0086] Based on the same inventive concept, the present invention also proposes a brain-inspired global-local dual-channel image classification system. Please refer to Figure 4 , Figure 4A structural block diagram of a brain-inspired global-local dual-channel image classification system provided by an embodiment of the present invention. The brain-inspired global-local dual-channel image classification system includes:
[0087] An image set partitioning module 110 for the image to be classified, configured to obtain and partition the image set to be classified into a training set and a test set;
[0088] A local path module 120, configured to select a CNN model, use its building unit as the local channel basic module of the dual-channel model, and extract local detailed information from the input features to obtain a feature representation with local information;
[0089] A global path module 130, configured to select a Transformer model, use its encoding layer component as the global channel basic module of the dual-channel model, and extract global topological correlation information from the input features to obtain a feature representation with global information;
[0090] A modulation module 140, configured to use the local channel basic module and the global channel basic module of the dual-channel model as parallel dual channels, and connect them to a modulator respectively to form a dual-channel building unit, and fuse the output features of the dual channels through the modulator to obtain the output features of the dual-channel module;
[0091] A model construction module 150, configured to stack a plurality of the dual-channel building units according to the hierarchical architecture of the CNN model to obtain a global-local dual-channel image classification model;
[0092] A model training module 160, configured to train the global-local dual-channel image classification model using the training set to obtain a trained global-local dual-channel image classification model;
[0093] An image classification module 170, configured to classify the test set using the trained global-local dual-channel image classification model to obtain an image classification result.
[0094] In the above implementation process, the image set to be classified is divided into a training set and a test set by the image set division module 110 to be classified; the local path module 120 selects a layer-structured deep convolutional neural network (CNN) model, uses its building unit as the local channel basic module of the dual-channel model, and extracts local detailed information from the input features to obtain a feature representation with local information; the global path module 130 selects a Transformer model, uses its encoding layer component as the basic module of the global channel of the dual-channel model, and extracts global topological correlation information from the input features to obtain a feature representation with global information; the modulation module 140 takes the CNN module of the local channel and the Transformer encoding layer module of the global path as parallel dual channels, and then fuses the output features of the two channels through a top-down modulator to form a building unit of the dual-channel image classification model, and finally obtains the output features of a layer of dual-channel modules; the model construction module 150 stacks the dual-channel building units and constructs the model according to the hierarchical architecture of the CNN model to obtain a global-local dual-channel image classification model; the model training module 160 trains using the training set according to the batch-based stochastic gradient descent method and some data augmentation methods. After the model is trained, the image classification module 170 tests the test set of the image to be classified to complete the classification prediction. By adding a new global path representation method and combining a top-down feature modulation mechanism, a global-local dual-path classification model is constructed, providing a computational model based on the dual-channel visual recognition mechanism of the human brain, called the Global-Local network (abbreviation: GLNet). It can make full use of the local fine-grained information and global spatial topological information of the image, and convert the typical layer-structured deep convolutional neural network (CNN) structure into the general principle of the global-local dual-path classification model structure, and apply this conversion to some representative baseline CNN models, thereby improving the classification accuracy of the model and enhancing the robustness and generalization ability of the model.
[0095] Please refer to Figure 5 , Figure 5A schematic structural block diagram of an electronic device provided by an embodiment of the present application. The electronic device includes a memory 101, a processor 102, and a communication interface 103. The memory 101, the processor 102, and the communication interface 103 are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory 101 can be used to store software programs and modules, such as program instructions / modules corresponding to a brain-inspired global-local dual-channel image classification system provided by an embodiment of the present application. The processor 102 executes various functional applications and data processing by executing the software programs and modules stored in the memory 101. The communication interface 103 can be used for signaling or data communication with other node devices.
[0096] Among them, the memory 101 can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0097] The processor 102 can be an integrated circuit chip with signal processing capabilities. The processor 102 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0098] It can be understood that Figure 5 The structure shown is only schematic, and the electronic device may further include more or fewer components than those shown Figure 5 herein, or have a different configuration from that shown Figure 5 herein. Figure 5 Each component shown herein can be implemented by hardware, software, or a combination thereof.
[0099] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0100] In addition, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0101] If the above functions are implemented in the form of software function modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, etc., which can store program codes.
[0102] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0103] It is obvious to those skilled in the art that the present application is not limited to the details of the above-described exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present application. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A brain-inspired global-local dual-channel image classification method, characterized in that Including the following steps: Obtain and divide the image set to be classified into a training set and a test set; Select a CNN model, use its building unit as the local channel basic module of the dual-channel model, and extract local detail information from the input features to obtain a feature representation with local information; Select a Transformer model, use its encoding layer component as the global channel basic module of the dual-channel model, and extract global topological correlation information from the input features to obtain a feature representation with global information; Use the local channel basic module of the dual-channel model and the global channel basic module of the dual-channel model as parallel dual channels, and connect them to the modulator respectively to form a dual-channel building unit. The output features of the dual channels are fused through the modulator to obtain the output features of the dual-channel module; Stack multiple said dual-channel building units according to the hierarchical architecture of the CNN model to obtain a global-local dual-channel image classification model; Train the global-local dual-channel image classification model with the training set to obtain a trained global-local dual-channel image classification model; Use the trained global-local dual-channel image classification model to classify the test set to obtain an image classification result; The extracting of global topological correlation information from the input features to obtain a feature representation with global information includes the following steps: mapping and converting the input features into N one-dimensional Tokens; Generating K groups of global topological representations from the N one-dimensional Tokens; Converting the N output Tokens of the K groups of global topological representations into a multi-dimensional output feature map to obtain a feature representation with global information; The fusing of the output features of the dual channels through the modulator to obtain the output features of the dual-channel module includes the following steps: The modulator fuses the feature representation with global information and the feature representation with local information in a learnable manner through a gating mechanism to generate the output features of the dual-channel module.
2. The brain-inspired global-local dual-channel image classification method according to claim 1, wherein The fusing of the feature representation with global information and the feature representation with local information to generate the output features of the dual-channel module includes the following steps: The modulator modulates the output features of the dual channels through a gating mechanism to obtain the output features of the dual-channel module. The expression of the output features of the dual-channel module is as follows: where σ(·) is the sigmoid activation function, is the feature representation with global information at the l-th layer, represents the feature representation with local information at the l-th layer, Y l is the output feature of the dual-channel module at the l-th layer.
3. The brain-inspired global-local dual-channel image classification method according to claim 1, wherein The modulator fuses the feature representation with global information and the feature representation with local information in a learnable manner to generate the output features of the dual-channel module, including the following steps: The modulator adopts an additive fusion method, and uses the learnable parameter λ to dynamically learn the proportional relationship between the features of the two channels, so as to obtain the output features of the two-channel module. The expression for obtaining the output features of the two-channel module is: Where is the ReLU activation function, is the feature representation with global information of the l-th layer, represents the feature representation with local information of the l-th layer, Y l is the output feature of the two-channel module of the l-th layer, and λ is to dynamically learn the proportional relationship between the features of the two channels.
4. The brain-inspired global-local dual-channel image classification method according to claim 1, wherein The local channel basic module in the dual-channel construction unit is composed of N L stacked CNN model construction units, and the global channel basic module in the dual-channel construction unit is composed of N G stacked encoding layer components of Transformer models.
5. A brain-inspired global-local dual-channel image classification system, characterized in that, Including: An image set division module to be classified, configured to obtain and divide the image set to be classified into a training set and a test set; A local path module, configured to select a CNN model, use its building unit as the local channel basic module of the dual-channel model, and extract local detail information from the input features to obtain a feature representation with local information; The global path module is used to select a Transformer model, use its encoding layer components as the global channel basic module of the dual-channel model, and extract global topological correlation information from the input features to obtain a feature representation with global information. The step of extracting global topological correlation information from the input features to obtain a feature representation with global information includes the following steps: mapping and converting the input features into N one-dimensional tokens; generating K groups of global topological representations from the N one-dimensional tokens; converting the N output tokens of the K groups of global topological representations into a multi-dimensional output feature map to obtain a feature representation with global information The modulation module is used to use the local channel basic module and the global channel basic module of the dual-channel model as parallel dual channels, and connect them to the modulator respectively to form a dual-channel construction unit. The modulator fuses the output features of the dual channels to obtain the output features of the dual-channel module. The step of fusing the output features of the dual channels by the modulator to obtain the output features of the dual-channel module includes the following steps: the modulator fuses the feature representation with global information and the feature representation with local information in a gated mechanism and in a learnable manner to generate the output features of the dual-channel module; The model construction module is used to stack multiple of the dual-channel construction units according to the hierarchical architecture of the CNN model to obtain a global-local dual-channel image classification model; The model training module is used to train the global-local dual-channel image classification model with the training set to obtain a trained global-local dual-channel image classification model; The image classification module is used to classify the test set with the trained global-local dual-channel image classification model to obtain an image classification result.
6. An electronic device, characterized in that, Comprising: A memory for storing one or more programs; A processor; When the one or more programs are executed by the processor, the method according to any one of claims 1-4 is implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Federal learning method and system for different agents in intelligent workshop
CN113255937A
Neural network relationship extraction method, computer device, and readable storage medium
WO2021174774A1