An image classification method and system based on binary neural network

By introducing modules such as feature amplification layer and segmented scaling layer into binary neural networks, the feature representation and model fitting capabilities of image classification are improved, the performance degradation problem of binary neural networks is solved, and higher image classification accuracy is achieved.

CN113936169BActive Publication Date: 2025-09-26INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111098800.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-18
Publication Date
2025-09-26
Estimated Expiration
2041-09-18

AI Technical Summary

Technical Problem

Existing binary neural networks have problems with insufficient feature representation and model fitting capabilities in image classification, especially after the weights and activation values ​​are binarized, which leads to performance degradation.

Method used

An improved module structure and method, including feature amplification layer, segmented scaling layer and binary convolution layer, is adopted to construct a binary neural network by stacking neural network modules, utilizing floating-point feature map information, increasing the number of channels and adjusting network performance.

Benefits of technology

On the ImageNet dataset, compared with ReActNet-A's 69.4%, the top 1 accuracy was improved by 2% to 71.46% with the same amount of computation, improving the accuracy of image classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113936169B_ABST
    Figure CN113936169B_ABST
Patent Text Reader

Abstract

The present invention proposes an image classification method and system based on a binary neural network, comprising: constructing a neural network module including a feature amplification layer, a binary convolution layer, an activation layer and a segmented scaling layer, and constructing a binary neural network by stacking the neural network modules; obtaining an image with an image category label marked as training data, inputting a floating-point feature map of the training data into the feature amplification layer of the first module in the binary neural network to amplify the number of channels of the floating-point feature map to obtain an amplified feature map, converting the amplified feature map into a binary feature map and inputting it into the convolution layer to obtain a convolution feature map of the binary feature map, normalizing the convolution feature map and inputting it into the segmented scaling layer, adjusting the floating-point feature map output by the scaling factor of the segmented scaling layer, and passing the result as input to the next neural network module, and using the image category of the training data obtained by the last neural network module as the training result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning and image classification, and in particular to an image classification method and system based on a binary neural network. Background Art

[0002] Binary neural networks (1-bit CNNs), whose weights and activations are both 1-bit binary, offer convolution computational efficiency exceeding 58 times that of full-precision floating-point (fp32) convolutions, while also reducing model storage overhead by 32 times, making them highly valuable for practical applications. However, binarization of activations significantly impacts the representational power of feature maps, significantly reducing the neural network's intermediate feature representation capabilities. Binarization of weights also leads to a single feature extraction method in the convolutional layer, resulting in reduced model fitting. Therefore, the challenge of binary neural networks lies in achieving good results within limited feature representation and model fitting capabilities. Due to the significant differences between binary networks and floating-point or quantized networks, current binarization methods have gradually shifted from general-purpose approaches to specifically designed binary-oriented network structures. For example, ReaActNet proposes a new nonlinear layer, the RPReLU, and ExpertConvNet enhances network performance through multiple weight groups and increased channel counts. Currently, a binary neural network with a stable output value range and distribution, improved feature representation, and improved model fitting capabilities is needed to improve the image classification performance of binary neural networks. Summary of the Invention

[0003] The improvements of the binary neural network structure design for image classification include:

[0004] Improvement point 1: Use a module block structure of convolution, activation and normalization to make the value range and distribution of block outputs in different layers more stable.

[0005] Improvement point 2: A method for efficiently obtaining binary feature maps with more channels is proposed. A feature amplification layer is inserted before the binary convolution to increase the number of channels. This reduces the amount of computation while increasing the representation capability of the multi-channel binary convolution and making full use of the information of the floating-point feature map.

[0006] Improvement 3: Add segmented learnable scaling factors between blocks, which are learned independently for the positive and negative half-axes of each channel. This adjusts the floating-point feature maps transmitted in the network, significantly improving the overall network performance with a very small amount of floating-point calculations.

[0007] Improvement 4: A more reasonable network structure layer stage arrangement was designed. On the basis of including the feature amplification layer, the feature map with a resolution of 14x14 in the network is parsed using 256 channels, making more effective use of binary convolution calculations.

[0008] Specifically, the present invention proposes an image classification method based on a binary neural network, which includes

[0009] Step 1: Construct a neural network module including a feature amplification layer, a binary convolution layer, an activation layer, and a segmented scaling layer, and construct a binary neural network by stacking the neural network module.

[0010] Step 2. Obtain an image marked with an image category label as training data, input the floating-point feature map of the training data into the feature amplification layer of the first module in the binary neural network to amplify the number of channels of the floating-point feature map to obtain an amplified feature map, convert the amplified feature map into a binary feature map and input it into the convolution layer to obtain a convolution feature map of the binary feature map, normalize the convolution feature map and input it into the segmented scaling layer, adjust the floating-point feature map output by the scaling factor module of the segmented scaling layer, and pass the result as input to the next neural network module, and use the image category of the training data obtained by the last neural network module as the training result.

[0011] Step 3: Construct a loss function based on the training results and the image category label, iteratively train the binary neural network until the loss function converges or reaches a preset number of iterations, save the current binary neural network as an image classification model, input the image to be classified into the image classification model, and obtain the image category of the image to be classified.

[0012] The image classification method based on binary neural network, wherein step 2 includes: using a symbol function to convert the amplified feature map into a binary feature map.

[0013] In the image classification method based on binary neural network, the feature amplification layer amplifies the number of channels of the floating-point feature map by adding an offset to the activation value.

[0014] In the binary neural network-based image classification method, the feature amplification layer uses convolution to amplify the number of channels of the floating-point feature map.

[0015] The present invention also proposes an image classification system based on a binary neural network, which includes

[0016] Module 1 is used to construct a neural network module including a feature amplification layer, a binary convolution layer, an activation layer and a segmented scaling layer, and to construct a binary neural network by stacking the neural network modules.

[0017] Module 2 is used to obtain images that have been marked with image category labels as training data, input the floating-point feature map of the training data into the feature amplification layer of the first module in the binary neural network to amplify the number of channels of the floating-point feature map, and obtain an amplified feature map. The amplified feature map is converted into a binary feature map and input into the convolution layer to obtain a convolution feature map of the binary feature map. The convolution feature map is normalized and input into the segmented scaling layer. The floating-point feature map output by the scaling factor adjustment module of the segmented scaling layer is passed, and the result is passed as input to the next neural network module. The image category of the training data obtained by the last neural network module is used as the training result.

[0018] Module 3 is used to construct a loss function based on the training results and the image category label, iteratively train the binary neural network until the loss function converges or reaches a preset number of iterations, save the current binary neural network as an image classification model, input the image to be classified into the image classification model, and obtain the image category of the image to be classified.

[0019] The image classification system based on binary neural network, wherein the module 2 includes: using a symbol function to convert the amplified feature map into a binary feature map.

[0020] In the binary neural network-based image classification system, the feature amplification layer amplifies the number of channels of the floating-point feature map by adding an offset to the activation value.

[0021] In the binary neural network-based image classification system, the feature amplification layer uses convolution to amplify the number of channels of the floating-point feature map.

[0022] It can be seen from the above scheme that the advantages of the present invention are:

[0023] On the ImageNet dataset, the top 1 accuracy is improved by 2% to 71.46% with the same amount of computation, compared with 69.4% of ReActNet-A. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is the binary block structure diagram proposed by the present invention. DETAILED DESCRIPTION

[0025] In order to make the above features and effects of the present invention more clearly understood, embodiments are given below and described in detail with reference to the accompanying drawings.

[0026] The structure diagram of efficient binary block is shown in Figure 1The feature maps in the network backbone and shortcuts are all floating-point; only the convolutional layer (conv) is binary. The Fexpand layer is a feature amplification layer, used to increase the number of channels in the floating-point feature map. A simple sign function is then used to convert the floating-point feature map into a corresponding binary feature map, allowing it to participate in the convolution operation. Following the residual structure is the PW-Scale layer, a piecewise scaling layer, which dynamically scales the floating-point feature map. The Fexpand layer, Sign layer, and PW-Scale layer together form the transformation from floating-point feature maps to binary feature maps. Feature amplification allows multiple binary feature maps to be extracted from a single floating-point feature map, fully utilizing the information in the floating-point feature maps and improving the representation capability of binary convolution. The entire binary network structure is stacked using the modules shown. During inference, an image is input into the network structure. After calculations in each module, the network ultimately outputs the image's classification category.

[0027] Among them, the feature amplification layer expands the number of channels based on convolution, or uses the method of adding a certain offset to the activation value for amplification. By adding different offsets to the activation value, floating-point feature maps with different offsets can be obtained, and the results of the sign function are also different, thereby realizing the extraction of multiple binary feature maps from a floating-point feature map.

[0028] Figure 1 The batchnorm layer in is a batch normalization layer commonly used in neural networks. Its input is a batch of features in the training process. The processing process is that each feature channel is processed independently. The mean, variance and two learnable parameters of each channel are used to adjust the input mean to near 0 and the variance to near 1.

[0029] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in conjunction with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0030] The present invention also proposes an image classification system based on a binary neural network, which includes

[0031] Module 1 is used to construct a neural network module including a feature amplification layer, a binary convolution layer, an activation layer and a segmented scaling layer, and to construct a binary neural network by stacking the neural network modules.

[0032] Module 2 is used to obtain images that have been marked with image category labels as training data, input the floating-point feature map of the training data into the feature amplification layer of the first module in the binary neural network to amplify the number of channels of the floating-point feature map, and obtain an amplified feature map. The amplified feature map is converted into a binary feature map and input into the convolution layer to obtain a convolution feature map of the binary feature map. The convolution feature map is normalized and input into the segmented scaling layer. The floating-point feature map output by the scaling factor adjustment module of the segmented scaling layer is passed, and the result is passed as input to the next neural network module. The image category of the training data obtained by the last neural network module is used as the training result.

[0033] Module 3 is used to construct a loss function based on the training results and the image category label, iteratively train the binary neural network until the loss function converges or reaches a preset number of iterations, save the current binary neural network as an image classification model, input the image to be classified into the image classification model, and obtain the image category of the image to be classified.

[0034] The image classification system based on binary neural network, wherein the module 2 includes: using a symbol function to convert the amplified feature map into a binary feature map.

[0035] In the binary neural network-based image classification system, the feature amplification layer amplifies the number of channels of the floating-point feature map by adding an offset to the activation value.

[0036] In the binary neural network-based image classification system, the feature amplification layer uses convolution to amplify the number of channels of the floating-point feature map.

Claims

1. An image classification method based on binary neural network, characterized in that: include Step 1: Construct a neural network module including a feature amplification layer, a binary convolution layer, an activation layer, and a segmented scaling layer, and construct a binary neural network by stacking the neural network modules; Step 2: Obtain an image marked with an image category label as training data, input the floating-point feature map of the training data into the feature amplification layer of the first module in the binary neural network to amplify the number of channels of the floating-point feature map to obtain an amplified feature map, convert the amplified feature map into a binary feature map and input it into the convolution layer to obtain a convolution feature map of the binary feature map, activate and normalize the convolution feature map, and input it into the segmented scaling layer through the addition function of the residual structure, adjust the floating-point feature map output by the scaling factor module of the segmented scaling layer, and pass the result as input to the next neural network module, and use the image category of the training data obtained by the last neural network module as the training result; Step 3: Construct a loss function based on the training results and the image category label, iteratively train the binary neural network until the loss function converges or reaches a preset number of iterations, save the current binary neural network as an image classification model, input the image to be classified into the image classification model, and obtain the image category of the image to be classified.

2. The image classification method based on binary neural network according to claim 1, characterized in that: The step 2 includes: using a sign function to convert the amplified feature map into a binary feature map.

3. The image classification method based on binary neural network according to claim 1, characterized in that: The feature amplification layer amplifies the number of channels of the floating-point feature map by adding an offset to the activation value.

4. The image classification method based on binary neural network according to claim 1, characterized in that: The feature amplification layer uses convolution to amplify the number of channels of the floating-point feature map.

5. An image classification system based on a binary neural network, characterized in that: include Module 1 is used to construct a neural network module including a feature amplification layer, a binary convolution layer, an activation layer, and a segmented scaling layer, and to construct a binary neural network by stacking the neural network modules; Module 2 is used to obtain images marked with image category labels as training data, input the floating-point feature map of the training data into the feature amplification layer of the first module in the binary neural network to amplify the number of channels of the floating-point feature map to obtain an amplified feature map, convert the amplified feature map into a binary feature map and input it into the convolution layer to obtain a convolution feature map of the binary feature map, after activation and normalization, the convolution feature map is input into the segmented scaling layer through the addition function of the residual structure, and the floating-point feature map output by the scaling factor adjustment module of the segmented scaling layer is passed as input to the next neural network module, and the image category of the training data obtained by the last neural network module is used as the training result; Module 3 is used to construct a loss function based on the training results and the image category label, iteratively train the binary neural network until the loss function converges or reaches a preset number of iterations, save the current binary neural network as an image classification model, input the image to be classified into the image classification model, and obtain the image category of the image to be classified.

6. The image classification system based on binary neural network according to claim 5, characterized in that: The module 2 includes: using a sign function to convert the amplified feature map into a binary feature map.

7. The image classification system based on binary neural network according to claim 5, characterized in that: The feature amplification layer amplifies the number of channels of the floating-point feature map by adding an offset to the activation value.

8. The image classification system based on binary neural network according to claim 5, characterized in that: The feature amplification layer uses convolution to amplify the number of channels of the floating-point feature map.

Citation Information

Patent Citations

  • Apparatus, method and computer program product for quantizing neural networks

    US20230412806A1