Image classification method and device, electronic device, and storage medium
By introducing fuzzy logic technology and the characteristics of human vision systems into the image classification method, the processing problems of image data noise and uncertainty in the big data environment are solved, the accuracy and reliability of image classification are improved, and the robustness of the system is enhanced.
Patent Information
- Application Number
- CN202411310991.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-09-20
AI Technical Summary
In a big data environment, image data may contain a lot of noise and unpredictable uncertainty, which increases the difficulty of data processing and reduces the accuracy and reliability of image classification tasks.
By introducing fuzzy logic technology and selective and regional characteristics of human vision systems, an image classification method is provided. The method includes preprocessing the input image, calculating the membership of the input feature as an additional feature using the Gaussian membership function, and blurring it in the convolutional layer and the pooling layer, and finally classifying it through the fully connected layer and the softmax function.
Effectively dealing with noise and uncertainty in image data in big data environments improves the accuracy and reliability of image classification tasks, especially in cases of serious feature ambiguity. At the same time, the system's tolerance for input data noise is enhanced, making the system more robust when processing complex, noisy image data.
Smart Images

Figure CN118823490B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of machine learning and image processing, and in particular to an image classification method and device, an electronic device, and a storage medium. Background Art
[0002] In the field of deep learning, Convolutional Neural Networks (CNN) are widely used in image classification and recognition tasks due to their excellent feature extraction capabilities. CNN simulates the hierarchical processing method of the human visual system through structures such as convolutional layers and pooling layers to effectively extract the spatial features of images. However, with the rapid growth of data scale, image data in big data environments may contain a lot of noise and unpredictable uncertainties, which not only increases the difficulty of data processing, but also puts higher requirements on the accuracy and reliability of image classification tasks. Summary of the invention
[0003] The present disclosure aims to solve at least one of the problems existing in the prior art, and provides an image classification method and device, an electronic device, and a storage medium to enhance the convolutional neural network's ability to handle uncertainty and noise in images, and improve the accuracy of extracting key feature information in images, thereby enabling the convolutional neural network to achieve a higher performance level in handwriting and clothing image classification tasks.
[0004] In one aspect of the present disclosure, there is provided an image classification method, the image classification method comprising:
[0005] Performing a preprocessing operation on the input image to adjust the pixel value of the input image to a preset range to obtain a corresponding normalized image;
[0006] The normalized image is input to a convolution layer for processing. Before each convolution operation, a Gaussian membership function is used to calculate the membership of the input feature to different fuzzy rules as the corresponding additional feature, the additional feature and the corresponding input feature are added to obtain an enhanced feature, and a convolution operation is performed on the enhanced feature to obtain a feature map after fuzzification processing by the convolution layer as the output result of the convolution layer;
[0007] The output result of the convolution layer is input into the pooling layer, and fuzzification is performed in the order of channel dimension and spatial dimension to obtain the feature vector output by the pooling layer;
[0008] The feature vector is input into a fully connected layer, and the output result of the fully connected layer is classified using a softmax function.
[0009] Optionally, the using of the Gaussian membership function to calculate the membership of the input feature to different fuzzy rules as the corresponding additional features includes:
[0010] Calculate the membership of the input feature: The input feature of each convolution Mapped to the corresponding fuzzy sets In the figure, each fuzzy set corresponds to a Gaussian membership function, generating the corresponding center and bandwidth. According to the calculation formula of the Gaussian membership function, the input features are calculated respectively. The corresponding membership degree, wherein the calculation formula of the Gaussian membership function is expressed as:
[0011] ;
[0012] in, Represents the tth input feature of the kth fuzzy rule The corresponding fuzzy set; Represents input features In fuzzy sets The degree of membership on ; Represents the tth input feature of the kth fuzzy rule The center of the Gaussian function; Represents the tth input feature of the kth fuzzy rule The bandwidth of the Gaussian function is d; the number of fuzzy rules and the number of input features are both d;
[0013] Calculate the activation degree of the fuzzy rule: Calculate the activation degree of each fuzzy rule according to the calculation formula of the rule activation degree, wherein the calculation formula of the rule activation degree is expressed as:
[0014] ;
[0015] in, Represents the activation degree of input feature x for the kth fuzzy rule; input feature x includes input feature ;The fuzzy rule layer performs the AND operator, which is implemented using the multiplication operation;
[0016] Calculate the additional features of the input features: Calculate the normalized membership of the input features for all fuzzy rules according to the calculation formula of the normalized membership, where the calculation formula of the normalized membership is expressed as:
[0017] ;
[0018] in, represents the normalized membership of the input feature x to the kth fuzzy rule, and also represents the input feature Additional features of Represents the sum of activations of the input feature x for all fuzzy rules; represents the fuzzy rule number and .
[0019] Optionally, the output result of the convolution layer is input to a pooling layer, and fuzzification is performed in the order of channel dimension and spatial dimension to obtain a feature vector output by the pooling layer, including:
[0020] Determine a preliminary enhanced feature map: in the channel dimension, use a Gaussian membership function to calculate the maximum fuzzy membership value map of each feature map in the output result of the convolution layer, calculate the main channel weight and the additional channel weight of each feature map according to the maximum fuzzy membership value map of each feature map, use the corresponding product of the main channel weight and the additional channel weight as the channel weight, multiply the channel weight by the corresponding feature map, and obtain a preliminary enhanced feature map after blur processing in the channel dimension;
[0021] Determine the secondary enhanced feature map: in the spatial dimension, use the Gaussian membership function to calculate the maximum fuzzy membership value map of the preliminary enhanced feature map, calculate the main spatial weight and the additional spatial weight of the preliminary enhanced feature map according to the maximum fuzzy membership value map of the preliminary enhanced feature map, use the corresponding product of the main spatial weight and the additional spatial weight as the spatial weight, multiply the spatial weight by the corresponding preliminary enhanced feature map, and obtain the secondary enhanced feature map after fuzzy processing in the spatial dimension;
[0022] The secondary enhanced feature map is pooled. For the pooling window, the downsampling mode is set to a summation operation, and the sum of all feature values in the pooling window is calculated as the feature after pooling to obtain the feature vector output by the pooling layer.
[0023] Optionally, determining the preliminary enhanced feature map specifically includes:
[0024] In the channel dimension, determine the fuzzy partition number V, set the number of Gaussian membership functions equal to the fuzzy partition number, generate the center and kernel width of the Gaussian membership function, use the Gaussian membership function to calculate the membership of all features on the feature map, and obtain V fuzzy membership value maps corresponding to each feature map; wherein, the calculation formula of the vth Gaussian membership function corresponding to the nth feature map is expressed as:
[0025] ;
[0026] Among them, Z represents the number of channels of the feature map; Represents the center of the vth Gaussian function corresponding to the feature map with n channels; Represents the maximum value of all elements in the feature map with n channels; Represents the minimum value of all elements in the feature map with n channels; Indicates that the position in the feature map with the number of channels n is Characteristic elements of Represents the fuzzy membership value graph of the vth Gaussian function corresponding to the feature map with n channels Characteristic elements The corresponding membership degree; Indicates the kernel width of the v-th Gaussian function corresponding to the feature map with n channels;
[0027] Spatial aggregation is performed on each fuzzy membership value graph using fuzzy algebra and operators , the spatial aggregation calculation formula of the fuzzy membership value graph is expressed as:
[0028] ;
[0029] in, Represents the fuzzy membership value graph of the vth Gaussian function corresponding to the nth feature map The spatial aggregation value of; H represents the height of the feature map; W represents the width of the feature map;
[0030] Select the fuzzy membership value map with the largest spatial aggregation value as the maximum membership value map of the corresponding feature map;
[0031] Calculate the main channel weights of all feature maps, where the main channel weight of the nth feature map is expressed as:
[0032] ;
[0033] in, Represents the main channel weight of the nth feature map; Represents the spatial aggregation value of the maximum membership value map of the nth feature map; Represents the sum of the spatial aggregation values of the maximum membership value maps of all feature maps;
[0034] Calculate the additional channel weights of all feature maps, where the additional channel weight of the nth feature map is expressed as:
[0035] ;
[0036] in, Represents the additional channel weight of the nth feature map;
[0037] Calculate the channel weights of all feature maps, where the calculation formula for the channel weight of the nth feature map is:
[0038] ;
[0039] in, Represents the channel weight of the nth feature map;
[0040] The channel weights of all feature maps are multiplied by the feature map output by the convolutional layer to obtain the preliminary enhanced feature map. The calculation formula of the preliminary enhanced feature map is expressed as:
[0041] ;
[0042] in, Representation feature map The channel weights of Represents dot multiplication operation; Represents the preliminary enhanced feature map.
[0043] Optionally, determining the secondary enhanced feature map specifically includes:
[0044] In the spatial dimension, the maximum fuzzy membership value map of the preliminary enhanced feature map is calculated, and the calculation formula is:
[0045] ;
[0046] in, Represents the maximum fuzzy membership value map of the preliminary enhanced feature map, whose size is ; The maximum fuzzy membership value map representing the initial enhanced feature map with n channels is ; Represents the vth fuzzy membership value map corresponding to the preliminary enhanced feature map with n channels The spatial aggregation value of
[0047] Calculate the main spatial weights of the preliminary enhanced feature map , whose size is , the calculation formula is expressed as:
[0048] ;
[0049] Calculate additional spatial weights for the preliminary enhanced feature map , whose size is , the calculation formula is expressed as:
[0050] ;
[0051] Among them, K represents the size A matrix whose elements are ; express The "same" mode is used for two-dimensional convolution operations; Represents the sum of the maximum fuzzy membership value map of the preliminary enhanced feature map in the channel dimension;
[0052] Calculate the spatial weights of the preliminary enhanced feature map , whose size is , the calculation formula is expressed as:
[0053] ;
[0054] The spatial weight of the preliminary enhanced feature map is multiplied by the preliminary enhanced feature map to obtain the secondary enhanced feature map. The calculation formula is expressed as:
[0055] ;
[0056] in, Represents the secondary enhanced feature map.
[0057] Optionally, the image classification method further includes:
[0058] The convolutional neural network composed of the convolutional layer, the pooling layer and the fully connected layer is trained, and the weight parameters of the convolutional neural network are updated and iterated using the gradient descent method.
[0059] Optionally, the training of the convolutional neural network composed of the convolutional layer, the pooling layer, and the fully connected layer, and updating and iterating the weight parameters of the convolutional neural network using a gradient descent method, includes:
[0060] All samples of the training set are shuffled and randomly sorted, and the sorted samples are used to train the convolutional neural network;
[0061] During the training process, the weight parameters of the convolutional neural network are iteratively updated by using a mini-batch gradient descent method and a momentum gradient descent method;
[0062] Among them, the activation function adopted by the convolution layer includes a sigmoid function or a ReLU function, and the loss function adopted includes a cross entropy loss function.
[0063] Another aspect of the present disclosure provides an image classification device, the image classification device comprising:
[0064] A preprocessing module, used to perform a preprocessing operation on the input image, adjust the pixel value of the input image to a preset range, and obtain a corresponding normalized image;
[0065] A convolution module, used for inputting the normalized image into a convolution layer for processing, and before each convolution operation, using a Gaussian membership function to calculate the membership of the input feature to different fuzzy rules as the corresponding additional feature, adding the additional feature and the corresponding input feature to obtain an enhanced feature, performing a convolution operation on the enhanced feature, and obtaining a feature map after fuzzification processing by the convolution layer as the output result of the convolution layer;
[0066] A pooling module, used to input the output result of the convolution layer into the pooling layer, perform fuzzy processing in the order of channel dimension and spatial dimension, and obtain a feature vector output by the pooling layer;
[0067] The classification module is used to input the feature vector into the fully connected layer and classify the output result of the fully connected layer using the softmax function.
[0068] Optionally, the image classification device further includes:
[0069] The iterative module is used to train the convolutional neural network composed of the convolutional layer, the pooling layer, and the fully connected layer, and to iterate and update the weight parameters of the convolutional neural network using the gradient descent method.
[0070] Another aspect of the present disclosure provides an electronic device, including:
[0071] at least one processor; and,
[0072] a memory communicatively connected to at least one processor; wherein,
[0073] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image classification method described above.
[0074] Another aspect of the present disclosure provides a computer-readable storage medium storing a computer program, which implements the image classification method described above when executed by a processor.
[0075] Compared with the prior art, the present invention has the following beneficial effects:
[0076] 1. By introducing fuzzy logic technology, it is possible to effectively handle the noise and unpredictable uncertainty contained in image data in a big data environment and reduce data uncertainty. Reducing data uncertainty is crucial to improving the accuracy and reliability of image classification tasks, especially when feature ambiguity is a serious problem.
[0077] 2. By drawing on the selectivity and regional characteristics of the human visual system, the image classification method provided by the present invention can better simulate the superior performance of the human visual system when processing complex image data, which not only improves the accuracy of key feature extraction, but also improves the overall processing efficiency, thereby achieving better results in various image classification tasks.
[0078] 3. The present disclosure introduces fuzzy logic technology, which not only helps to process uncertain information, but also enhances the system's tolerance to input data noise, making the system more robust when processing complex, noisy image data. Compared with conventional CNN models, the image classification method provided by the present disclosure shows more outstanding performance in the classification tasks of different types of images such as handwriting and clothing. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0080] Figure 1 A flowchart of an image classification method provided in one embodiment of the present disclosure;
[0081] Figure 2 The overall architecture diagram of a convolutional neural network improved based on fuzzy logic technology provided by another embodiment of the present disclosure;
[0082] Figure 3 An overall architecture diagram of a channel and spatial attention mechanism provided for another embodiment of the present disclosure;
[0083] Figure 4 A sub-architecture diagram of a channel attention mechanism provided for another embodiment of the present disclosure;
[0084] Figure 5 A sub-architecture diagram of a spatial attention mechanism provided for another embodiment of the present disclosure;
[0085] Figure 6 A schematic diagram of the structure of an image classification device provided by another embodiment of the present disclosure;
[0086] Figure 7 A schematic structural diagram of an electronic device provided in another embodiment of the present disclosure. DETAILED DESCRIPTION
[0087] In order to make the purpose, technical scheme and advantages of the embodiments of the present disclosure clearer, the embodiments of the present disclosure will be described in detail below in conjunction with the accompanying drawings. However, it can be understood by those skilled in the art that in each embodiment of the present disclosure, many technical details are proposed in order to enable readers to better understand the present disclosure. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed for protection in the present disclosure can also be implemented. The division of the following embodiments is for the convenience of description and should not constitute any limitation on the specific implementation of the present disclosure. The various embodiments can be combined and referenced with each other without contradiction.
[0088] One embodiment of the present disclosure relates to an image classification method, which introduces fuzzy logic technology and human visual activity characteristics and is implemented based on a convolutional neural network.
[0089] Fuzzy logic technology is a method that can handle uncertainty. It can effectively handle the uncertainty in the original data by introducing fuzzy sets and fuzzy logic, thereby improving the accuracy of data processing. Combining fuzzy logic technology with convolutional neural networks can improve the performance of convolutional neural networks, especially in image classification, and effectively handle uncertainty and noise in the original data.
[0090] The human visual system has two key characteristics: selectivity and regionality. The selectivity characteristic can selectively focus on specific features and locations through channel attention and spatial attention, thereby significantly improving processing efficiency. The regional characteristics are divided into central vision and peripheral vision. The former is responsible for high-resolution detail perception, such as reading and recognizing faces, and the latter is responsible for extensive environmental monitoring and dynamic perception, enabling the network to process information at different scales and regions, thereby achieving more sophisticated and comprehensive image analysis. Based on these two characteristics, the human visual system can efficiently process information in complex visual environments, and can focus on subtle visual details while fully perceiving changes in the surrounding environment.
[0091] Based on the above content, this embodiment provides a convolutional neural network image classification method based on fuzzy logic technology and human visual activity characteristics. This method not only enhances the network's ability to handle uncertainty and noise in images, but also improves the accuracy of extracting key feature information in images, thereby significantly improving the performance of image classification. Compared with the traditional CNN model, the convolutional neural network image classification method based on fuzzy logic technology proposed in this embodiment shows better performance in the classification tasks of different types of pictures such as handwriting and clothing.
[0092] Combined Figure 1 , the image classification method provided in this embodiment includes:
[0093] Step 1: Perform preprocessing operations on the input image, adjust the pixel values of the input image to a preset range, and obtain the corresponding normalized image.
[0094] Specifically, step 1 mainly performs data preprocessing and normalizes the input data. The pixel values of the input image can be adjusted to a range of 0 to 1 through the normalization operation to obtain a corresponding normalized image.
[0095] Step 2: Input the normalized image into the convolution layer for processing. Before each convolution operation, use the Gaussian membership function to calculate the membership of the input feature to different fuzzy rules as its corresponding additional feature. Add the additional feature and its corresponding input feature to obtain the enhanced feature. Perform convolution operation on the enhanced feature to obtain the feature map after fuzzification processing by the convolution layer as the output result of the convolution layer.
[0096] Specifically, step 2 is mainly for the operation of the convolution layer. Figure 2 Before each convolution operation, the additional features of the input features are calculated, and the additional features are added to the corresponding input features, that is, the original input features, to obtain the enhanced feature representation. After that, the convolution operation is performed on the enhanced features.
[0097] Step 3: Input the output result of the convolution layer to the pooling layer, and perform blurring in the order of channel dimension and spatial dimension to obtain the feature vector output by the pooling layer.
[0098] Specifically, step 3 is mainly for the operation of the pooling layer. Figure 2 ,The pooling layer uses the channel attention mechanism and the spatial attention mechanism in turn to blur the output results of the convolutional layer.
[0099] Exemplarily, step 3 specifically includes:
[0100] Step 3.1: Determine the preliminary enhanced feature map: In the channel dimension, use the Gaussian membership function to calculate the maximum fuzzy membership value map of each feature map in the output result of the convolution layer, and calculate the main channel weight and additional channel weight of each feature map according to the maximum fuzzy membership value map of each feature map. The corresponding product of the main channel weight and the additional channel weight is used as the channel weight, and the channel weight is multiplied by the corresponding feature map to obtain the preliminary enhanced feature map after blur processing in the channel dimension, that is, the preliminary enhanced feature map.
[0101] Step 3.2: Determine the secondary enhanced feature map: In the spatial dimension, use the Gaussian membership function to calculate the maximum fuzzy membership value map of the preliminary enhanced feature map, calculate the main spatial weight and additional spatial weight of the preliminary enhanced feature map based on the maximum fuzzy membership value map of the preliminary enhanced feature map, take the corresponding product of the main spatial weight and the additional spatial weight as the spatial weight, multiply the spatial weight with the corresponding preliminary enhanced feature map, and obtain the secondary enhanced feature map after fuzzy processing in the spatial dimension, that is, the secondary enhanced feature map.
[0102] Step 3.3: Pool the secondary enhanced feature map. For the pooling window, set the downsampling method to the summation operation, calculate the sum of all eigenvalues in the pooling window as the pooled feature, and obtain the feature map after the pooling layer blur processing, that is, the feature vector output by the pooling layer.
[0103] Step 4: Input the feature vector into the fully connected layer and use the softmax function to classify the output of the fully connected layer.
[0104] Specifically, combined with Figure 2 , Step 4 mainly inputs the feature vector obtained after processing by the convolution layer and the pooling layer into the fully connected layer, and uses the softmax function to classify the input feature vector.
[0105] Exemplarily, in step 2, using a Gaussian membership function to calculate the membership of the input feature to different fuzzy rules as the corresponding additional features, including:
[0106] Calculate the membership of the input features: The input features of each convolution Mapped to the corresponding fuzzy sets Each fuzzy set corresponds to a membership function, which uses the Gaussian membership function to generate the corresponding center and bandwidth , used to calculate the input features The membership degree of the input features is calculated according to the calculation formula of the Gaussian membership function. The corresponding membership degree, where the calculation formula of the Gaussian membership function is expressed as:
[0107] .
[0108] in, Represents the tth input feature of the kth fuzzy rule The corresponding fuzzy set. Represents input features In fuzzy sets The degree of membership on . Represents the tth input feature of the kth fuzzy rule The center of the Gaussian function can be generated by randomly selecting one of the center points {0, 0.25, 0.5, 0.75, 1.0} of the data in its feature map. Represents the tth input feature of the kth fuzzy rule The bandwidth of the Gaussian function can be uniformly set to 1. The number of fuzzy rules and the number of input features are both d, which is manifested in that the value range of k and the value range of t are both 1 to d.
[0109] Calculate the activation degree of the fuzzy rules: According to the calculation formula of the rule activation degree, calculate the activation degree of each fuzzy rule respectively, where the calculation formula of the rule activation degree is expressed as:
[0110]
[0111] in, Represents the activation degree of input feature x for the kth fuzzy rule. Input feature x includes input feature The fuzzy rule layer performs the AND operator, which is implemented using a multiplication operation.
[0112] Calculate the additional features of the input features: According to the calculation formula of the normalized membership, calculate the normalized membership of the input features for all fuzzy rules, where the calculation formula of the normalized membership is expressed as:
[0113]
[0114] in, represents the normalized membership of the input feature x to the kth fuzzy rule, and also represents the input feature Additional features. It represents the sum of activations of the input feature x for all fuzzy rules. represents the fuzzy rule number and .
[0115] For example, the feature map after the convolutional layer blurring is recorded as , where H represents the height of the convolutional layer output feature map, W represents the width of the convolutional layer output feature map, and Z represents the number of channels of the convolutional layer output feature map. Figure 3 and Figure 4 , step 3.1 determines the preliminary enhanced feature map, specifically including:
[0116] In the channel dimension, first determine the fuzzy partition number V, set the number of Gaussian membership functions equal to the fuzzy partition number, and then generate the center of the Gaussian membership function and kernel width , the membership of all features on the feature map is calculated using the Gaussian membership function, and V fuzzy membership value maps corresponding to each feature map are obtained; among which, the calculation formula of the vth Gaussian membership function corresponding to the nth feature map is expressed as:
[0117] .
[0118] in, Represents the center of the vth Gaussian function corresponding to the feature map with n channels. Represents the maximum value of all elements in the feature map with n channels. Represents the minimum value of all elements in the feature map with n channels. Indicates that the i-th row and j-th column in the feature map with n channels is located at characteristic elements. Represents the fuzzy membership value graph of the vth Gaussian function corresponding to the feature map with n channels Characteristic elements The corresponding membership degree. The kernel width of the v-th Gaussian function corresponding to the feature map with n channels can be uniformly set to 1.
[0119] At this point, V fuzzy membership values corresponding to all elements in all feature maps, that is, Z feature maps, are obtained, and each feature map has V fuzzy membership value maps.
[0120] Spatial aggregation is performed on each fuzzy membership value graph using fuzzy algebra and operators , the spatial aggregation calculation formula of the fuzzy membership value graph is expressed as:
[0121] .
[0122] in, Represents the fuzzy membership value graph of the vth Gaussian function corresponding to the nth feature map The spatial aggregation value of . H represents the height of the feature map. W represents the width of the feature map.
[0123] The fuzzy membership value map with the largest spatial aggregation value is selected as the maximum membership value map of the corresponding feature map. The fuzzy membership value map of is taken as the maximum fuzzy membership value map of the nth feature map, and the fuzzy membership value map with the largest value among the V spatial aggregation values corresponding to each feature map is called the maximum fuzzy membership value map:
[0124] .
[0125] in, It represents the maximum fuzzy membership value map among the V fuzzy membership value maps corresponding to all feature maps, and its size is . Represents the maximum fuzzy membership value map of the nth feature map (i.e., the fuzzy membership value map set The fuzzy membership value map with the largest spatial aggregation value in the middle) is . Represents the vth fuzzy membership value map corresponding to the nth feature map The spatial aggregation value of .
[0126] Calculate the main channel weights of all feature maps , whose size is , where the main channel weight of the nth feature map is expressed as:
[0127] .
[0128] in, Represents the main channel weight of the nth feature map. Represents the spatial aggregation value of the maximum membership value map of the nth feature map. represents the sum of the spatial aggregation values of the maximum membership value maps of all feature maps, that is, the sum of the spatial aggregation values of the maximum membership value maps of the 1st to Zth feature maps, where o represents the feature map number and its value range is .
[0129] Calculate the additional channel weights for all feature maps , whose size is , where the additional channel weight of the nth feature map is expressed as:
[0130] .
[0131] in, Represents the additional channel weight of the nth feature map.
[0132] Secondly, calculate the channel weights of all feature maps , whose size is , where the calculation formula for the channel weight of the nth feature map is:
[0133] .
[0134] in, Represents the channel weight of the nth feature map.
[0135] Finally, the channel weights of all feature maps The feature map output by the convolutional layer is the original input feature map Corresponding multiplication is performed to obtain the initial enhanced feature map after blurring in the channel dimension. , the calculation formula of the preliminary enhanced feature map is expressed as:
[0136] .
[0137] in, Representation feature map The channel weights of . Represents the dot product operation. Represents the preliminary enhanced feature map. Indicates the height of the preliminary enhanced feature map. Represents the width of the preliminary enhanced feature map. The number of channels of the preliminary enhanced feature map is the same as the number of channels of the feature map output by the convolutional layer, both of which are Z.
[0138] Exemplary, combined Figure 3 and Figure 5 , step 3.2 determines the secondary enhanced feature map, specifically including:
[0139] In the spatial dimension, first, the initial enhanced feature map after blurring in the channel dimension is , using the same operation as step 3.1 to calculate the maximum fuzzy membership value map, calculate each preliminary enhanced feature map The maximum fuzzy membership value map among all V fuzzy membership value maps is calculated as follows:
[0140] .
[0141] in, Represents the maximum fuzzy membership value map of the preliminary enhanced feature map, whose size is . A set of fuzzy membership value maps representing the initial enhanced feature map with n channels The maximum fuzzy membership value map in is . Represents the vth fuzzy membership value map corresponding to the preliminary enhanced feature map with n channels The spatial aggregation value of .
[0142] Calculate the main spatial weights of the preliminary enhanced feature map , whose size is , the calculation formula is expressed as:
[0143] .
[0144] Calculate additional spatial weights for the preliminary enhanced feature map , whose size is , the calculation formula is expressed as:
[0145] .
[0146] Among them, K represents the size A matrix whose elements are ,Right now . express The "same" mode is used for two-dimensional convolution operations. The maximum fuzzy membership value map representing the preliminary enhanced feature map (i.e., the fuzzy membership value map set The sum of the fuzzy membership value graph with the largest spatial aggregation value in the channel dimension is .
[0147] Secondly, calculate the spatial weights of all preliminary enhanced feature maps , whose size is , the calculation formula is expressed as:
[0148] .
[0149] Finally, the spatial weights of the feature maps are initially enhanced With the preliminary enhanced feature map Corresponding multiplication, we get the secondary enhanced feature map after fuzzy processing in the spatial dimension, that is, the secondary enhanced feature map , all Z preliminary enhanced feature maps Share the same spatial weight , the calculation formula is expressed as:
[0150] .
[0151] in, Represents the secondary enhanced feature map.
[0152] Exemplarily, the image classification method further includes:
[0153] Step 5: Train the convolutional neural network composed of convolutional layer, pooling layer, and fully connected layer, and use the gradient descent method to iteratively update the weight parameters of the convolutional neural network to further improve the performance of the convolutional neural network.
[0154] Exemplarily, step 5 specifically includes:
[0155] All samples in the training set are shuffled and randomly sorted, and the sorted samples are used to train the convolutional neural network.
[0156] During the training process, the back-propagation algorithm is used to calculate the gradients of each weight parameter in the convolutional neural network. The small batch gradient descent method and momentum gradient descent method are used to iteratively update the weight parameters of the convolutional neural network, and L2 regularization is adopted.
[0157] Among them, the activation function used in the convolution layer includes the sigmoid function or the ReLU function, and the loss function used includes the cross entropy loss function.
[0158] Compared with the prior art, the image classification method provided by the present invention has the following beneficial effects:
[0159] 1. By introducing fuzzy logic technology, it is possible to effectively handle the noise and unpredictable uncertainty contained in image data in a big data environment and reduce data uncertainty. Reducing data uncertainty is crucial to improving the accuracy and reliability of image classification tasks, especially when feature ambiguity is a serious problem.
[0160] 2. By drawing on the selectivity and regional characteristics of the human visual system, it can better simulate the superior performance of the human visual system when processing complex image data, which not only improves the accuracy of key feature extraction, but also improves the overall processing efficiency, thereby achieving better results in various image classification tasks.
[0161] 3. By introducing fuzzy logic technology, it not only helps to process uncertain information, but also enhances the system's tolerance to input data noise, making the system more robust when processing complex and noisy image data. It shows better performance than conventional CNN models in the classification tasks of different types of images such as handwriting and clothing.
[0162] In order to verify the robustness of the image classification method provided by the embodiment of the present disclosure, a comparative analysis was conducted using the test set accuracy index and the LetNet structure of the conventional CNN network architecture. See Table 1 below for specific data.
[0163] Table 1 Comparison of experimental results of the LetNet structure of the conventional CNN network architecture and the embodiments of the present disclosure
[0164]
[0165] During the experiment, the test set used the MNIST dataset and the Fashion-MNIST dataset. The MNIST dataset consists of handwritten images. Fashion-MNIST consists of clothing category images. The network parameters were set, the loss function used the sigmoid function, the number of fuzzy divisions of the pooling layer was set to 3, the momentum gradient descent method parameters were set to 0.90, the batch size used by the small batch gradient descent method was 8, the learning rate was set to 0.1, the weight decay rate was set to 0.0001, the size of the convolution kernel was 5×5, the number of convolution kernels was 6, and the pooling window size was 2×2. After the implementation of the present disclosure uses fuzzy logic technology to improve the LetNet structure of the conventional CNN network architecture, in the performance test of the MNIST dataset and the Fashion-MNIST dataset, the improved network image classification performance reached 98.41% and 87.61% accuracy on the MNIST dataset and the Fashion-MNIST dataset, respectively, and the performance shown was significantly higher than the LetNet structure of the conventional CNN network architecture. Obviously, compared with the traditional CNN architecture, the improved convolutional neural network image classification method based on fuzzy logic technology proposed in the embodiment of the present disclosure can achieve a higher performance level in the classification tasks of different types of pictures such as handwriting and clothing.
[0166] Another embodiment of the present disclosure relates to an image classification device, such as Figure 6 As shown, it includes a preprocessing module 610, a convolution module 620, a pooling module 630, and a classification module 640.
[0167] The preprocessing module 610 is used to perform a preprocessing operation on the input image, adjust the pixel value of the input image to a preset range, and obtain a corresponding normalized image.
[0168] The convolution module 620 is used to input the normalized image into the convolution layer for processing. Before each convolution operation, the membership of the input feature to different fuzzy rules is calculated using the Gaussian membership function as its corresponding additional feature. The additional feature and its corresponding input feature are added to obtain an enhanced feature. The enhanced feature is convolved to obtain a feature map after fuzzification by the convolution layer as the output result of the convolution layer.
[0169] The pooling module 630 is used to input the output result of the convolution layer into the pooling layer, perform fuzzy processing in the order of channel dimension and spatial dimension, and obtain the feature vector output by the pooling layer.
[0170] The classification module 640 is used to input the feature vector into the fully connected layer and classify the output result of the fully connected layer using the softmax function.
[0171] Exemplarily, the image classification device further includes an iteration module. The iteration module is used to train a convolutional neural network composed of a convolutional layer, a pooling layer, and a fully connected layer, and to iteratively update the weight parameters of the convolutional neural network using a gradient descent method.
[0172] The specific implementation method of the image classification device provided in the embodiment of the present disclosure can refer to the image classification method provided in the embodiment of the present disclosure, which will not be repeated here.
[0173] Compared with the prior art, the image classification device provided by the present disclosure has the following beneficial effects:
[0174] 1. By introducing fuzzy logic technology, it is possible to effectively handle the noise and unpredictable uncertainty contained in image data in a big data environment and reduce data uncertainty. Reducing data uncertainty is crucial to improving the accuracy and reliability of image classification tasks, especially when feature ambiguity is a serious problem.
[0175] 2. By drawing on the selectivity and regional characteristics of the human visual system, it can better simulate the superior performance of the human visual system when processing complex image data, which not only improves the accuracy of key feature extraction, but also improves the overall processing efficiency, thereby achieving better results in various image classification tasks.
[0176] 3. By introducing fuzzy logic technology, it not only helps to process uncertain information, but also enhances the system's tolerance to input data noise, making the system more robust when processing complex and noisy image data. It shows better performance than conventional CNN models in the classification tasks of different types of images such as handwriting and clothing.
[0177] Another embodiment of the present disclosure relates to an electronic device, such as Figure 7 As shown, including:
[0178] at least one processor 701; and,
[0179] A memory 702 is communicatively connected to at least one processor 701; wherein,
[0180] The memory 702 stores instructions that can be executed by at least one processor 701 . The instructions are executed by at least one processor 701 so that the at least one processor 701 can execute the image classification method described in the above embodiment.
[0181] Among them, the memory and the processor are connected in a bus manner, and the bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. The data processed by the processor is transmitted on a wireless medium via an antenna, and further, the antenna also receives data and transmits the data to the processor.
[0182] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory can be used to store data used by the processor when performing operations.
[0183] Another embodiment of the present disclosure relates to a computer-readable storage medium storing a computer program, which implements the image classification method described in the above embodiment when executed by a processor.
[0184] That is, those skilled in the art can understand that all or part of the steps in the method described in the above embodiments can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including a number of instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0185] Those skilled in the art will appreciate that the above-mentioned embodiments are specific embodiments for implementing the present disclosure, and in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present disclosure.
Claims
1. An image classification method, characterized in that: The image classification method comprises: Performing a preprocessing operation on the input image to adjust the pixel value of the input image to a preset range to obtain a corresponding normalized image; The normalized image is input to a convolution layer for processing. Before each convolution operation, a Gaussian membership function is used to calculate the membership of the input feature to different fuzzy rules as the corresponding additional feature, the additional feature and the corresponding input feature are added to obtain an enhanced feature, and a convolution operation is performed on the enhanced feature to obtain a feature map after fuzzification processing by the convolution layer as the output result of the convolution layer; The output result of the convolution layer is input into the pooling layer, and fuzzification is performed in the order of channel dimension and spatial dimension to obtain the feature vector output by the pooling layer; Input the feature vector into a fully connected layer, and classify the output result of the fully connected layer using a softmax function; The output result of the convolution layer is input to the pooling layer, and fuzzification is performed in the order of channel dimension and spatial dimension to obtain the feature vector output by the pooling layer, including: Determine a preliminary enhanced feature map: in the channel dimension, use a Gaussian membership function to calculate the maximum fuzzy membership value map of each feature map in the output result of the convolution layer, calculate the main channel weight and the additional channel weight of each feature map according to the maximum fuzzy membership value map of each feature map, use the corresponding product of the main channel weight and the additional channel weight as the channel weight, multiply the channel weight by the corresponding feature map, and obtain a preliminary enhanced feature map after blur processing in the channel dimension; Determine the secondary enhanced feature map: in the spatial dimension, use the Gaussian membership function to calculate the maximum fuzzy membership value map of the preliminary enhanced feature map, calculate the main spatial weight and the additional spatial weight of the preliminary enhanced feature map according to the maximum fuzzy membership value map of the preliminary enhanced feature map, use the corresponding product of the main spatial weight and the additional spatial weight as the spatial weight, multiply the spatial weight by the corresponding preliminary enhanced feature map, and obtain the secondary enhanced feature map after fuzzy processing in the spatial dimension; The secondary enhanced feature map is pooled. For the pooling window, the downsampling mode is set to a summation operation, and the sum of all feature values in the pooling window is calculated as the feature after pooling to obtain the feature vector output by the pooling layer.
2. The image classification method according to claim 1, characterized in that: The method of calculating the membership of the input feature to different fuzzy rules using the Gaussian membership function as the corresponding additional feature includes: Calculate the membership of the input feature: The input feature of each convolution Mapped to the corresponding fuzzy sets In the figure, each fuzzy set corresponds to a Gaussian membership function, generating the corresponding center and bandwidth. According to the calculation formula of the Gaussian membership function, the input features are calculated respectively. The corresponding membership degree, wherein the calculation formula of the Gaussian membership function is expressed as: ; in, Represents the tth input feature of the kth fuzzy rule The corresponding fuzzy set; Represents input features In fuzzy sets The degree of membership on ; Represents the tth input feature of the kth fuzzy rule The center of the Gaussian function; Represents the tth input feature of the kth fuzzy rule The bandwidth of the Gaussian function is d; the number of fuzzy rules and the number of input features are both d; Calculate the activation degree of the fuzzy rule: Calculate the activation degree of each fuzzy rule according to the calculation formula of the rule activation degree, wherein the calculation formula of the rule activation degree is expressed as: ; in, Represents the activation degree of input feature x for the kth fuzzy rule; input feature x includes input feature ;The fuzzy rule layer performs the AND operator, which is implemented using the multiplication operation; Calculate the additional features of the input features: Calculate the normalized membership of the input features for all fuzzy rules according to the calculation formula of the normalized membership, where the calculation formula of the normalized membership is expressed as: ; in, represents the normalized membership of the input feature x to the kth fuzzy rule, and also represents the input feature Additional features of Represents the sum of activations of the input feature x for all fuzzy rules; represents the fuzzy rule number and .
3. The image classification method according to claim 1, characterized in that: The determining of the preliminary enhanced feature map specifically includes: In the channel dimension, determine the fuzzy partition number V, set the number of Gaussian membership functions equal to the fuzzy partition number, generate the center and kernel width of the Gaussian membership function, use the Gaussian membership function to calculate the membership of all features on the feature map, and obtain V fuzzy membership value maps corresponding to each feature map; wherein, the calculation formula of the vth Gaussian membership function corresponding to the nth feature map is expressed as: ; Among them, Z represents the number of channels of the feature map; Represents the center of the vth Gaussian function corresponding to the feature map with n channels; Represents the maximum value of all elements in the feature map with n channels; Represents the minimum value of all elements in the feature map with n channels; Indicates that the position in the feature map with the number of channels n is Characteristic elements of Represents the fuzzy membership value graph of the vth Gaussian function corresponding to the feature map with n channels Characteristic elements The corresponding membership degree; Indicates the kernel width of the v-th Gaussian function corresponding to the feature map with n channels; Spatial aggregation is performed on each fuzzy membership value graph using fuzzy algebra and operators , the spatial aggregation calculation formula of the fuzzy membership value graph is expressed as: ; in, Represents the fuzzy membership value graph of the vth Gaussian function corresponding to the nth feature map The spatial aggregation value of; H represents the height of the feature map; W represents the width of the feature map; Select the fuzzy membership value map with the largest spatial aggregation value as the maximum membership value map of the corresponding feature map; Calculate the main channel weights of all feature maps, where the main channel weight of the nth feature map is expressed as: ; in, Represents the main channel weight of the nth feature map; Represents the spatial aggregation value of the maximum membership value map of the nth feature map; Represents the sum of the spatial aggregation values of the maximum membership value maps of all feature maps; Calculate the additional channel weights of all feature maps, where the additional channel weight of the nth feature map is expressed as: ; in, Represents the additional channel weight of the nth feature map; Calculate the channel weights of all feature maps, where the calculation formula for the channel weight of the nth feature map is: ; in, Represents the channel weight of the nth feature map; The channel weights of all feature maps are multiplied by the feature map output by the convolutional layer to obtain the preliminary enhanced feature map. The calculation formula of the preliminary enhanced feature map is expressed as: ; in, Representation feature map The channel weights of Represents dot multiplication operation; Represents the preliminary enhanced feature map.
4. The image classification method according to claim 3, characterized in that: The determining of the secondary enhanced feature map specifically includes: In the spatial dimension, the maximum fuzzy membership value map of the preliminary enhanced feature map is calculated, and the calculation formula is: ; in, Represents the maximum fuzzy membership value map of the preliminary enhanced feature map, whose size is ; The maximum fuzzy membership value map representing the initial enhanced feature map with n channels is ; Represents the vth fuzzy membership value map corresponding to the preliminary enhanced feature map with n channels The spatial aggregation value of Calculate the main spatial weights of the preliminary enhanced feature map , whose size is , the calculation formula is expressed as: ; Calculate additional spatial weights for the preliminary enhanced feature map , whose size is , the calculation formula is expressed as: ; Among them, K represents the size A matrix whose elements are ; express The "same" mode is used for two-dimensional convolution operations; Represents the sum of the maximum fuzzy membership value map of the preliminary enhanced feature map in the channel dimension; Calculate the spatial weights of the preliminary enhanced feature map , whose size is , the calculation formula is expressed as: ; The spatial weight of the preliminary enhanced feature map is multiplied by the preliminary enhanced feature map to obtain the secondary enhanced feature map. The calculation formula is expressed as: ; in, Represents the secondary enhanced feature map.
5. The image classification method according to any one of claims 1 to 4, characterized in that: The image classification method further comprises: The convolutional neural network composed of the convolutional layer, the pooling layer and the fully connected layer is trained, and the weight parameters of the convolutional neural network are updated and iterated using the gradient descent method.
6. The image classification method according to claim 5, characterized in that: The training of the convolutional neural network composed of the convolutional layer, the pooling layer, and the fully connected layer, and updating and iterating the weight parameters of the convolutional neural network using a gradient descent method, includes: All samples of the training set are shuffled and randomly sorted, and the sorted samples are used to train the convolutional neural network; During the training process, the weight parameters of the convolutional neural network are iteratively updated by using a mini-batch gradient descent method and a momentum gradient descent method; Among them, the activation function adopted by the convolution layer includes a sigmoid function or a ReLU function, and the loss function adopted includes a cross entropy loss function.
7. An image classification device, characterized in that: The image classification device comprises: A preprocessing module, used to perform a preprocessing operation on the input image, adjust the pixel value of the input image to a preset range, and obtain a corresponding normalized image; A convolution module, used for inputting the normalized image into a convolution layer for processing, and before each convolution operation, using a Gaussian membership function to calculate the membership of the input feature to different fuzzy rules as the corresponding additional feature, adding the additional feature and the corresponding input feature to obtain an enhanced feature, performing a convolution operation on the enhanced feature, and obtaining a feature map after fuzzification processing by the convolution layer as the output result of the convolution layer; A pooling module, used to input the output result of the convolution layer into the pooling layer, perform fuzzy processing in the order of channel dimension and spatial dimension, and obtain a feature vector output by the pooling layer; A classification module, used for inputting the feature vector into a fully connected layer and classifying the output result of the fully connected layer using a softmax function; The output result of the convolution layer is input to the pooling layer, and fuzzification is performed in the order of channel dimension and spatial dimension to obtain the feature vector output by the pooling layer, including: Determine a preliminary enhanced feature map: in the channel dimension, use a Gaussian membership function to calculate the maximum fuzzy membership value map of each feature map in the output result of the convolution layer, calculate the main channel weight and the additional channel weight of each feature map according to the maximum fuzzy membership value map of each feature map, use the corresponding product of the main channel weight and the additional channel weight as the channel weight, multiply the channel weight by the corresponding feature map, and obtain a preliminary enhanced feature map after blur processing in the channel dimension; Determine the secondary enhanced feature map: in the spatial dimension, use the Gaussian membership function to calculate the maximum fuzzy membership value map of the preliminary enhanced feature map, calculate the main spatial weight and the additional spatial weight of the preliminary enhanced feature map according to the maximum fuzzy membership value map of the preliminary enhanced feature map, use the corresponding product of the main spatial weight and the additional spatial weight as the spatial weight, multiply the spatial weight by the corresponding preliminary enhanced feature map, and obtain the secondary enhanced feature map after fuzzy processing in the spatial dimension; The secondary enhanced feature map is pooled. For the pooling window, the downsampling mode is set to a summation operation, and the sum of all feature values in the pooling window is calculated as the feature after pooling to obtain the feature vector output by the pooling layer.
8. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the image classification method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the image classification method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Medical data-oriented deep convolutional fuzzy neural network and training method thereof
CN116384450A
Image classification method and device, electronic equipment and storage medium
CN118587512A