Fruit image classification method based on shuttle-shaped dynamic neuron model
Through the fruit image classification method based on the Shunshi dynamic neuron model, the problems of large in-class differences, complex background interference, insufficient data volume and low computing efficiency in the prior art are solved, and the effects of high precision, robustness and real-time classification are achieved.
Patent Information
- Application Number
- CN202510019919.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-16
AI Technical Summary
The existing fruit image classification technology faces problems such as large in-class differences, complex background interference, insufficient data volume and low computing efficiency, and it is difficult to achieve high-precision and real-time classification in practical applications.
The fruit image classification method based on the Shunshi dynamic neuron model is adopted, and efficient feature extraction and classification of fruit images is achieved through steps such as image block division, multi-head self-attention mechanism, residual connection and normalization, feedforward network processing and classifier classification, combined with the synaptic layer, dendritic layer and membrane layer dynamic adjustment mechanism of the Shunshi dynamic neuron model.
It significantly improves the accuracy and robustness of fruit image classification, enhances the generalization ability in small sample scenarios, improves the biological interpretability and computational efficiency of the model, and is suitable for real-time classification tasks.
Smart Images

Figure CN120014628A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of classification decision of fruit images, and in particular to a fruit image classification method based on a fusiform dynamic neuron model. Background Art
[0002] Fruit image classification is an important part of fruit product sales and management. Its purpose is to automatically identify the category of fruit through machine learning or computer vision technology, thereby improving the efficiency of fruit sorting, packaging, storage and sales. Traditional fruit classification relies on manual operation, which is not only inefficient, but also prone to inconsistency in classification results due to human errors. In the context of modern agricultural production and intelligent logistics, automated and high-precision fruit classification methods have become an urgent need in the industry. The core of fruit image classification lies in how to effectively extract the diverse features of fruit such as color, shape, size and texture, and how to build an efficient classification model to accurately identify the extracted features. However, there are many kinds of fruits in nature, and both intra-class differences (such as differences in maturity, color and shape of fruits of the same type) and inter-class similarities (such as different fruits with similar colors or shapes) put forward higher requirements for classification technology. In addition, complex backgrounds (such as lighting changes and occlusions under different shooting conditions) and limited data (such as small sample problems) also pose great challenges to the accuracy, generalization and robustness of fruit image classification.
[0003] In response to the challenges in the field of fruit image classification, domestic and foreign scholars have proposed a variety of technologies and methods, mainly including classification technologies based on traditional image processing methods and classification models based on deep learning. Traditional fruit image classification methods are usually based on manually designed feature extraction technologies and machine learning classifiers. For example, some methods extract the texture features of fruit images through color complete local binary patterns (CLBP) and use the nearest neighbor classifier (KNN) to achieve the final classification. This method has a certain effect in texture feature extraction, but it is not robust enough for complex backgrounds. Some methods use Gaussian filtering to smooth fruit images to varying degrees, and then combine simulated annealing particle swarm algorithm (PSO-SA) for classification. This method relies on the adjustment of filter parameters, is sensitive to changes in illumination, and has limited practical application effects.
[0004] With the development of deep learning, convolutional neural network (CNN) has become the mainstream method for fruit image classification, and has derived a number of improved models and strategies. For example:
[0005] The invention patent with the patent publication number CN118154967A proposes a fruit classification method based on the MobileViT network, which uses image enhancement technology (such as random rotation, translation, and noise injection) to expand sample data and remove noise through bilateral filtering to improve the classification accuracy of the model. This method has certain advantages in fruit surface defect detection, but the network structure is relatively complex and the reasoning efficiency is relatively low.
[0006] The invention patent with patent publication number CN114818931B proposes a fruit image classification method based on small sample meta-learning. It solves the classification problem of small sample data sets through the MAML meta-learning framework and the DenseNet-121 network combined with the feature pyramid network (FPN), but requires a complex internal and external loop algorithm and has a large training overhead.
[0007] The invention patent with the patent publication number CN114881155B proposes a classification model based on deep transfer learning, which realizes feature extraction by freezing the low-level network parameters, and optimizes the high-level network parameters, and uses the transfer model (TL-VGG16, TL-InceptionV3 and TL-ResNet50) to improve the classification effect. However, this method still faces the problem of overfitting in the case of small data sets and is not adaptable enough to new categories of fruits.
[0008] Although the above technologies and patents have improved the accuracy and efficiency of fruit image classification to a certain extent, there are still the following deficiencies in practical applications:
[0009] Intra-class differences and inter-class similarities: For the same type of fruit, differences in maturity, size, and color can significantly increase the difficulty of classification; similar characteristics of different types of fruit (such as apples and pears) can easily lead to misclassification.
[0010] Interference from complex background: In natural collection scenes, fruits are often accompanied by complex backgrounds (such as leaves, branches, or other debris). These background information may interfere with the extraction of classification features and reduce the classification accuracy.
[0011] Insufficient data and small sample problems: The amount of image data for some rare fruits is limited, and traditional deep learning methods rely on large-scale data, which limits their performance in small sample scenarios.
[0012] Insufficient computational efficiency and generalization ability: Although some methods (such as transfer learning and meta-learning) have improved classification performance, their models are complex and inference efficiency is low, making it difficult to meet the needs of real-time classification. Some methods are too dependent on specific training data sets and lack sufficient generalization ability, making it difficult to adapt to various fruit image classification tasks in different scenarios. Summary of the invention
[0013] In order to solve the problems faced by fruit image classification in the prior art, such as large intra-class differences, complex background interference, insufficient data volume, and low computational efficiency, the present invention proposes a fruit image classification method based on a fusiform dynamic neuron model. The technical solution of the present invention comprises the following steps:
[0014] S1: Image block division, specifically including the following steps:
[0015] S1-1: Input 3D fruit image Among them, H is the image height, W is the image width, and C is the number of channels;
[0016] S1-2: Divide the input image into blocks of fixed size P×P to obtain a set of image blocks X patch ={x1,x2,...,x N}, where each image block N is the number of image blocks;
[0017] S1-3: Embed each image block into a high-dimensional feature space through linear projection and calculate the embedded representation Z, which is calculated as follows:
[0018] Z=[x1W E ,x2W E ,...,x N W E ]
[0019] in, Embedding matrix, D is the embedding dimension, is the embedding representation matrix;
[0020] S2: Multi-head self-attention mechanism, input the embedded representation Z of the image block, and apply the multi-head self-attention mechanism to capture the long-distance dependencies between image blocks. Specifically, it includes the following steps:
[0021] S2-1: Generate query matrix Q, key matrix K and value matrix V according to the embedded representation Z. The calculation formulas are:
[0022] Q=ZW Q ,K=ZW K ,V=ZW V
[0023] Among them, W Q , W K , are the projection matrices for query, key, and value respectively;
[0024] S2-2: Calculate the attention weight matrix A, which is calculated as follows:
[0025]
[0026] Among them, D embed The dimensions of the key matrix and query matrix are used for scaling to avoid excessive inner product values of high-dimensional data;
[0027] S2-3: Calculate the attention output Z′ based on the weight matrix, and the calculation formula is:
[0028] Z′=A⊙V
[0029] Among them, ⊙ represents the matrix multiplication operation;
[0030] S2-4: Perform the above processing on h different attention heads respectively, and finally concatenate the outputs of multiple heads into the multi-head attention result, whose calculation formula is:
[0031] Z′ multi_head =[Z1,Z2,...,Z h ]W O
[0032] Among them, Z i is the attention result of the i-th head, is the output weight matrix, h represents the number of heads;
[0033] S3: Residual connection and normalization: multi-head attention result Z′ multi_head Perform residual connection with the embedded representation Z to obtain the residual result: Z residual =Z+Z′ multi_head ; Normalize the residual result and calculate the normalized result Z″, the calculation formula is:
[0034]
[0035] Among them, μ is the mean and σ is the standard deviation;
[0036] S4: Feedforward network processing: Input the normalized embedding representation Z″ and perform nonlinear transformation through the feedforward neural network, which specifically includes the following steps:
[0037] S4-1: Use the fully connected layer and activation function to calculate the feature transformation, and the calculation formula is:
[0038] f(Z″)=W2(max(0,Z″W1+b1))+b2
[0039] in, is the weight matrix of the feedforward network, b1, b2 are bias terms, D hidden is the hidden layer dimension;
[0040] S4-2: Perform a residual connection between the feedforward network output and the input Z″ to obtain the final feedforward network processing result, which is calculated as follows:
[0041]
[0042] Among them, μ′ is the mean after residual connection, σ′ is the standard deviation;
[0043] S5: Classifier classification: Input feature representation Z after being processed by the feedforward network out , the fruit image classification is completed through the classifier; the classifier is based on Z out The feature distribution of the fruit is output, and the prediction result of the corresponding fruit category is output.
[0044] As a preferred technical solution of the present invention, the classifier in step S5 completes the classification of data through the following steps:
[0045] S5-1: Maximum average pooling operation: The output data Z processed by the multi-head self-attention mechanism and the feedforward network is out Perform maximum average pooling processing, the calculation formula is:
[0046]
[0047] Where n represents the nth sample, c represents the number of channels, h and w represent the height and width indices respectively, h′ and w′ represent the dimensions after pooling, and k h and k w is the height and width of the pooling kernel;
[0048] S5-2: Synaptic layer processing: The data after maximum average pooling is transmitted to the synaptic layer of the fusiform dynamic neuron model (FV-DNM); the synaptic layer calculates the synaptic response value S according to the input ij , and its calculation formula is:
[0049]
[0050] Among them, k is the distance parameter; w ij is the synaptic weight, which is randomly generated by the normal distribution function; tanh represents the hyperbolic tangent function, and its calculation formula is:
[0051] S5-3: Dendritic layer summation: The output of the synaptic layer is transmitted to the dendritic layer, and the dendritic layer calculates the synaptic response value S on a single dendrite. ij The calculation formula is:
[0052]
[0053] Where N is the number of synapses, D j represents the sum of the j-th dendritic branch;
[0054] S5-4: Membrane layer summation: The output of the dendritic layer is transmitted to the membrane layer, and the membrane layer sums the output of all dendritic branches. The calculation formula is:
[0055]
[0056] Where M is the number of dendritic branches, and E is the output of the membrane layer;
[0057] S5-5: Somatic layer regulation: The output of the membrane layer is transmitted to the somatic layer, which makes adjustments based on the historical activity data of the synapses and finally calculates the output O. The calculation formula is:
[0058]
[0059] Among them, k v is a learnable scaling parameter, mean j (w ij ) is the mean value calculated along the column, exp represents the natural exponential function, and σ is the activation function;
[0060] S5-6: Axon layer output: The somatic layer's regulatory output O is transmitted to the axon layer together with the membrane layer's output E. The axon layer summarizes and generates the final classification result, which is calculated as follows:
[0061] T=E⊙O
[0062] Among them, T is the final output of the classifier, which is used for fruit image classification.
[0063] As a preferred technical solution of the present invention, the fusiform dynamic neuron model described in step S5-2 has the following system structure:
[0064] T1: Input layer: The input layer receives preprocessed fruit images. The input data shape is (batch_size, 3, 64, 64), where batch_size is the batch size, 3 represents the number of RGB channels, and 64×64 represents the height and width of the image. The preprocessing steps include image normalization and standardization to eliminate the effects of image brightness and contrast.
[0065] T2: Convolutional layer and pooling layer, including the following features:
[0066] The convolution layer is composed of a number of stacked convolution units, each of which extracts features of a local area of the input feature map through a convolution kernel;
[0067] The output of the convolutional layer is connected to the maximum pooling layer and the average pooling layer. The pooling operation is used to reduce the resolution of the feature map and compress the data size.
[0068] The outputs of the maximum pooling layer and the average pooling layer are fused through the maximum average pooling mechanism to uniformly express local features and global features;
[0069] The fused feature map output by the convolutional layer and the pooling layer is directly connected to the fusiform dynamic neuron layer.
[0070] T3: Dynamic neuronal layer (DNM), including the following features:
[0071] The DNM layer dynamically adjusts the synaptic weights according to the input feature map to simulate the dynamic response process of biological spindle neurons;
[0072] The DNM layer includes a synaptic sublayer, a dendrite sublayer and a membrane sublayer:
[0073] The output feature representation of the fusiform dynamic neuron layer is directly connected to the classifier.
[0074] As a technical preferred solution of the present invention, the synaptic sublayer generates dynamic synaptic connection weights by weight calculation based on input feature distribution;
[0075] The dendritic sublayer performs weighted summation on multi-synaptic response values on a single dendrite to generate a dendritic output;
[0076] The membrane sublayer summarizes the outputs of all dendritic branches to obtain a high-order feature representation after feature fusion;
[0077] The fusiform dynamic neuron layer is directly connected to the outputs of the convolutional layer and the pooling layer, and receives the fused feature map as input;
[0078] The output of the fusiform dynamic neuron layer is directly used as the input of the classifier to generate the final classification result.
[0079] As a preferred technical solution of the present invention, the pretreatment step described in structure T1 includes the following steps:
[0080] Resizing: resizing the original fruit image to a fixed-size image, uniformly resizing to 64×64 pixels; wherein, the resizing adopts a bilinear interpolation method or a nearest neighbor interpolation method;
[0081] Data augmentation: Use data augmentation techniques to randomly transform image samples to increase the diversity and richness of the data set. Data augmentation methods include random rotation, random flipping, random cropping, and color adjustment.
[0082] Normalization: Normalize the pixel values of the fruit image and map the pixel values from the interval [0,255] to [0,1]. The calculation formula is:
[0083]
[0084] Among them, μ is the pixel mean and σ is the pixel standard deviation.
[0085] As a technical preferred solution of the present invention, the fruit image classification method further includes an actual classification process based on a fusiform dynamic neuron model, and the classification process includes the following steps:
[0086] Feature extraction: The input fruit image is predicted by the trained fusiform dynamic neuron model; the image first passes through the convolutional layer and the pooling layer to extract low-level and mid-level features from the input data; the extracted mid-level features are further processed by the self-attention mechanism to generate a high-level feature representation that contains global context information and dynamic responses;
[0087] Category prediction: The high-level feature representation is input to the fusiform dynamic neuron model (FV-DNM), which is gradually processed through the synaptic layer, dendrite layer, and membrane layer in its fusiform neuron layer to achieve the mapping of fruit categories; the axon layer of the fusiform dynamic neuron model converts the mapping results into the category prediction probability of the output layer, generating the confidence value corresponding to each category;
[0088] Top-N prediction: Combined with the prediction probability output by the model to support the Top-N classification results, the model will output the N categories with the highest confidence.
[0089] Compared with the prior art, the present invention has the following beneficial effects:
[0090] Improve classification accuracy and robustness: The present invention introduces the fusiform dynamic neuron model (FV-DNM), combined with the dynamic synaptic adjustment mechanism and nonlinear transformation ability of fusiform neurons, to achieve efficient learning of complex features of fruit images; the application of the self-attention mechanism can capture the global and local features in the image, especially for the problems of large intra-class differences and similarities between classes; the maximum average pooling mechanism combined with the dynamic weight adjustment method effectively eliminates the interference of complex background and improves the robustness and accuracy of classification.
[0091] Enhanced generalization ability in small sample scenarios: Through dynamic weight updates and biologically inspired synaptic mechanisms, the present invention has excellent learning ability under small sample data sets; the dynamic adjustment mechanism of fusiform neurons adapts to the situation of insufficient data volume, greatly improving the generalization performance of the model in multiple scenarios and different data distributions.
[0092] Improve the biological interpretability of the model: The FV-DNM model is based on the characteristics of biological neurons and adopts a structure similar to synaptic dynamic adjustment and dendritic information integration, making the model biologically interpretable; this mechanism is not only suitable for image classification tasks, but also can provide inspiration for research in neuroscience and other biological fields.
[0093] Improve computing efficiency and reduce training difficulty: The modular design of the fusiform dynamic neuron model makes the model more lightweight, reducing the number of parameters and computing costs; the dynamic adjustment and feature aggregation mechanism are combined in the classification process to significantly improve the inference speed, making it suitable for real-time classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] Figure 1 It is a flow chart of a fruit image classification method based on a fusiform dynamic neuron model of the present invention;
[0095] Figure 2 It is a schematic diagram of the model structure of FV-DNM according to an embodiment of the present invention;
[0096] Figure 3 1 is a comparison chart of 5 model index values provided in the embodiment of the present invention;
[0097] Figure 4 It is a comparison chart of 2 values of 5 model indicators provided in the embodiments of the present invention. DETAILED DESCRIPTION
[0098] The present invention is further described below in conjunction with the accompanying drawings and examples. However, the present invention can be implemented in many different ways and should not be construed as being limited to the embodiments shown; on the contrary, these embodiments provide those skilled in the art with implementation methods that meet applicable legal requirements.
[0099] Example 1: According to Figure 1 As shown, this embodiment provides a specific implementation process of a fruit image classification method based on a fusiform dynamic neuron model. The steps are as follows:
[0100] S1: Image block division, specifically including the following steps:
[0101] S1-1: Input 3D fruit image Among them, H is the image height, W is the image width, and C is the number of channels;
[0102] S1-2: Divide the input image into blocks of fixed size P×P to obtain a set of image blocks X patch ={x1,x2,...,x N}, where each image block N is the number of image blocks;
[0103] S1-3: Embed each image block into a high-dimensional feature space through linear projection and calculate the embedded representation Z, which is calculated as follows:
[0104] Z=[x1W E ,x2W E ,...,x N W E ]
[0105] in, Embedding matrix, D is the embedding dimension, is the embedding representation matrix;
[0106] S2: Multi-head self-attention mechanism, input the embedded representation Z of the image block, and apply the multi-head self-attention mechanism to capture the long-distance dependencies between image blocks. Specifically, it includes the following steps:
[0107] S2-1: Generate query matrix Q, key matrix K and value matrix V according to the embedded representation Z. The calculation formulas are:
[0108] Q=ZW Q ,K=ZW K ,V=ZW V
[0109] Among them, W Q , W K , are the projection matrices for query, key, and value respectively;
[0110] S2-2: Calculate the attention weight matrix A, which is calculated as follows:
[0111]
[0112] Among them, D embed The dimensions of the key matrix and query matrix are used for scaling to avoid excessive inner product values of high-dimensional data;
[0113] S2-3: Calculate the attention output Z′ based on the weight matrix, and the calculation formula is:
[0114] Z′=A⊙V
[0115] Among them, ⊙ represents the matrix multiplication operation;
[0116] S2-4: Perform the above processing on h different attention heads respectively, and finally concatenate the outputs of multiple heads into the multi-head attention result, whose calculation formula is:
[0117] Z′ multi_head =[Z1,Z2,...,Z h ]W O
[0118] Among them, Z i is the attention result of the i-th head, is the output weight matrix, h represents the number of heads;
[0119] S3: Residual connection and normalization: multi-head attention result Z′ multi_head Perform residual connection with the embedded representation Z to obtain the residual result: Zresidual =Z+Z′ multi_head ; Normalize the residual result and calculate the normalized result Z″, the calculation formula is:
[0120]
[0121] Among them, μ is the mean and σ is the standard deviation;
[0122] S4: Feedforward network processing: Input the normalized embedding representation Z″ and perform nonlinear transformation through the feedforward neural network, which specifically includes the following steps:
[0123] S4-1: Use the fully connected layer and activation function to calculate the feature transformation, and the calculation formula is:
[0124] f(Z″)=W2(max(0,Z″W1+b1))+b2
[0125] in, is the weight matrix of the feedforward network, b1, b2 are bias terms, D hidden is the hidden layer dimension;
[0126] S4-2: Perform a residual connection between the feedforward network output and the input Z″ to obtain the final feedforward network processing result, which is calculated as follows:
[0127]
[0128] Among them, μ′ is the mean after residual connection, σ′ is the standard deviation;
[0129] S5: Classifier classification: Input feature representation Z after being processed by the feedforward network out , the fruit image classification is completed through the classifier; the classifier is based on Z out The feature distribution of the output corresponding fruit category prediction results. The classifier completes the classification of data through the following steps:
[0130] S5-1: Maximum average pooling operation: The output data Z processed by the multi-head self-attention mechanism and the feedforward network is out Perform maximum average pooling processing, the calculation formula is:
[0131]
[0132] Where n represents the nth sample, c represents the number of channels, h and w represent the height and width indices respectively, h′ and w′ represent the dimensions after pooling, and k h and k w is the height and width of the pooling kernel;
[0133] S5-2: Synaptic layer processing: The data after maximum average pooling is transmitted to the synaptic layer of the fusiform dynamic neuron model (FV-DNM); the synaptic layer calculates the synaptic response value S according to the input ij , and its calculation formula is:
[0134]
[0135] Among them, k is the distance parameter; w ij is the synaptic weight, which is randomly generated by the normal distribution function; tanh represents the hyperbolic tangent function, and its calculation formula is:
[0136] S5-3: Dendritic layer summation: The output of the synaptic layer is transmitted to the dendritic layer, and the dendritic layer calculates the synaptic response value S on a single dendrite. ij The calculation formula is:
[0137]
[0138] Where N is the number of synapses, D j represents the sum of the j-th dendritic branch;
[0139] S5-4: Membrane layer summation: The output of the dendritic layer is transmitted to the membrane layer, and the membrane layer sums the output of all dendritic branches. The calculation formula is:
[0140]
[0141] Where M is the number of dendritic branches, and E is the output of the membrane layer;
[0142] S5-5: Somatic layer regulation: The output of the membrane layer is transmitted to the somatic layer, which makes adjustments based on the historical activity data of the synapses and finally calculates the output O. The calculation formula is:
[0143]
[0144] Among them, k v is a learnable scaling parameter, mean j (w ij ) is the mean value calculated along the column, exp represents the natural exponential function, and σ is the activation function;
[0145] S5-6: Axon layer output: The somatic layer's regulatory output O is transmitted to the axon layer together with the membrane layer's output E. The axon layer summarizes and generates the final classification result, which is calculated as follows:
[0146] T=E⊙O
[0147] Among them, T is the final output of the classifier, which is used for fruit image classification.
[0148] like Figure 2 As shown, the fusiform dynamic neuron model has the following system structure:
[0149] T1: Input layer: The input layer receives preprocessed fruit images. The input data shape is (batch_size, 3, 64, 64), where batch_size is the batch size, 3 represents the number of RGB channels, and 64×64 represents the height and width of the image. The preprocessing steps include image normalization and standardization to eliminate the effects of image brightness and contrast, specifically including the following steps:
[0150] Resizing: resizing the original fruit image to a fixed-size image, uniformly resizing to 64×64 pixels; wherein, the resizing adopts a bilinear interpolation method or a nearest neighbor interpolation method;
[0151] Data augmentation: Use data augmentation techniques to randomly transform image samples to increase the diversity and richness of the data set. Data augmentation methods include random rotation, random flipping, random cropping, and color adjustment.
[0152] Normalization: Normalize the pixel values of the fruit image and map the pixel values from the interval [0,255] to [0,1]. The calculation formula is:
[0153]
[0154] Among them, μ is the pixel mean and σ is the pixel standard deviation.
[0155] T2: Convolutional layer and pooling layer, including the following features:
[0156] The convolution layer consists of several stacked convolution units, each of which extracts features of the local area of the input feature map through the convolution kernel;
[0157] The output of the convolutional layer is connected to the maximum pooling layer and the average pooling layer. The pooling operation is used to reduce the resolution of the feature map and compress the data size.
[0158] The outputs of the maximum pooling layer and the average pooling layer are fused through the maximum average pooling mechanism to express local features and global features in a unified way;
[0159] The fused feature map output by the convolutional layer and the pooling layer is directly connected to the fusiform dynamic neuron layer.
[0160] T3: Dynamic Fusiform Neuron Layer (DNM), including the following features: The DNM layer dynamically adjusts the synaptic weights according to the input feature map to simulate the dynamic response process of biological fusiform neurons;
[0161] The DNM layer includes the synaptic sublayer, the dendritic sublayer, and the membrane sublayer:
[0162] The output feature representation of the fusiform dynamic neuron layer is directly connected to the classifier.
[0163] The fusiform dynamic neuron model further includes the following features:
[0164] The synaptic sublayer generates dynamic synaptic connection weights by weight calculation based on the input feature distribution;
[0165] The dendritic sublayer performs weighted summation of multi-synaptic response values on a single dendrite to generate the dendritic output;
[0166] The membrane sublayer summarizes the outputs of all dendritic branches to obtain a high-order feature representation after feature fusion;
[0167] The fusiform dynamic neuron layer is directly connected to the output of the convolutional layer and the pooling layer, and receives the fused feature map as input;
[0168] The output of the fusiform dynamic neuron layer is directly used as the input of the classifier to generate the final classification result.
[0169] The classification method of the present invention further includes an actual classification process based on the fusiform dynamic neuron model, and the classification process includes the following steps:
[0170] Feature extraction: The input fruit image is predicted by the trained fusiform dynamic neuron model; the image first passes through the convolutional layer and the pooling layer to extract low-level and mid-level features from the input data; the extracted mid-level features are further processed by the self-attention mechanism to generate a high-level feature representation that contains global context information and dynamic responses;
[0171] Category prediction: The high-level feature representation is input to the fusiform dynamic neuron model (FV-DNM), which is gradually processed through the synaptic layer, dendrite layer, and membrane layer in its fusiform neuron layer to achieve the mapping of fruit categories; the axon layer of the fusiform dynamic neuron model converts the mapping results into the category prediction probability of the output layer, generating the confidence value corresponding to each category;
[0172] Top-N prediction: Combined with the prediction probability output by the model to support the Top-N classification results, the model will output the N categories with the highest confidence.
[0173] Example 2: This example provides a fruit image classification method based on a fusiform dynamic neuron model. Its specific implementation process covers key links such as data preprocessing, model design and training, classification, and evaluation optimization. It aims to achieve high-precision fruit image classification by simulating the dynamic response mechanism of biological fusiform neurons.
[0174] In this embodiment, the source of fruit image data is the Fruits-100 data set, which covers a variety of different types of fruit images to ensure data diversity. In order to unify the data format, all images are first resized and normalized to 64×64 pixels to meet the requirements of the neural network input layer. In order to improve the model's adaptability to complex scenes, the RandAugment technology is used in the embodiment to enhance the data, including operations such as random rotation, flipping, cropping, and color adjustment, which significantly improves sample diversity. At the same time, by normalizing the image pixel values, the original pixel values are mapped to the standardized interval to accelerate model convergence and improve training stability. The data preprocessing stage of the present invention provides a high-quality data foundation for subsequent feature extraction and classification through resizing, data enhancement and normalization.
[0175] The core of the present invention is to introduce a neural network framework based on the fusiform dynamic neuron model (FV-DNM), which simulates the behavioral characteristics of biological neurons. In the specific implementation, the model mainly includes the following modules: Input layer: Receive the preprocessed fruit image and ensure that the input shape matches the network structure. Convolution and pooling module: Use the standard convolution layer to extract the local features of the fruit image, combine the maximum pooling and average pooling mechanisms, and realize the fusion of global and local features through the maximum average pooling operation. This module effectively improves the model's adaptability to intra-class differences and inter-class similarity problems. Fusiform dynamic neuron layer (DNM layer): As the core module of the model, the DNM layer simulates the adaptive response characteristics of neurons by dynamically adjusting the synaptic weights. This layer uses the LeCun Tanh activation function to process nonlinear relationships, and enhances the ability to capture complex features through a variance-based weight update mechanism. Axon layer output: The axon layer summarizes the high-dimensional features after processing to generate the final classification results, supporting multi-category probability prediction.
[0176] This embodiment realizes the extraction and learning of multi-level features of fruit images through modular design, especially the fusiform dynamic neuron layer significantly improves the adaptability and generalization performance of the model to complex scenes. The implementation process of model training is as follows:
[0177] In the training stage, the present invention uses the cross-entropy loss function (Cross-Entropy Loss) as the optimization target, which can effectively evaluate the performance of the model in multi-classification tasks. The optimization algorithm uses the AdamW optimizer, and its adaptive learning rate mechanism and weight decay strategy ensure the rapid convergence and stability of the model training. The training data is divided into training set, validation set and test set in a ratio of 8:1:1. During the training process, the input image is processed by each layer, the model outputs the predicted category, and the network weight is adjusted by the back propagation algorithm. In the verification stage, the model performance is evaluated by the verification set, and the hyperparameter settings, such as learning rate, weight decay coefficient, etc., are optimized according to the results to avoid overfitting. The implementation process of classification is as follows:
[0178] After completing the model training, the input fruit image is classified and predicted by the trained fusiform dynamic neuron model. The classification process is divided into two parts: feature extraction and category prediction: Feature extraction: The input image is sequentially subjected to the convolution and pooling modules to extract low-level and intermediate features, and then the high-level dynamic features are further extracted through the self-attention mechanism; Category prediction: The extracted high-level features are mapped by the fusiform dynamic neuron layer, and finally the prediction probability of each category is generated in the axon layer. The classification module of the present invention supports Top-N prediction, especially Top-1 and Top-3 classification accuracy, which can achieve efficient and accurate classification in practical applications, while providing a certain fault tolerance. The implementation process of model evaluation and optimization is as follows:
[0179] In the model testing phase, the model performance is comprehensively evaluated through the test set, and the indicators include Top-1Accuracy and Top-3 Accuracy. In addition, this embodiment supports the use of transfer learning technology, combining pre-trained deep convolutional neural networks (such as ResNet, Inception, etc.) as feature extractors with the fusiform dynamic neuron model to further improve classification performance, especially for small-scale data sets.
[0180] In order to verify the effectiveness of the proposed FV-DNM model in fruit image classification, the present invention conducted a large number of experimental analyses and comparisons. The present invention compared five models: residual neural network (ResNet), visual transformer (ViT), SwimTransformer, CoAtNet, and CaiT. The present invention selected fruits100 as the data set, which is a famous data set on the kaggle website. It contains 100 kinds of fruit pictures and has been divided into data sets, training sets, validation sets, and test sets in a ratio of 8:1:1.
[0181] In order to scientifically and accurately evaluate the classification ability of each model for fruit images, the present invention uses TOP-1Accuracy (first accuracy) and TOP-3Accuracy (top three accuracy) as evaluation indicators of the present invention. They are:
[0182] Indicator 1: Top-1 Accuracy:
[0183]
[0184] Among them, Number ofCorrectPredictions is the number of samples whose prediction results of the model for the test samples completely match the actual labels. Total Number ofPredictions is the total number of samples in the test set.
[0185] Indicator 2: top-3Accuracy:
[0186]
[0187] Among them, Number of Correct Predictions in Top-3 is the number of samples containing the true category in the top 3 highest probability categories predicted by the model for each sample. Total Number of Predictions is the total number of samples in the test set.
[0188] Table 1: Experimental results
[0189] Model TOP-1Accuracy TOP-3Accuracy ResNet 68.75 79.63 ViT 73.64 81.15 SwimTransformer 75.59 83.97 CoAtNet 75.13 85.37 Cai 75.37 86.11 FV-DNM 77.63 88.79
[0190] From Table 1, Figure 3 and Figure 4 It can be seen that in the experimental comparison results of different models, the model proposed in the present invention has achieved better prediction results in both TOP-1Accuracy and TOP-3Accuracy. Experiments have shown that the FV-DNM model proposed in the present invention can more accurately identify the category to which the fruit belongs, thereby improving the classification efficiency in the fruit market with a wide variety of categories and effectively improving the trade efficiency in the fruit product market.
[0191] The above embodiments only express several implementation methods of the present invention, and the description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A fruit image classification method based on a fusiform dynamic neuron model, characterized in that: The following steps are involved: S1: Image block division, specifically including the following steps: S1-1: Input 3D fruit image Among them, H is the image height, W is the image width, and C is the number of channels; S1-2: Divide the input image into blocks of fixed size P×P to obtain an image block set X patch ={x1,x2,...,x N }, where each image block N is the number of image blocks; S1-3: Embed each image block into a high-dimensional feature space through linear projection and calculate the embedded representation Z, which is calculated as follows: Z=[x1W E ,x2W E ,...,x N W E ] in, Embedding matrix, D is the embedding dimension, is the embedding representation matrix; S2: Multi-head self-attention mechanism, input the embedded representation Z of the image block, and apply the multi-head self-attention mechanism to capture the long-distance dependencies between image blocks. Specifically, it includes the following steps: S2-1: Generate query matrix Q, key matrix K and value matrix V according to the embedded representation Z. The calculation formulas are: Q=ZW Q ,K=ZW K ,V=ZW V Among them, W Q , W K , are the projection matrices for query, key, and value respectively; S2-2: Calculate the attention weight matrix A, which is calculated as follows: Among them, D embed The dimensions of the key matrix and query matrix are used for scaling to avoid excessive inner product values of high-dimensional data; S2-3: Calculate the attention output Z′ based on the weight matrix, and the calculation formula is: Z′=A⊙V Among them, ⊙ represents the matrix multiplication operation; S2-4: Perform the above processing on h different attention heads respectively, and finally concatenate the outputs of multiple heads into the multi-head attention result, whose calculation formula is: WITH' multi_head =[Z1,Z2,...,Z h ]IN O Among them, Z i is the attention result of the i-th head, is the output weight matrix, h represents the number of heads; S3: Residual connection and normalization: multi-head attention result Z′ multi_head Perform residual connection with the embedded representation Z to obtain the residual result: Z residual =Z+Z′ multi_head ; Normalize the residual result and calculate the normalized result Z″, the calculation formula is: Among them, μ is the mean and σ is the standard deviation; S4: Feedforward network processing: Input the normalized embedding representation Z″ and perform nonlinear transformation through the feedforward neural network, which specifically includes the following steps: S4-1: Use the fully connected layer and activation function to calculate the feature transformation, and the calculation formula is: f(Z″)=W2(max(0,Z″W1+b1))+b2 in, is the weight matrix of the feedforward network, b1, b2 are bias terms, D hidden is the hidden layer dimension; S4-2: Perform a residual connection between the feedforward network output and the input Z″ to obtain the final feedforward network processing result, which is calculated as follows: Among them, μ′ is the mean after residual connection, σ′ is the standard deviation; S5: Classifier classification: Input feature representation Z after being processed by the feedforward network out , the fruit image classification is completed through the classifier; the classifier is based on Z out The feature distribution of the fruit is output, and the prediction result of the corresponding fruit category is output.
2. A method for fruit image classification based on a fusiform dynamic neuron model according to claim 1, characterized in that: The classifier in step S5 completes the classification of the data through the following steps: S5-1: Maximum average pooling operation: The output data Z processed by the multi-head self-attention mechanism and the feedforward network is out Perform maximum average pooling processing, the calculation formula is: Where n represents the nth sample, c represents the number of channels, h and w represent the height and width indices respectively, h′ and w′ represent the dimensions after pooling, and k h and k w is the height and width of the pooling kernel; S5-2: Synaptic layer processing: The data after maximum average pooling is transmitted to the synaptic layer of the fusiform dynamic neuron model FV-DNM; the synaptic layer calculates the synaptic response value S according to the input ij , and its calculation formula is: Among them, k is the distance parameter; w ij is the synaptic weight, which is randomly generated by the normal distribution function; tanh represents the hyperbolic tangent function, and its calculation formula is: S5-3: Dendritic layer summation: The output of the synaptic layer is transmitted to the dendritic layer, and the dendritic layer calculates the synaptic response value S on a single dendrite. ij The calculation formula is: Where N is the number of synapses, D j represents the sum of the j-th dendritic branch; S5-4: Membrane layer summation: The output of the dendritic layer is transmitted to the membrane layer, and the membrane layer sums the output of all dendritic branches. The calculation formula is: Where M is the number of dendritic branches, and E is the output of the membrane layer; S5-5: Somatic layer regulation: The output of the membrane layer is transmitted to the somatic layer, which makes adjustments based on the historical activity data of the synapses and finally calculates the output O. The calculation formula is: Among them, k v is a learnable scaling parameter, mean j (w ij ) is the mean value calculated along the column, exp represents the natural exponential function, and σ is the activation function; S5-6: Axon layer output: The somatic layer's regulatory output O is transmitted to the axon layer together with the membrane layer's output E. The axon layer summarizes and generates the final classification result, which is calculated as follows: T=E⊙O Among them, T is the final output of the classifier, which is used for fruit image classification.
3. A method for fruit image classification based on a fusiform dynamic neuron model according to claim 2, characterized in that: The fusiform dynamic neuron model in step S5-2 has the following system structure: T1: Input layer: The input layer receives preprocessed fruit images. The input data shape is (batch_size, 3, 64, 64), where batch_size is the batch size, 3 represents the number of RGB channels, and 64×64 represents the height and width of the image. The preprocessing steps include image normalization and standardization to eliminate the effects of image brightness and contrast. T2: Convolutional layer and pooling layer, including the following features: The convolution layer is composed of a number of stacked convolution units, each of which extracts features of a local area of the input feature map through a convolution kernel; The output of the convolutional layer is connected to the maximum pooling layer and the average pooling layer. The pooling operation is used to reduce the resolution of the feature map and compress the data size. The outputs of the maximum pooling layer and the average pooling layer are fused through the maximum average pooling mechanism to uniformly express local features and global features; The fused feature map output by the convolutional layer and the pooling layer is directly connected to the fusiform dynamic neuron layer; T3: Fusiform dynamic neuron layer DNM, including the following features: The DNM dynamically adjusts the synaptic weights according to the input feature map to simulate the dynamic response process of biological spindle neurons; The DNM layer includes a synaptic sublayer, a dendritic sublayer and a membrane sublayer; The output feature representation of the fusiform dynamic neuron layer is directly connected to the classifier.
4. A method for fruit image classification based on a fusiform dynamic neuron model according to claim 3, characterized in that: The synaptic sublayer generates dynamic synaptic connection weights by weight calculation based on input feature distribution; The dendritic sublayer performs weighted summation on multi-synaptic response values on a single dendrite to generate a dendritic output; The membrane sublayer summarizes the outputs of all dendritic branches to obtain a high-order feature representation after feature fusion; The fusiform dynamic neuron layer is directly connected to the outputs of the convolutional layer and the pooling layer, and receives the fused feature map as input; The output of the fusiform dynamic neuron layer is directly used as the input of the classifier to generate the final classification result.
5. A method for classifying fruit images based on a fusiform dynamic neuron model according to claim 3, characterized in that: The preprocessing step in T1 includes the following steps: Resizing: resizing the original fruit image to a fixed-size image, uniformly resizing to 64×64 pixels; wherein, the resizing adopts a bilinear interpolation method or a nearest neighbor interpolation method; Data augmentation: Use data augmentation techniques to randomly transform image samples to increase the diversity and richness of the data set. Data augmentation methods include random rotation, random flipping, random cropping, and color adjustment. Normalization: Normalize the pixel values of the fruit image and map the pixel values from the interval [0,255] to [0,1]. The calculation formula is: Among them, μ is the pixel mean and σ is the pixel standard deviation.
6. The fruit image classification method based on the fusiform dynamic neuron model according to any one of claims 1 to 5, characterized in that: The method further comprises an actual classification process based on the fusiform dynamic neuron model, the classification process comprising the following steps: Feature extraction: The input fruit image is predicted by the trained fusiform dynamic neuron model; the image first passes through the convolutional layer and the pooling layer to extract low-level and mid-level features from the input data; the extracted mid-level features are further processed by the self-attention mechanism to generate a high-level feature representation that contains global context information and dynamic responses; Category prediction: The high-level feature representation is input to the fusiform dynamic neuron model FV-DNM, which is gradually processed through the synaptic layer, dendrite layer and membrane layer in its fusiform neuron layer to achieve the mapping of fruit categories; the axon layer of the fusiform dynamic neuron model converts the mapping results into the category prediction probability of the output layer, and generates the confidence value corresponding to each category; Top-N prediction: Combined with the prediction probability output by the model to support the Top-N classification results, the model will output the N categories with the highest confidence.
Citation Information
Patent Citations
A fruit image classification method based on small sample meta-learning
CN114818931B
Fruit image classification method based on deep transfer learning
CN114881155B
Fruit image classification method based on deep learning
CN118154967A