Hulless oat seed fine-grained image classification method based on improved ResNet18
By improving the ResNet18 model, introducing a multi-scale feature pyramid enhancement module, a detail enhancement network, and a feature complementary network, and performing adaptive collaborative fusion, the problem of low accuracy in naked oat seed classification was solved, and efficient naked oat seed image classification was achieved.
Patent Information
- Application Number
- CN202510699434.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing oat seed classification methods mainly rely on manual inspection, which is time-consuming and labor-intensive. In addition, the accuracy of machine vision-based methods is low and it is difficult to meet actual needs.
An improved ResNet18 model was adopted, and the multi-scale feature pyramid enhancement module (MSFPE) was introduced to construct the detail reinforcement network (DRN) and feature complementary network (FCN). The features were weightedly fused through the adaptive collaborative fusion module (ACFM) to improve the model's classification ability for fine-grained images of naked oat seeds.
The accuracy and robustness of naked oat seed image classification were significantly improved, with the accuracy rate increased from 78.03% to 88.63%, meeting the needs of practical applications.
Smart Images

Figure CN120599355A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of oat seed classification, and in particular to an oat seed fine-grained image classification method based on an improved ResNet18. Background Art
[0002] Naked oats are a herbaceous plant in the genus Avena in the Poaceae family. With the booming seed market, competition is intensifying. Reduced crop yields due to mixed seed varieties are a common occurrence, and seed purity is a growing concern. Traditional seed purity testing methods rely primarily on manual destructive inspection, which is time-consuming and labor-intensive, making it difficult to scale up in practical applications. Currently, a growing number of non-destructive testing technologies are gaining traction, including machine vision, near-infrared spectroscopy, and hyperspectral imaging. However, existing machine vision-based classification methods have relatively low accuracy, making them unsuitable for widespread use. Summary of the Invention
[0003] The technical problem to be solved by the present invention is how to provide a method for accurately classifying fine-grained images of naked oat seeds.
[0004] To solve the above technical problems, the technical solution adopted by the present invention is: a fine-grained image classification method of naked oat seeds based on an improved ResNet18, comprising the following steps:
[0005] Image acquisition: An optical image data acquisition system is used to collect fine-grained images of naked oat seeds. After processing, several single-grain naked oat seed images are obtained to form a data set.
[0006] Dataset processing: Divide the dataset into training set and test set according to the set ratio, and preprocess the training set and test set respectively;
[0007] Build an improved ResNet18 image classification model: Based on the backbone network of the ResNet18 model, first, introduce the fused multi-scale feature pyramid enhancement module (MSFPE) afterwards. Secondly, build a DRFC network consisting of a detail reinforcement network (DRN) and a feature complementation network (FCN). Finally, use the adaptive collaborative fusion module (ACFM) to perform adaptive weighted fusion of the features extracted by the DRN and FCN.
[0008] Image classification: The improved ResNet18 image classification model was trained and used to classify the collected single-grain naked oat seed images.
[0009] The beneficial effects of adopting the above technical solution are as follows: when constructing an improved ResNet18 image classification model, the method described in this application first introduces a fusion multi-scale feature pyramid enhancement module MSFPE to enhance the model's ability to capture and fuse the multi-scale fine-grained features of oatmeal seeds; secondly, a DRFC network composed of a detail enhancement network DRN and a feature complementary network FCN is constructed to improve the ability to extract local key features and edge detail features of oatmeal seeds; finally, an adaptive collaborative fusion module ACFM is used to perform adaptive weighted fusion on the features extracted by DRN and FCN to enhance the model's attention to overall features. Experimental results show that the constructed model can significantly improve the accuracy of single-grain oatmeal seed image classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0011] Figure 1 is an overall flow chart of the method according to an embodiment of the present invention;
[0012] Figure 2a-2f 1 is a raw data diagram of naked oats seeds in an embodiment of the present invention (a, Baiyan No. 18; b, Bayou No. 5; c, Bayou No. 9; d, Huazao No. 2; e, Jinyan No. 18; f, Zhangyou No. 15);
[0013] Figure 3 1 is a schematic structural diagram of an image data acquisition system according to an embodiment of the present invention;
[0014] Figure 4 is a processing flow chart of the MSFPE module in an embodiment of the present invention;
[0015] Figure 5 is a processing flow chart of a DRFC network in an embodiment of the present invention;
[0016] Figure 6 is a processing flow chart of the ACFM module in an embodiment of the present invention;
[0017] Among them: 1. Sealed box; 2. Industrial camera; 3. Camera; 4. Pallet; 5. Computer; 6. LED light source. DETAILED DESCRIPTION
[0018] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.
[0019] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0020] Overall, such as Figure 1 As shown, the embodiment of the present invention discloses a fine-grained image classification method for naked oat seeds based on an improved ResNet18, comprising the following steps:
[0021] S1, image acquisition: Use an optical image data acquisition system to collect fine-grained images of naked oat seeds. After processing, several single-grain naked oat seed images are obtained to form a data set;
[0022] S2, data set processing: divide the data set into training set and test set according to the set ratio, and preprocess the training set and test set respectively;
[0023] S3, building an improved ResNet18 image classification model: Based on the backbone network of the ResNet18 model, first, the fused multi-scale feature pyramid enhancement module MSFPE is introduced afterwards. Secondly, a DRFC network consisting of a detail enhancement network DRN and a feature complementation network FCN is constructed. Finally, an adaptive collaborative fusion module ACFM is used to perform adaptive weighted fusion of the features extracted by DRN and FCN.
[0024] S4, image classification: Train the improved ResNet18 image classification model and use the trained improved ResNet18 image classification model to classify the collected single-grain naked oat seed images.
[0025] Application object: This application uses 6 kinds of naked oats seeds (Baiyan No. 18, Bayou No. 5, Bayou No. 9, Huazao No. 2, Jinyan No. 18 and Zhangyou No. 15, described as VAR1, VAR2, VAR3, VAR4, VAR5 and VAR6 respectively) as experimental application objects, and stores them at normal room temperature and in an environment with guaranteed ventilation and breathability. Figure 2a-2f Sample images of 6 types of seeds are given. The main colors of the experimental seeds are white and yellow, with an average seed length of 0.8 cm and an average width of 0.3 cm.
[0026] Image acquisition: Optical image data acquisition system, such as Figure 3As shown, the optical image data acquisition system in the image acquisition includes: an industrial camera 2 is fixed on the top of a cubic black sealed box 1, and a camera 3 is provided on the industrial camera 2. The lens of the camera 3 faces a tray 4 at the bottom center of the sealed box 1. A plurality of naked oat seed samples are placed on the tray 4. The distance between the lens and the naked oat seed samples is a set value. Two LED light sources are fixed on the periphery of the lens for filling light for the camera. The industrial camera 2 is controlled by a computer 5 and is used to operate under the control of the computer 5 and transmit the collected data to the computer 5 for processing.
[0027] Furthermore, a camera (model S10S-4K, Huiboshi Technology Co., Ltd., Shenzhen, China) was fixed on the top of a black cubic sealed box with a side length of 80 cm. The lens was facing the center of the bottom, the distance between the lens and the sample was 20 cm, and two square LED light sources were fixed on the periphery of the lens. Parameters such as aperture, focal length, and sample position were adjusted. The camera sensor calibration process was as follows: first, the aperture was adjusted to the maximum and the lens was set to telephoto mode. The focus was adjusted to make the sample image clear. Then the lens was adjusted to a wide angle. At this time, the panoramic image already included all the samples to be collected. The focus was adjusted again until the image was clearest.
[0028] Twenty-five samples of each oat variety were selected and arranged in a 5×5 array on a white tray, with a ruler used to maintain consistent spacing between seeds. Image data for the six oat varieties was collected using an optical image data acquisition system. Single-seed extraction was then performed on the original images: color space was converted based on the HSV color features of the oat seeds, and noise was removed through mask creation and morphological operations (opening and closing). A contour detection algorithm was used to identify target regions whose areas met a threshold, and the bounding box was expanded to capture single-seed images. This process ultimately yielded 3,300 high-quality single-seed oat images, totaling 550 for each variety. This provided a standardized dataset for subsequent model training.
[0029] Improved ResNet18 image classification model:
[0030] ResNet18 was chosen as the base model. With its moderate number of layers and fast convergence, ResNet18 effectively overcomes vanishing and exploding gradients using its residual block structure, providing strong support for deep network training. Its core skip connection mechanism further ensures network performance. ResNet18 consists of initial convolutional and pooling layers, four residual blocks (each containing two residual units), a global average pooling layer, and fully connected layers.
[0031] The oat seed images involved in this application belong to the category of fine-grained images, and the image differences between different types of oat seeds are extremely subtle. The shallow structure of ResNet18 has limitations in extracting such detailed features. To solve this problem, this application makes the following improvements to the ResNet18 model: 1) Based on the pre-trained ResNet18 model, a fused multi-scale feature pyramid enhancement module MSFPE is introduced; 2) Detail reinforcement networks (DRN) and feature complementary networks (FCN) are constructed; (3) an adaptive collaborative fusion module (ACFM) is designed to perform adaptive weighted fusion of the features extracted by DRN and FCN.
[0032] The improved ResNet18 network model operates as follows: First, preprocessed naked oat seed images are fed into the ResNet18 backbone network for preliminary learning. This network is then combined with Multi-Scale Feature Extraction (MSFPE) to enhance multi-scale feature extraction. The backbone network generates four layers of multi-scale feature maps, with scales of 2, 4, 8, and 16, corresponding to 64, 128, 256, and 512 output channels. These feature maps are fed through the bidirectional MSFPE to achieve a deep fusion of high-level semantic information and low-level detail information, and feature representation is optimized through convolution operations. The processed features are then fed into the Detail Reinforcement Network (DRN) and the Feature Complementation Network (FCN), enabling the model to simultaneously focus on important local features and edge detail information. Finally, the adaptive collaborative fusion module (ACFM) performs an adaptive weighted fusion of the features extracted by the DRN and FCN, further improving classification accuracy and robustness, enabling precise recognition and classification of different naked oat seed varieties.
[0033] Multi-scale feature pyramid enhancement module MSFPE:
[0034] In this application, the multi-scale feature pyramid enhancement module MSFPE is closely connected to the ResNet-18 backbone network for oat seed image feature extraction and enhancement. This module innovatively integrates multi-scale information interaction and feature enhancement technology, effectively making up for the shortcomings of traditional feature extraction methods in dealing with fine-grained images, and significantly improving the model's ability to perceive and capture subtle features of oat seeds. Its unique structural design and operating mechanism play a key role in improving the overall performance of the model. The feature fusion process of the MSFPE module in this model is as follows: Figure 4 shown.
[0035] 1) Multi-scale feature extraction and preliminary fusion:
[0036] The multi-scale feature maps C2 to C5 output by the ResNet-18 backbone network are imported into the MSFPE module. The module first processes the feature maps of different scales using dilated convolutions with different dilation rates. For the C2 feature map (which is large, rich in details, but relatively weak in semantic information), a 3×3 dilated convolution with a dilation rate of 1 is used to preserve the original details while enhancing the expression of local features. For the C3 feature map, a 3×3 dilated convolution with a dilation rate of 2 is used to further explore the associations between features while expanding the receptive field. The C4 and C5 feature maps correspond to dilated convolutions with dilation rates of 3 and 4, respectively, to enhance the extraction of large-scale semantic information.
[0037] After the feature map is processed by the dilated convolution, the number of channels is compressed by 1×1 convolution to reduce the amount of calculation and integrate the feature information, and the feature maps P2' to P5' are obtained respectively. Taking P2' as an example, its calculation formula is:
[0038] P2'=Conv 1×1 (Conv 3×3,dilation=1 (C2));
[0039] Among them, Conv 1×1 Represents a 1×1 convolution operation, Conv 3×3 Represents a 3×3 dilated convolution operation with a dilation rate of 1.
[0040] 2) Bidirectional feature fusion path:
[0041] To achieve deep fusion and enhancement of features, the MSFPE module constructs a bidirectional feature fusion path.
[0042] In the top-down path, the high-level feature map P5' is first expanded to the same size as the feature map P4' through an upsampling operation (using bilinear interpolation), and then fused with the feature map P4' element by element to obtain the preliminary fused feature map P4'. The process can be expressed as:
[0043] P4'=P4'+Upsample(P5');
[0044] Here, Upsample represents a bilinear interpolation upsampling operation. P4' is then optimized for feature representation using a 3x3 convolution to obtain P4*. This process is repeated until the fused P2* feature map is obtained.
[0045] The bottom-up path is the opposite. The low-level feature map P2 is downsampled (max pooling) and reduced in size to align with P3. The two are then element-wise added and then subjected to 3×3 convolution to obtain the N3 feature map. The same process is repeated to obtain the feature maps N4 and N5. For example, the calculation process of the feature map N3 is:
[0046] N3=Conv3×3 (P3*+MaxPool(P2*));
[0047] Among them, MaxPool represents the maximum pooling downsampling operation.
[0048] 3) Final feature fusion and output
[0049] Feature maps N3, N4, and N5 are upsampled to the same size as feature map N2 using bilinear interpolation and then concatenated along the channel dimension. The concatenated feature maps now have 1024 channels and a size of (1024, 56, 56). To further integrate features and improve their effectiveness and distinguishability, a 3×3 convolutional layer is used to convolve the concatenated feature maps, resizing the number of channels back to 256. The final fused feature map is output with a size of (256, 56, 56). This feature map carries rich multi-scale and multi-semantic information, providing a high-quality data foundation for feature extraction in the subsequent DRFC network.
[0050] The MSFPE module's bidirectional feature fusion path and multi-scale dilated convolution operations fully exploit the potential of features at different scales, achieving a deep fusion of detail and semantic information, significantly improving the model's ability to express subtle features. While enhancing feature expression, the MSFPE module balances computational cost and model performance through a rational design of convolution and sampling operations, effectively avoiding the computational burden associated with overly complex structures. This ensures the model's efficient operation in practical applications and provides a reliable feature extraction and enhancement solution for the complex naked oat seed image classification task.
[0051] Reinforcement Complementary Learning Network (DRFC):
[0052] This application constructs a reinforcement complementary learning network structure DRFC based on the ResNet18 model, which consists of a detail reinforcement network DRN and a feature complementary network FCN. In the fine-grained oat seed image classification task, DRN aims to enhance the ability to capture the key features of oat seeds, while FCN focuses on mining edge and non-significant features that are ignored by DRN. The two work together to improve the overall classification performance.
[17] .
[0053] Figure 5 The image, after the multi-scale deep features are extracted by the MSFPE module, enters the DRFC network processing process. Specifically, the key area mask is generated through the adaptive threshold method to guide the DRN and FCN to focus on different feature areas.
[0054] 1) Adaptive threshold key area mask generation:
[0055] For the feature map F output by the MSFPE module, first calculate its global mean μ and standard deviation σ:
[0056]
[0057] Where H, W, and C are the height, width, and number of channels of the feature map respectively, and F i,j,k Represents the value of the feature map at position (i, j) channel k. Set the adaptive threshold τ = μ + α × σ, where α is an adjustable parameter and its optimal value is determined through experiments. For each element F in the feature map i,j,k If F i,j,k ≥τ, the corresponding position in the mask M is set to 1, indicating that the position is a key area; otherwise it is set to 0.
[0058] 2) Detailed Enhanced Network (DRN) processing:
[0059] Multiply the generated key area mask M by the feature map F element by element to obtain the feature map F focused on the key area DRN =F×M. DRN to F DRN For processing, a series of convolutional layers and pooling layers are used to enhance the ability to express key features. For example, a 3×3 convolution kernel is used for convolution operation, while batch normalization and ReLU activation function are used, as shown below:
[0060] F DRN1 =RELU(BatchNorm(Conv 3×3 (F DRN )));
[0061] After multiple layers of processing, the detail enhancement network DRN outputs the feature map F DRN-out And adjust its size to (256,14,14) through adaptive average pooling to meet the subsequent fusion requirements.
[0062] 3) Adaptive threshold key area mask generation:
[0063] In order to make the feature complementary network FCN focus on non-critical areas, a reverse mask M is generated complement = 1-M, and multiply it element-wise with the feature map F to get F FCN =F×M complement FCN to F FCN For processing, the same 3×3 convolution kernel is used for convolution operation, combined with batch normalization and ReLU activation function:
[0064] F FCN1 =RELU(BatchNorm(Conv 3×3 (F FCN )))
[0065] After multiple layers of processing, FCN outputs the feature map F FCN-out and resizes it to (256, 14, 14) via adaptive average pooling.
[0066] 4) Feature fusion preparation:
[0067] After the above processing, the detail enhancement network DRN and the feature complementary network FCN output feature maps F DRN-out and F FCN-out , with sizes of (256, 14, 14), preparing for feature fusion in the subsequent ACFM module. The DRFC network, through the collaborative work of the DRN and FCN, comprehensively captures key and non-significant features, enhancing the model's discriminative capabilities. The adaptive threshold key region mask generation method dynamically determines key regions based on the statistical properties of feature maps, enabling more targeted feature extraction by the DRN and FCN, effectively improving the model's ability to learn the characteristics of different naked oat seeds.
[0068] ACFM adaptive collaborative fusion module:
[0069] This application proposes an ACFM adaptive collaborative fusion module. This module is based on dynamic attention allocation and cross-channel feature interaction mechanism, aiming to solve the problem of information redundancy and key feature weakening when fusing features of different network branches, and to achieve deep integration and optimization of oat seed image features. Its structure and workflow are as follows Figure 6 shown.
[0070] 1) Feature preprocessing and attention weight generation:
[0071] Feature map F output by detail enhancement network DRN and feature complementary network FCN DRN-out 、F FCN-out First, they pass through an independent feature transformation module. This module consists of two parallel 1×1 convolutional layers, one of which converts F DRN Converted to query vector Q, the calculation formula is:
[0072] Q = Conv 1×1 (F DRN-out );
[0073] Second, F FCN Mapped to key vector K and value vector V respectively:
[0074] K = Conv 1×1 (F FCN-out );
[0075] V=Conv 1×1 (F FCN-out );
[0076] Then, the original attention matrix A is generated by calculating the cosine similarity between the query vector Q and the key vector K raw :
[0077]
[0078] Where i and j represent the spatial position index of the feature map. To enhance the response to the key feature area of naked oat seeds, an adaptive adjustment factor γc based on channel statistics is introduced. This factor is dynamically calculated based on the variance and mean of each channel in the feature map:
[0079]
[0080] where σ c 、μ c are the standard deviation and mean of channel c respectively, and ∈ is a very small constant to avoid the denominator being zero. c With A raw Multiply channel by channel to get the final attention weight matrix A: A = γ c ×A raw .
[0081] 2) Final feature fusion and output:
[0082] Use the attention weight matrix A to perform weighted summation on the value vector V to generate the fusion feature Z:
[0083] Z=∑ j A(:,j)·V(j);
[0084] To further promote the interaction of features from different branches, a cross-channel feature interaction unit is designed. This unit performs a cyclic shift operation on Z and Q along the channel dimension, then adds them element-by-element, and then reconstructs the features through a 3×3 convolution:
[0085] Z'=Conv 3×3 (Z+Shift(Q));
[0086] Shift represents a cyclic shift operation, which forces the feature information to flow across channels by disrupting the channel order. Finally, Z' and Q are concatenated in the channel dimension, and the final fusion feature F is output after dimensionality reduction through the fully connected layer. final :F final =FC([Z';Q]).
[0087] 3) Final feature fusion and output
[0088] Dynamic Attention Allocation: Unlike the global, unified weighting of traditional attention mechanisms, ACFM uses channel-wise statistical adaptive adjustment factors to dynamically adjust the model's focus on key areas such as subtle textures and edge contours based on the distribution of features across different categories of oatmeal seeds. For example, when processing oatmeal varieties with minimal differences in seed coat texture, the model automatically increases the attention weight on features such as the seed tip shape.
[0089] Cross-channel feature interaction: The cyclic shift operation breaks the information barriers between feature channels, promoting the deep fusion of local detail features extracted by DRN and global structural features captured by FCN, avoiding the problem of feature information fragmentation caused by simple splicing, and significantly improving the model's ability to express the complex features of naked oats seeds.
[0090] Computational efficiency optimization: Through lightweight operations such as 1×1 convolution and circular shift, while achieving efficient feature fusion, the additional computational overhead is controlled within 5% of the original model's computational workload, ensuring the real-time and practicality of the improved model.
[0091] Model training and testing:
[0092] 1) Experimental platform and data parameters:
[0093] The computer processor model used in this application is Intel(R) Core(TM) i5-12400F 2.50GHz, the GPU model is NVIDAGeForce RTX4060, the video memory is 16GB, the operating system is Windows 11, the deep learning framework is Pytorch 2.1.0+cu121, the development environment is Python 3.9.21, and the computing architecture is cuda 12.1. The number of training rounds is set to 50 rounds, the number of input images per batch is 32, the initial learning rate is 0.001, the input oats seed image resolution is 224×224 pixels after image preprocessing, and the cross entropy loss function is used as the model loss evaluation function.
[0094] 2) Evaluation indicators:
[0095] To verify the effectiveness of this model, the evaluation indicators used in this experiment include: accuracy (A,%), recall (R,%), precision (P,%), and F1 score (F1,%).
[0096] Experimental results and analysis:
[0097] In order to further verify the effectiveness and synergy of the three improved modules of MSFPE, DRFC, and ACFM, this application designed and carried out a series of ablation experiments. The experimental results are shown in Table 1.
[0098] Experimental results show that the introduction of the MSFPE module alone to the original ResNet18 model increased the model's accuracy in the oat seed image classification task from 78.03% to 82.15%, an improvement of 4.12 percentage points. This demonstrates that the MSFPE module, through bidirectional feature fusion and multi-scale dilated convolution, effectively enhances the model's ability to capture subtle features of oat seeds. Furthermore, the introduction of the DRFC network alone increased the model's accuracy, recall, precision, and F1 score by 3.78, 3.81, 3.82, and 3.83 percentage points, respectively, demonstrating that the DRFC network effectively enhances the extraction of key regions and edge details in oat seeds.
[0099] When the MSFPE module is combined with the DRFC network, model performance is further improved, reaching an accuracy of 85.72%, a 7.69 percentage point increase compared to the original model. However, this improvement falls short of the combined effect of the two modules alone. Analysis reveals that this is because conventional fusion methods fail to fully leverage the strengths of the two modules, resulting in the weakening of some detailed features during the fusion process.
[0100] Finally, building on the aforementioned improvements, the ACFM module was introduced to produce the final improved ResNet18 model (ResNet18-MSFPE-DRFC-ACFM). This model performed exceptionally well across all metrics, with accuracy increasing to 88.63%, a 10.6 percentage point improvement over the original model. Recall, precision, and F1 scores also increased to 88.53%, 88.74%, and 88.63%, respectively. This demonstrates that the synergistic effect of the three improved modules significantly improved the model's classification performance for fine-grained naked oat seed images.
[0101] In terms of model complexity, the introduction of various improved modules has reduced the inference speed of the improved ResNet18 model from the original 120 frames per second to 98 frames per second, and the model size has increased from 85.3MB to 108.5MB. Despite this, the improved model can still be successfully deployed on computers with standard computing power, meeting the needs of practical applications.
[0102] Table 1. Comparison of ablation test performance
[0103]
[0104] This application systematically improves the ResNet18 network to solve the problem of fine-grained oat seed image classification. By introducing the fused multi-scale feature pyramid enhancement module (MSFPE), constructing a Detail Reinforcement Network (DRN) and a Feature Complementary Network (FCN) composed of a Detail Reinforcement Network (DRN), and combining it with the Adaptive Collaborative Fusion Module (ACFM), the application achieves a comprehensive improvement in the feature extraction, enhancement, and fusion capabilities of oat seed images.
[0105] In summary, this application uses six types of oat seeds as classification objects, constructs a dedicated image dataset for model training and testing. Experimental results show that the improved ResNet18 model achieves significant results in the oat seed classification task. Compared with the original model, the accuracy rate is increased from 78.03% to 88.63%, and the precision, recall rate, and F1 score are increased to 88.74%, 88.53%, and 88.63%, respectively. Although the improved model's structural optimization results in a reduction in inference speed of 98 frames per second and an increase in model size to 108.5MB, it can still be stably deployed on conventional computing power devices.
Claims
1. A fine-grained image classification method for naked oat seeds based on improved ResNet18, characterized by The steps include: Image acquisition: An optical image data acquisition system is used to collect fine-grained images of naked oat seeds. After processing, several single-grain naked oat seed images are obtained to form a data set. Dataset processing: Divide the dataset into training set and test set according to the set ratio, and preprocess the training set and test set respectively; Build an improved ResNet18 image classification model: Based on the backbone network of the ResNet18 model, first, introduce the fused multi-scale feature pyramid enhancement module (MSFPE) afterwards. Secondly, build a DRFC network consisting of a detail reinforcement network (DRN) and a feature complementation network (FCN). Finally, use the adaptive collaborative fusion module (ACFM) to perform adaptive weighted fusion of the features extracted by the DRN and FCN. Image classification: The improved ResNet18 image classification model was trained and used to classify the collected single-grain naked oat seed images.
2. The naked oat seed fine-grained image classification method based on improved ResNet18 according to claim 1, characterized in that: The optical image data acquisition system in the image acquisition includes: An industrial camera (2) is fixed on the top of a cubic black sealed box (1). A camera (3) is provided on the industrial camera (2). The lens of the camera (3) faces a tray (4) at the center of the bottom of the sealed box (1). A plurality of naked oat seed samples are placed on the tray (4). The distance between the lens and the naked oat seed samples is a set value. Two LED light sources (6) are fixed on the periphery of the lens for supplementary lighting for the camera. The industrial camera (2) is controlled by a computer (5) and is used to operate under the control of the computer (5) and transmit collected data to the computer (5) for processing.
3. The naked oat seed fine-grained image classification method based on improved ResNet18 according to claim 1, characterized in that: The method for obtaining several single oat seed images includes the following steps: converting the color space based on the HSV color features of the oat seeds, removing noise through mask creation and morphological operations, screening the target area whose area meets the threshold using a contour detection algorithm, and capturing the single oat seed image after expanding the bounding box.
4. The naked oat seed fine-grained image classification method based on improved ResNet18 according to claim 1, characterized in that The data set processing includes the following steps: The dataset is divided into training and test sets in a ratio of 8:2, and data enhancement is performed through data preprocessing. The data preprocessing includes random scaling and cropping, image enhancement, Tensor conversion, and normalization operations on the training set; and scaling, cropping, Tensor conversion, and normalization operations on the test set.
5. The naked oat seed fine-grained image classification method based on improved ResNet18 according to claim 1, characterized in that The processing method of the improved ResNet18 image classification model includes the following steps: The preprocessed oat seed images enter the backbone network of the ResNet18 model for preliminary learning. The backbone network generates four layers of multi-scale feature maps with scales of 2, 4, 8 and 16 times, corresponding to 64, 128, 256 and 512 output channels respectively; the above multi-scale feature maps are deeply integrated with high-level semantic information and low-level detail information through the bidirectional path of the multi-scale feature pyramid enhancement module MSFPE, and the feature expression is optimized through convolution operation; then, the processed feature maps are input into the DRFC network composed of the detail enhancement network DRN and the feature complementation network FCN for processing, so that the model can focus on local important features and edge detail information at the same time; finally, the features extracted by the detail enhancement network DRN and the feature complementation network FCN are adaptively weighted and fused through the ACFM module to realize the recognition and classification of different types of oat seeds.
6. The naked oat seed fine-grained image classification method based on improved ResNet18 according to claim 1, characterized in that: The processing method of the multi-scale feature pyramid enhancement module MSFPE includes the following steps: 1) Multi-scale feature extraction and preliminary fusion: The multi-scale feature maps C2 to C5 output by the ResNet-18 backbone network are respectively imported into the MSF PE module. For feature maps of different scales, the module first uses dilated convolutions with different dilation rates for processing. For the C2 feature map, a 3×3 dilated convolution with a dilation rate of 1 is used to retain the original details while enhancing the expression of local features. For the C3 feature map, a 3×3 dilated convolution with a dilation rate of 2 is used to further explore the correlation between features while expanding the receptive field. For the C4 and C5 feature maps, dilated convolutions with dilation rates of 3 and 4 are used, respectively, to enhance the extraction of large-scale semantic information. After the feature map is processed by the dilated convolution, the number of channels is compressed by 1×1 convolution to obtain the feature maps P2' to P5' respectively; Taking P2' as an example, its calculation formula is: P2'=Conv 1×1 (Conv 3×3,dilation=1 (C2)) Among them, Conv 1×1 Represents a 1×1 convolution operation, Conv 3×3 represents a 3×3 dilated convolution operation with a dilation rate of 1; 2) Bidirectional feature fusion path: In the top-down path of the MSFPE module, the high-level feature map P5' is first expanded to the same size as the feature map P4' through an upsampling operation, and then fused with the feature map P4' element by element to obtain the preliminary fused feature map P4'. The process is expressed as follows: P4'=P4'+Upsample(P5'); Among them, Upsample represents the bilinear interpolation upsampling operation; the initial fusion feature map P4" is then optimized by 3×3 convolution to obtain the P4* feature map. This process is repeated until the fused P2* feature map is obtained; The bottom-up path of the MSFPE module is the opposite. After the low-level feature map P2 is downsampled and reduced in size, it is aligned with the feature map P3. The two are added element by element, and then a 3×3 convolution is performed to obtain the N3 feature map. The feature maps N4 and N5 are obtained in the same way. The calculation process of the feature map N3 is: N3=Conv 3×3 (P3*+MaxPool(P2*)); Among them, MaxPool represents the maximum pooling downsampling operation; 3) Final feature fusion and output The feature maps N3, N4, and N5 are upsampled to the same size as the feature map N2 using bilinear interpolation, and then concatenated in the channel dimension. At this time, the number of channels of the concatenated feature map becomes 1024, and the size is (1024, 56, 56). A 3×3 convolutional layer is used to perform a convolution operation on the concatenated feature map, and the number of channels is adjusted back to 256. Finally, the fused feature map is output with a size of (256, 56, 56).
7. The naked oat seed fine-grained image classification method based on improved ResNet18 according to claim 1, characterized in that: The DRFC network processing method includes the following steps: Generate key area masks through adaptive thresholding method to guide the detail enhancement network DRN and feature complementation network FCN to focus on different feature areas respectively; 1) Adaptive threshold key area mask generation: For the feature map F output by the multi-scale feature pyramid enhancement module MSFPE, first calculate its global mean μ and standard deviation σ: Where H, W, and C are the height, width, and number of channels of the feature map respectively, and F i,j,k Represents the value of the feature map at position (i, j) channel k; set the adaptive threshold τ = μ + α × σ, where α is an adjustable parameter and its optimal value is determined through experiments; for each element F in the feature map i,j,k , if F i,j,k ≥τ, the corresponding position in the mask M is set to 1, indicating that the position is a key area; otherwise it is set to 0; 2) Detailed Enhanced Network (DRN) processing: Multiply the generated key area mask M by the feature map F element by element to obtain the feature map F focused on the key area DRN =F×M; detail enhancement network DRN to F DRN For processing, convolutional layers and pooling layers are used to enhance the expression of key features. A 3×3 convolution kernel is used for convolution operation, and batch normalization and ReLU activation functions are used, as shown below: F DRN1 =RELU(BatchNorm(Conv 3×3 (F DRN ))) After multiple layers of processing, the detail enhancement network DRN outputs the feature map F DRN-out , and adjust its size to (256, 14, 14) through adaptive average pooling to meet the subsequent fusion requirements; 3) Adaptive threshold key area mask generation: In order to make the feature complementary network FCN focus on non-critical areas, a reverse mask M is generated complement = 1-M, and multiply it element-wise with the feature map F to get F FCN =F×M complement ; Feature complementary network FCN to F FCN For processing, a 3×3 convolution kernel is used for convolution operation, combined with batch normalization and ReLU activation function: F FCN1 =RELU(BatchNorm(Conv 3×3 (F FCN ))) After multi-layer processing, the feature complementary network FCN outputs the feature map F FCN-out , and resize it to (256,14,14) through adaptive average pooling; 4) Feature fusion preparation: After the above processing, the detail enhancement network DRN and the feature complementary network FCN output feature maps F DRN-out and F FCN-out , the sizes are all (256,14,14).
8. The naked oat seed fine-grained image classification method based on improved ResNet18 according to claim 1, characterized in that: The processing method of the adaptive collaborative fusion module ACFM includes the following steps: 1) Feature preprocessing and attention weight generation Feature map F output by detail enhancement network DRN and feature complementary network FCN DRN-out 、F FCN-ou t first passes through an independent feature transformation module, which consists of two parallel 1×1 convolutional layers, one of which transforms the feature map F DRN-out Converted to query vector Q, the calculation formula is: Q=Conv 1×1 (F DRN-out ); Second, F FCN-out Mapped to key vector K and value vector V respectively: K=Conv 1×1 (F FCN-out ); V=Conv 1×1 (F FCN-out ); Then, the original attention matrix A is generated by calculating the cosine similarity between the query vector Q and the key vector K raw : Where i and j represent the spatial position index of the feature map respectively; the adaptive adjustment factor γ based on channel statistics is introduced c , this factor is dynamically calculated based on the variance and mean of each channel of the feature map: where σ c 、μ c are the standard deviation and mean of channel c respectively, ∈ is a very small constant to avoid the denominator being zero, and γ c With A raw Multiply channel by channel to get the final attention weight matrix A: A=γ c ×A raw ; 2) Final feature fusion and output: Use the attention weight matrix A to perform weighted summation on the value vector V to generate the fusion feature Z: Z =∑ j A(:,j)·V(j); Design a cross-channel feature interaction unit, which performs a cyclic shift operation on the channel dimension of the fused feature Z and the query vector Q, then adds them element by element, and then reorganizes the features through 3×3 convolution: Z'=Conv 3×3 (Z+Shift(Q)); Shift represents a circular shift operation, which forces feature information to flow across channels by disrupting the order of channels. Finally, Z' is concatenated with the query vector Q in the channel dimension, and the final fusion feature F is output after dimensionality reduction through the fully connected layer. final : F final = FC([Z';Q]).
9. The naked oat seed fine-grained image classification method based on improved ResNet18 according to claim 1, characterized in that: When training the improved ResNet18 image classification model, the number of training rounds was set to 50, the number of input images per batch was 32, the initial learning rate was 0.001, the resolution of the input oat seed image was 224×224 pixels after image preprocessing, and the cross entropy loss function was used as the model loss evaluation function.
10. The naked oat seed fine-grained image classification method based on improved ResNet18 according to claim 9, characterized in that: To verify the effectiveness of the improved ResNet18 image classification model, the evaluation indicators used include: accuracy, recall, precision and F1 score.