SF1-BAGAN image data category balance enhancement method and image classification method

By improving the SF1-BAGAN image data balancing enhancement method and lightweight convolutional neural network, the problems of dataset imbalance and insufficient feature extraction by lightweight network in garbage image classification are solved, and high-precision garbage image classification with efficient data augmentation and lightweight classification model is achieved.

CN121033618APending Publication Date: 2025-11-28UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510954985.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing garbage image classification models suffer from class imbalance when processing regional garbage datasets. Lightweight networks lack the ability to extract features from difficult-to-classify samples, and existing data augmentation methods have limited effectiveness, making it difficult to improve the model's generalization ability and classification accuracy.

Method used

An SF1-BAGAN image data balancing enhancement method is adopted. A low-dimensional feature space is generated through an improved autoencoder pre-training process. DeepSMOTE interpolation is performed between the autoencoder and generator during the generative adversarial stage. Combined with the F1-score obtained from transfer learning, the F1-score for each category is dynamically determined, and the sampling quantity for each category is dynamically calculated to generate a balanced image dataset. This method is applied to a lightweight convolutional neural network. Furthermore, the CBAM module and Focal Loss loss function are integrated into the lightweight convolutional neural network to improve feature representation and classification performance.

Benefits of technology

It effectively alleviates the problem of class imbalance in datasets, improves the feature extraction ability and classification accuracy of lightweight networks for difficult-to-classify samples, and achieves high-precision garbage image classification while maintaining lightweight design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033618A_ABST
    Figure CN121033618A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, in particular to an SF1-BAGAN image data category balance enhancement method and an image classification method. Aiming at the problem of class imbalance of a junk image data set, an SF1-BAGAN enhancement model is proposed, the quality of a feature space is improved by reconstructing an auto-encoder pre-training process, SMOTE interpolation is introduced in a generative adversarial stage, and meanwhile, the sampling number is dynamically calculated based on a class F1-score obtained by transfer learning, so that adaptive sample expansion is realized, and the robustness of the method is improved. Category deviation is relieved from the source; for the defect that a lightweight model is insufficient in recognition of difficult-to-distinguish samples, a CBAM mixed attention module is fused in a MobileNetV3-Marge network to replace an SE module, feature expression is enhanced through channel and space two-dimensional weighting, a Focal Loss loss function is adopted to replace cross entropy, and the sensitivity of the model to the difficult-to-distinguish samples is improved by adjusting the weight of the difficult samples. While the imbalance problem of the data set is effectively suppressed, the feature extraction capability of the lightweight network on the difficult-to-distinguish samples is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, specifically to an SF1-BAGAN image data category balancing enhancement method and an image classification method. Background Technology

[0002] Currently, many methods exist for garbage image classification, but a sufficiently perfect and highly effective solution still lacks a complete framework. Most deep learning-based garbage image classification schemes use publicly available garbage datasets for network training. However, these datasets often fail to fully represent the true state of household waste in a given area, exhibiting issues such as limited image categories, poor image quality, or failure to meet local waste classification standards, thus restricting the model's generalization ability. While traditional oversampling or generative adversarial networks (such as BAGAN) are used for data augmentation, the augmented image quality is unstable, and it is difficult to alleviate class imbalance.

[0003] Existing classification models face a trade-off between lightweight design and high accuracy: lightweight networks (such as MobileNetV3) can meet real-time requirements, but lack the ability to extract features from difficult-to-classify samples; while high-accuracy models often have a large number of parameters and are computationally expensive. Although attempts have been made to improve performance through attention mechanisms (such as SE modules) or loss function optimization, the effects of a single improvement are limited, especially when dealing with complex samples after data augmentation, where classification accuracy still faces significant bottlenecks. Therefore, there is an urgent need for a method that integrates efficient data augmentation with lightweight and high-accuracy classification. Summary of the Invention

[0004] The technical problem solved by this invention is to provide an SF1-BAGAN image data class balancing enhancement method and an image classification method to solve the technical problems of imbalanced datasets and insufficient feature extraction ability of lightweight networks for difficult-to-classify samples.

[0005] The basic solution provided by this invention is an SF1-BAGAN image data class balancing enhancement method, applied to class-imbalanced image data processing, comprising the following steps:

[0006] Autoencoder pre-training process: Features are extracted from the original imbalanced dataset through an improved autoencoder pre-training process to generate a low-dimensional feature space;

[0007] DeepSMOTE interpolation operation: In the generative adversarial phase, DeepSMOTE interpolation operation is performed on the low-dimensional feature space of the autoencoder output.

[0008] Dynamic sampling based on F1-score: This method uses transfer learning to obtain the F1-score evaluation metric for each category, and dynamically determines the number of interpolation vectors for each category based on the F1-score. The sampling formula is as follows:

[0009]

[0010] In the formula, c represents the category, and n sanmpc f represents the number of samples to be taken from category c, m is the number of categories in the dataset, and f 1c For category c, the evaluation metric is F1-score, 1-f 1c This represents the error rate for class c, where N is the total number of samples in the original dataset;

[0011] Generate a balanced image dataset: Input the interpolated feature vectors into the generator, and merge the output enhanced image with the original imbalanced dataset to form a class-balanced image dataset.

[0012] Furthermore, the improved autoencoder pre-training process includes:

[0013] The front-end codec reconstructs the original data and calculates the reconstruction loss;

[0014] The backend adjusts the order of reconstructed images of the same category and calculates the penalty loss.

[0015] The sum of reconstruction loss and penalty loss is used as the pre-training optimization objective.

[0016] Furthermore, the DeepSMOTE interpolation operation is performed between the encoder and generator during the generative adversarial phase, and the generator parameters are initialized by the autoencoder decoder during the pre-training phase.

[0017] An image classification method includes the following steps:

[0018] S1: Obtain the raw garbage image dataset and identify its imbalanced class distribution;

[0019] S2: The recognized dataset is augmented using one of the SF1-BAGAN image data class balancing enhancement methods described in any of the above items;

[0020] S3: Input the processed dataset into the improved lightweight convolutional neural network classification model for training;

[0021] S4: Deploy a trained lightweight convolutional neural network classification model for real-time garbage image classification.

[0022] Furthermore, the lightweight convolutional neural network classification model in S3 adopts the MobileNetV3 series network model, and its improvements include:

[0023] When the number of input channels and output channels of the inverse residual structure are the same and the convolution stride is 1, the SE module in the inverse residual structure is replaced with the CBAM module.

[0024] Furthermore, the improvement to the lightweight convolutional neural network classification model also includes replacing the cross-entropy loss function with the Focal Loss loss function.

[0025] Furthermore, the CBAM module includes a channel attention submodule and a spatial attention submodule connected in sequence. The channel attention submodule adopts a parallel structure of global average pooling and max pooling, and the spatial attention submodule adopts a channel compression convolutional layer.

[0026] The principles and advantages of this invention are as follows: Addressing the class imbalance problem in garbage image datasets, this invention proposes an SF1-BAGAN enhancement model. It improves the feature space quality by reconstructing the autoencoder pre-training process (joint reconstruction loss and order penalty loss), and introduces SMOTE interpolation during the generative adversarial stage. Simultaneously, it dynamically calculates the sampling quantity based on the class F1-score obtained through transfer learning, achieving adaptive sample expansion and mitigating class bias at its source. For the deficiency of lightweight models in recognizing difficult samples, it integrates a CBAM hybrid attention module into the MobileNetV3-Large network to replace the SE module. This enhances feature representation through channel and spatial dual-dimensional weighting, and replaces cross-entropy with Focal Loss. By adjusting the weights of difficult samples, it improves the model's sensitivity to difficult samples. While effectively suppressing the dataset imbalance problem, it improves the feature extraction capability of lightweight networks for difficult samples. Attached Figure Description

[0027] Figure 1 This is a flowchart illustrating the steps of an embodiment of the SF1-BAGAN image data category balancing enhancement method of the present invention.

[0028] Figure 2 This is a flowchart illustrating the steps of a second embodiment of the image classification method of the present invention. Detailed Implementation

[0029] The following detailed description illustrates the specific implementation method:

[0030] The specific implementation process is as follows:

[0031] Example 1

[0032] Example 1 is attached. Figure 1 As shown, an SF1-BAGAN image data class balancing enhancement method includes the following steps:

[0033] Autoencoder pre-training process: Features are extracted from the original imbalanced dataset through an improved autoencoder pre-training process to generate a low-dimensional feature space;

[0034] DeepSMOTE interpolation operation: In the generative adversarial phase, DeepSMOTE interpolation operation is performed on the low-dimensional feature space of the autoencoder output.

[0035] Dynamic sampling based on F1-score: This method uses transfer learning to obtain the F1-score evaluation metric for each category, and dynamically determines the number of interpolation vectors for each category based on the F1-score. The sampling formula is as follows:

[0036]

[0037] In the formula, c represents the category, and n sanmpc f represents the number of samples to be taken from category c, m is the number of categories in the dataset, and f 1c For category c, the evaluation metric is F1-score, 1-f 1c This represents the error rate for class c, where N is the total number of samples in the original dataset;

[0038] Generate a balanced image dataset: Input the interpolated feature vectors into the generator, and merge the output enhanced image with the original imbalanced dataset to form a class-balanced image dataset.

[0039] Specifically, in this embodiment, the SF1-BAGAN data augmentation model is mainly used to augment the imbalanced image dataset:

[0040] 1. A low-dimensional feature space is generated through an improved autoencoder pre-training process.

[0041] The improved autoencoder pre-training process includes:

[0042] ①. The front-end codec reconstructs the original data and calculates the reconstruction loss, thereby ensuring the fidelity of the feature space.

[0043] ②. The backend adjusts the order of reconstructed images of the same category and calculates the penalty loss, thereby enhancing the diversity of samples of the same category through the order penalty loss.

[0044] ③. Using the sum of reconstruction loss and penalty loss as the pre-training optimization objective, the joint optimization of the two significantly improves the semantic rationality of the generated samples.

[0045] 2. In the generative adversarial phase, DeepSMOTE interpolation is performed on the low-dimensional feature space of the autoencoder output. This DeepSMOTE interpolation is performed between the encoder and generator in the generative adversarial phase, and the generator parameters are initialized by the autoencoder decoder in the pre-training phase.

[0046] 3. Obtain the F1-score evaluation metric for each category based on transfer learning, and dynamically determine the number of interpolation vectors for each category according to the F1-score evaluation metric. The sampling formula is as follows:

[0047]

[0048] In the formula, c represents the category, and n sanmpc f represents the number of samples to be taken from category c, m is the number of categories in the dataset, and f 1c Let F1 be the evaluation metric for category c, and N be the total number of samples in the original dataset, 1-f 1c This represents the error rate for category c. A higher F1 score indicates better classification performance for that category. The F1 score is measured between 0 and 1, hence 1-f 1c A higher F1 score indicates that class c performs poorly within the dataset. In this embodiment, the F1 score for each class is obtained by testing an existing classification model trained on the original dataset. By calculating the difference (number of samples required) according to the formula, classes with high error rates can obtain more generated samples, specifically alleviating class imbalance. This also avoids the blindness of manually setting sampling ratios, achieving a closed-loop optimization of data augmentation and classification performance.

[0049] Regarding data augmentation, this embodiment verifies the effectiveness of each improvement point of the present invention through image generation ablation experiments. As shown in Table 1, the image generation ablation experiment results table, this embodiment performs image generation training on the original dataset. The three generation models generate 500 images for each class in the dataset. The 40 categories are divided into few-sample classes and many-sample classes according to the number of samples, and the FID and SSIM of these generated images are calculated.

[0050] Table 1. Results of Image Generation and Ablation Experiments

[0051]

[0052] Experimental results show that SF1-BAGAN reduces the FID of images generated by data augmentation by 18.21 compared to BAGAN, while improving SSIM by 0.066. Furthermore, as shown in Table 2, this embodiment further designed an image classification comparison experiment to verify the classification performance of this model when combined with different CNN classification models on the augmented dataset. The experimental results show that when SF1-BAGAN data augmentation is combined with a CNN network, the accuracy improvement is higher than the other control methods, verifying that SF1-BAGAN data augmentation can effectively suppress the dataset imbalance problem.

[0053] Table 2 Experimental Results of MobileNetV3 Model

[0054]

[0055] Example 2

[0056] Example 2 is attached. Figure 2 As shown, an image classification method includes the following steps:

[0057] S1: Obtain the raw garbage image dataset and identify its imbalanced class distribution;

[0058] S2: The above-mentioned SF1-BAGAN image data class balancing enhancement method is used to expand the recognized dataset;

[0059] S3: Input the processed dataset into the improved lightweight convolutional neural network classification model for training;

[0060] S4: Deploy a trained lightweight convolutional neural network classification model for real-time garbage image classification.

[0061] Specifically, in S1, the original garbage dataset in this embodiment uses the Huawei municipal solid waste dataset, which identifies categories with significant differences in sample size. By locating the few categories that need to be strengthened, it provides target guidance for subsequent dynamic sampling.

[0062] In S2, the imbalanced dataset is augmented using one of the SF1-BAGAN image data class balancing enhancement methods described above.

[0063] S3 inputs the processed dataset into an improved lightweight convolutional neural network classification model for training, involving collaborative improvements to the lightweight classification model. Its lightweight convolutional neural network classification model is achieved through the following improvements:

[0064] ① The lightweight convolutional neural network classification model adopts the MobileNetV3 series model, and this embodiment uses the MobileNetV3-Large network. The inverse residual structure SE module in MobileNetV3-Large is replaced with a CBAM module; the attention mechanism is upgraded, replacing the SE module with CBAM in specific inverse residual blocks (when the input / output channels are the same and the stride is 1). The CBAM module contains sequentially connected channel attention sub-modules and spatial attention sub-modules. The channel attention sub-module uses a parallel structure of global average pooling and max pooling to focus on important feature channels; the spatial attention sub-module uses channel compression convolutional layers to capture key spatial locations. This dual-dimensional attention enhances the ability to distinguish details of similar waste (such as glass bottles / plastic bottles), while replacing the SE module only when the input / output channels are the same and the stride is 1 balances computational efficiency and feature extraction performance.

[0065] ②. The Focal Loss loss function is used instead of the cross-entropy loss function, which automatically reduces the weight of easily distinguishable samples and focuses on training difficult samples (such as deformed or occluded junk images). In addition, it can produce a synergistic effect with CBAM. CBAM strengthens the feature representation of difficult samples, while Focal Loss amplifies the gradient feedback of difficult samples. The combination of the two significantly improves the model's sensitivity to edge cases.

[0066] Finally, the lightweight convolutional neural network classification model trained in S4 is deployed for real-time garbage image classification. It is trained using an enhanced balanced dataset, which improves the recognition rate of minority and hard-to-classify samples while maintaining lightweight design.

[0067] Regarding classification performance, the combination of SF1-BAGAN data augmentation and MobileNetV3-Large performed poorly on some difficult-to-classify categories. To further improve the model's classification performance, this embodiment conducted an in-depth investigation of the model's key components, analyzed the possible reasons affecting model performance, and designed improvement schemes accordingly. Finally, two sets of experiments verified the effectiveness of each improvement point and whether it improved the classification accuracy of difficult-to-classify samples and the overall classification performance.

[0068] As shown in Table 3, the first group of experiments was the ablation experiment. The results showed that compared with the baseline MobileNetV3-Large model, the Top-1 accuracy and F1-score of the models that introduced Focal Loss alone and the models that introduced CBAM alone were both improved. The complete improved model that combined the two improvements improved the Top-1 accuracy by 3.94 percentage points and the F1-score by 3.33 percentage points compared with the baseline model, and performed the best among all experimental groups. Therefore, all improvements are helpful in improving the model performance.

[0069] Table 3 Ablation Experiment Results

[0070]

[0071] The second group consists of comparative experiments between the improved model and commonly used lightweight models, as shown in Table 4. The results show that the improved model CFL_MobileNetV3-Large achieved the highest Top-1 accuracy in all experimental groups, and the number of parameters and algorithm complexity were only 0.4M higher than the baseline model, demonstrating that the improved model has excellent overall performance.

[0072] Table 4 Comparison of experimental results

[0073]

[0074] In summary, this invention addresses the class imbalance problem in garbage image datasets by proposing the SF1-BAGAN enhancement model. It improves the feature space quality by reconstructing the autoencoder pre-training process (joint reconstruction loss and order penalty loss), and introduces SMOTE interpolation during the generative adversarial stage. Simultaneously, it dynamically calculates the sampling quantity based on the class F1-score obtained through transfer learning, achieving adaptive sample expansion and mitigating class bias at its source. To address the weakness of lightweight models in identifying difficult samples, the invention integrates a CBAM hybrid attention module into the MobileNetV3-Large network to replace the SE module. This enhances feature representation through channel and spatial dual-dimensional weighting, and replaces cross-entropy with Focal Loss, improving the model's sensitivity to difficult samples by adjusting their weights. This effectively suppresses the dataset imbalance problem while improving the feature extraction capability of lightweight networks for difficult samples.

[0075] The above are merely embodiments of the present invention. Commonly known structures and characteristics are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. An SF1-BAGAN image data class balancing enhancement method, characterized in that, The method is applied to class imbalance image data processing, and comprises the following steps: An improved autoencoder pre-training process is used to extract features from the original unbalanced data set and generate a low-dimensional feature space; DeepSMOTE interpolation operation: in the generation stage, the low-dimensional feature space output by the autoencoder is subjected to DeepSMOTE interpolation operation; Dynamic sampling based on F1-score: based on the F1-score evaluation index of each class obtained by transfer learning, the number of interpolation vectors of each class is dynamically determined according to the evaluation index F1-score, and the sampling formula is: In the formula, c represents a category, n sanmpc represents the number of samples required for category c, m is the number of categories of the data set, f 1c is the evaluation index F1-score of category c, 1-f 1c represents the error rate of category c, and N is the total number of samples of the original data set; Generating a balanced image data set: input the interpolated feature vector into the generator to output enhanced images and original unbalanced data sets, and form a class-balanced image data set.

2. The SF1-BAGAN image data class balancing enhancement method of claim 1, wherein: The improved autoencoder pre-training process comprises: The front-end codec reconstructs the original data and calculates the reconstruction loss; The rear-end adjusts the order of the reconstructed images of the same class and calculates the penalty loss; The sum of the reconstruction loss and the penalty loss is used as the pre-training optimization target.

3. The SF1-BAGAN image data class balancing enhancement method of claim 1, wherein: The DeepSMOTE interpolation operation is performed between the encoder and the generator in the generation stage, and the generator parameters are initialized by the autoencoder decoder in the pre-training stage.

4. An image classification method characterized by, The method is applied to garbage image classification, and comprises the following steps: S1: obtaining an original garbage image data set and identifying the class imbalance distribution thereof; S2: using the SF1-BAGAN image data class balance enhancement method of any one of claims 1-3 to expand the identified data set; S3: inputting the processed data set into an improved lightweight convolutional neural network classification model for training; S4: deploying the trained lightweight convolutional neural network classification model for real-time garbage image classification.

5. The image classification method of claim 4, wherein: The lightweight convolutional neural network classification model in S3 uses a MobileNetV3 series network model, and the improvements thereof comprise: When the input channel number and the output channel number of the reverse residual structure are the same and the convolution step is 1, the SE module in the reverse residual structure is replaced by a CBAM module.

6. The image classification method of claim 5, wherein: The improvement of the lightweight convolutional neural network classification model also comprises: using a Focal Loss loss function instead of a cross-entropy loss function.

7. The image classification method of claim 6, wherein: The CBAM module comprises a channel attention submodule and a spatial attention submodule connected in sequence, the channel attention submodule adopts a parallel structure of global average pooling and maximum pooling, and the spatial attention submodule adopts a channel compression convolution layer.