Image fusion classification method and device based on segmentation guide data enhancement

The method leverages FastSAM network segmentation and information gain optimization to address data scarcity and imbalance in lightweight models, enhancing model performance and robustness by generating semantically consistent augmented data, achieving substantial accuracy gains.

CN120318557AActive Publication Date: 2025-07-15UNIV OF SCI & TECH BEIJING
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510335372.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-15
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

In the case of scarcity of data and resource limitation, the generalization ability of lightweight image classification models is insufficient. Traditional data augmentation methods cannot effectively improve the classification accuracy and generalization ability of the model, and the generative data augmentation method may lead to data distribution offset and overfitting.

Method used

Image segmentation is used to generate segmentation graphs, balanced segmentation graphs are selected through cosine similarity and information gain optimization, enhancement data sets are built, and training is performed using ResNet50 model to ensure consistency between the distribution of the enhanced data and the original data, and dynamically adjust the number of enhanced data to avoid redundancy.

Benefits of technology

The classification performance of the model has been significantly improved. The Top-1 accuracy of ResNet-50 has increased from 77.02% to 81.72%, and the Top-1 accuracy of RepViT has increased from 85.61% to 87.03%, reducing the workload of manual screening and reducing the risk of overfitting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318557A_ABST
    Figure CN120318557A_ABST
Patent Text Reader

Abstract

The invention provides an image fusion classification method and device based on segmentation guide data enhancement, and relates to the technical field of computer vision. The method comprises the following steps: segmenting each image in a training data set by using a FastSAM network; calculating the cosine similarity of the image and the segmented image, selecting a balanced segmented image, and constructing a segmented data set; distance norm measurement and information gain are constructed to determine the segmentation image enhancement number, an enhancement data set is obtained, and the classification model is trained. According to the method, the image segmentation technology is combined, accurate segmentation feature extraction is carried out on the image, key semantic regions and structural features in the image are recognized, and therefore the semantic consistency and structural integrity of the image are better kept in the data enhancement process. A dynamic interpolation strategy is introduced, parameters and modes of enhancement operation are adaptively adjusted according to segmentation features and context information of an image, and diversity and rationality of enhanced data are realized under the condition that original data distribution is not changed as far as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to an image fusion classification method and device based on segmentation-guided data augmentation. Background Art

[0002] With the development of deep learning technology, image classification tasks have been widely applied in the field of computer vision and occupy a core position, and are widely used in many practical application scenarios such as autonomous driving, medical image analysis, and face recognition. Deep neural networks, with their powerful feature extraction capabilities and end-to-end learning mechanisms, have significantly improved the accuracy and efficiency of image classification. However, the performance improvement of deep learning models largely depends on high-quality and diverse training data, enabling them to maintain a high classification accuracy when facing unseen data. In practical applications, obtaining a large amount of high-quality labeled data often faces many challenges. On the one hand, the data collection and annotation process is usually time-consuming and laborious, especially in fields that require professional knowledge, and the cost of obtaining high-quality labeled data is even more expensive. On the other hand, the distribution of training data is often uneven, and the data of some categories may be very scarce, resulting in significant deficiencies in the performance of the model on these categories. In addition, with the popularization of resource-constrained platforms such as mobile and embedded devices, the application demand for lightweight models (such as MobileNet, SqueezeNet, etc.) is increasing day by day. However, due to the limitations of the number of parameters and computational complexity, their learning ability is relatively weak, and they are easily negatively affected by the scarcity and uneven distribution of training data, resulting in insufficient generalization ability of the model and affecting the actual application effect.

[0003] Data augmentation technology, as an important means to improve model performance, has received extensive attention. Traditional data augmentation methods mainly include image flipping, cropping, rotation, scaling, color transformation, etc. By performing simple geometric and color transformations on the original image, the training data set is expanded, thus alleviating the data scarcity problem to a certain extent. However, traditional data augmentation methods often ignore the semantic information and structural features of the image, which may lead to deviations in semantic consistency and structural integrity of the augmented image. In recent years, emerging generative data augmentation methods, such as AutoAugment, Adversarial Example Generation, etc., have significantly improved data diversity and model robustness by learning the data distribution or generating new samples. However, these generative methods usually involve a large computational overhead, especially in the training stage, and additional computational resources are required to generate and process the augmented data. At the same time, the generative methods may introduce a shift in the data distribution in some cases, that is, the generated data may deviate from the distribution of the real data, resulting in the performance of the model in the real environment not necessarily being improved, and may even decline.

[0004] The current lightweight image classification algorithms have achieved good performance, but they perform poorly when the data sample size is small and unevenly distributed. Therefore, how to improve the classification accuracy and generalization ability of the model by improving the problems of small data sample size and uneven distribution when the lightweight model or computing resources are limited.

[0005] With the application of data augmentation technology in the field of computer vision, data augmentation methods have shifted from traditional methods to various advanced generative methods, and have made certain progress in the field of image classification.

[0006] First, there are the problems of data scarcity and uneven distribution. Generative data augmentation methods sometimes generate samples that are inconsistent with the real data distribution, resulting in a decrease in the prediction performance of the model in practical applications. In addition, fixed augmentation strategies cannot be adjusted according to the dynamic changes of data, which limits the adaptability of the model under different data distributions, thereby affecting the generalization ability and robustness of the model.

[0007] Secondly, the semantic consistency of the enhanced data is insufficient. Traditional data augmentation methods mainly focus on the geometric and color attribute transformations of images, ignoring the semantic information and structural features of images. Excessive geometric transformation or color transformation may lead to the loss or distortion of key features, destroy the overall semantic consistency and structural integrity of the image, and may generate invalid or noisy data, thereby affecting the model's correct understanding of the image content and classification performance.

[0008] Finally, the computational overhead and overfitting risk of the model. Although advanced data augmentation methods can significantly increase data diversity, their complex computational process and high computing resource requirements limit their widespread application in real-time applications and resource-constrained devices, and it is difficult to balance the amount of augmented data with the improvement of model performance. Summary of the invention

[0009] In order to solve the technical problem of how to use data enhancement, image segmentation and fusion technology in the prior art to improve the image classification accuracy in a lightweight model with less data volume and limited resources, the embodiment of the present invention provides an image fusion classification method and device based on segmentation-guided data enhancement. The technical solution is as follows:

[0010] On the one hand, an image fusion classification method based on segmentation-guided data enhancement is provided, the method is implemented by an image fusion classification device, and the method includes:

[0011] S1. Obtain a training data set; wherein the training data set includes multiple images of multiple categories.

[0012] S2. For each image in the training dataset, use the FastSAM network for segmentation to obtain the segmentation map corresponding to each image; the segmentation map contains the key semantic regions and structural features of the image.

[0013] S3. Calculate the cosine similarity between each image and its corresponding segmentation map, select the balanced segmentation map according to the cosine similarity, and construct the segmentation dataset based on the balanced segmentation map.

[0014] S4. Construct the distance norm metric and information gain, determine the number of enhanced segmentation maps according to the distance norm metric and information gain, select the enhanced images from the segmentation dataset according to the number of enhanced segmentation maps, and fuse the enhanced images with the training dataset to obtain the enhanced dataset.

[0015] S5. Adopt ResNet50 as the classification model, and train the classification model according to the enhanced dataset to obtain the trained classification model.

[0016] S6. Obtain the image to be classified, input the image to be classified into the trained classification model, and obtain the image classification result.

[0017] Optionally, calculating the cosine similarity between each image and its corresponding segmentation map in S3, selecting the balanced segmentation map according to the cosine similarity, and constructing the segmentation dataset based on the balanced segmentation map includes:

[0018] S31. Calculate the cosine similarity between each image and its corresponding segmentation map through the following formula (1):

[0019] (1)

[0020] In the formula, represents the cosine similarity, represents the th image in the training dataset, represents the th segmentation map corresponding to the image.

[0021] S32. Divide the segmentation maps into under-segmented maps, balanced segmentation maps, and over-segmented maps according to a preset cosine similarity threshold, and select the balanced segmentation maps to construct the segmentation dataset.

[0022] Optionally, constructing the distance norm metric and information gain in S4, and determining the number of enhanced segmentation maps according to the distance norm metric and information gain includes:

[0023] S41. Construct the distance norm metric.

[0024] S42. Construct the information gain.

[0025] S43. Construct the loss metric according to the distance norm metric and information gain.

[0026] S44. Minimize the loss metric and determine the number of augmented segmentation maps.

[0027] Optionally, the distance norm metric in S41 is as shown in the following formula (2):

[0028] (2)

[0029] In the formula, represents the distance norm metric, represents the number of augmented images, represents the average feature vector of the augmented dataset, represents the average feature vector of the training dataset, represents the standard deviation vector of the augmented dataset, represents the standard deviation vector of the training dataset.

[0030] Optionally, the information gain in S42 is as shown in the following formula (3):

[0031] (3)

[0032] In the formula, represents the information gain function, represents the number of augmented images, represents the first hyperparameter for adjustment, represents the size of the training dataset, represents the second hyperparameter for adjustment.

[0033] Optionally, the loss metric in S43 is as shown in the following formula (4):

[0034] (4)

[0035] In the formula, represents the loss metric, represents the number of augmented images, represents the third hyperparameter, represents the distance norm metric, represents the fourth hyperparameter, represents the information gain function.

[0036] Optionally, in S4, select augmented images from the segmentation dataset according to the number of augmented segmentation maps, and fuse the augmented images with the training dataset to obtain an augmented dataset, including:

[0037] Arrange the balanced segmentation maps in the segmentation dataset in descending order according to the cosine similarity, select augmented images from the arrangement according to the number of augmented segmentation maps, and fuse the augmented images with the corresponding images in the training dataset to obtain an augmented dataset.

[0038] On the other hand, an image fusion classification device based on segmentation-guided data augmentation is provided. This device is applied to an image fusion classification method based on segmentation-guided data augmentation. The device includes:

[0039] An acquisition module for acquiring a training data set; wherein, the training data set includes multiple images of multiple categories.

[0040] A segmentation module for segmenting each image in the training data set using the FastSAM network to obtain a segmentation map corresponding to each image; the segmentation map contains the key semantic regions and structural features of the image.

[0041] A segmentation data set construction module for calculating the cosine similarity between each image and the corresponding segmentation map, selecting balanced segmentation maps according to the cosine similarity, and constructing a segmentation data set according to the balanced segmentation maps.

[0042] An augmented data set construction module for constructing a distance norm metric and an information gain, determining the number of segmentation map augmentations according to the distance norm metric and the information gain, selecting augmented images from the segmentation data set according to the number of segmentation map augmentations, and fusing the augmented images with the training data set to obtain an augmented data set.

[0043] A training module for using ResNet50 as a classification model and training the classification model according to the augmented data set to obtain a trained classification model.

[0044] A classification module for acquiring an image to be classified, inputting the image to be classified into the trained classification model, and obtaining an image classification result.

[0045] Optionally, the segmentation data set construction module is further configured to:

[0046] S31. Calculate the cosine similarity between each image and the corresponding segmentation map through the following formula (1):

[0047] (1)

[0048] In the formula, represents the cosine similarity, represents the th image in the training data set, represents the th segmentation map corresponding to the

[0049] S32. Divide the segmentation maps into under-segmented maps, balanced segmentation maps, and over-segmented maps according to a preset cosine similarity threshold, and select the balanced segmentation maps to construct a segmentation data set.

[0050] Optionally, the enhanced dataset construction module is further configured to:

[0051] S41. Construct a distance norm metric.

[0052] S42. Construct an information gain.

[0053] S43. Construct a loss metric based on the distance norm metric and the information gain.

[0054] S44. Minimize the loss metric to determine the number of enhanced segmentation maps.

[0055] Optionally, the distance norm metric is as shown in the following formula (2):

[0056] (2)

[0057] In the formula, represents the distance norm metric, represents the number of enhanced images, represents the average feature vector of the enhanced dataset, represents the average feature vector of the training dataset, represents the standard deviation vector of the enhanced dataset, represents the standard deviation vector of the training dataset.

[0058] Optionally, the information gain is as shown in the following formula (3):

[0059] (3)

[0060] In the formula, represents the information gain function, represents the number of enhanced images, represents the first hyperparameter for adjustment, represents the size of the training dataset, represents the second hyperparameter for adjustment.

[0061] Optionally, the loss metric is as shown in the following formula (4):

[0062] (4)

[0063] In the formula, represents the loss metric, represents the number of enhanced images, represents the third hyperparameter, represents the distance norm metric, represents the fourth hyperparameter, represents the information gain function.

[0064] Optionally, the enhanced dataset construction module is further configured to:

[0065] Arrange the balanced segmentation maps in the segmented dataset in descending order according to the cosine similarity, select the enhanced images from the arrangement according to the number of enhanced segmentation maps, and fuse the enhanced images with the corresponding images in the training dataset to obtain an enhanced dataset.

[0066] On the other hand, provided is an image fusion classification device, including: a processor; a memory storing computer-readable instructions thereon, and when the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned image fusion classification method based on segmentation-guided data augmentation is implemented.

[0067] On the other hand, provided is a computer-readable storage medium storing at least one instruction, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned image fusion classification method based on segmentation-guided data augmentation.

[0068] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:

[0069] In the embodiments of the present invention, in the image classification problem, the classification performance depends to a large extent on the quality and diversity of the training data. In the case of scarce data, how to improve the generalization ability of the model and avoid overfitting without destroying the original data distribution is the core issue. Currently, traditional data augmentation methods such as flipping, random cropping, and changing color contrast are mainly used to solve the problem of small data volume. These methods can achieve good results when the data is sufficient and can increase data diversity to a certain extent. However, there are still limitations in the classification problem of less datasets. The present invention believes that the segmentation maps generated by the segmentation network can further extract the structural and semantic features of the image, generate diverse enhanced data, help the model learn richer features, can reduce the ineffective augmentation that may be caused by traditional augmentation methods, effectively expand the training dataset, and alleviate the problem of insufficient data.

[0070] Since the data generated by the segmentation network also has the risk of introducing ineffective augmented data and noise, the present invention further improves the quality of the data augmentation data by measuring the consistency between the augmented data and the original data distribution through the distance norm, ensuring that the augmented data does not deviate from the semantic distribution of the original data, and further reducing the overfitting risk. At the same time, an information gain optimization strategy is adopted to evaluate the value of the augmented data. By dynamically adjusting the number of augmented data, it is ensured that the generated augmented data has high information content, and only the segmentation maps that are most helpful for improving the model performance are added to the training set, avoiding the introduction of redundant data, thereby optimizing the model performance in the case of scarce data.

[0071] Through the optimization of cosine similarity and information gain, the present invention can intelligently select high-quality augmented data, reducing the workload of manual screening.

[0072] The performance of the present invention is significantly improved on multiple classical image classification models. The Top-1 accuracy of ResNet-50 is increased from 77.02% to 81.72%, and the Top-1 accuracy of RepViT is increased from 85.61% to 87.03%. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0074] Figure 1 It is a flowchart of an image fusion classification method based on segmentation-guided data augmentation provided by an embodiment of the present invention;

[0075] Figure 2 It is an overall structure diagram of an image fusion classification method based on segmentation-guided data augmentation provided by an embodiment of the present invention;

[0076] Figure 3 It is a schematic flowchart of an image fusion classification method based on segmentation-guided data augmentation provided by an embodiment of the present invention;

[0077] Figure 4 It is a FastSAM segmentation flowchart provided by an embodiment of the present invention;

[0078] Figure 5 It is a cosine similarity filtering flowchart provided by an embodiment of the present invention;

[0079] Figure 6 It is a flowchart of the best data selection strategy provided by an embodiment of the present invention;

[0080] Figure 7 It is a block diagram of an image fusion classification device based on segmentation-guided data augmentation provided by an embodiment of the present invention;

[0081] Figure 8 It is a schematic structural diagram of an image fusion classification device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0082] The following will describe the technical solutions in the present invention with reference to the drawings.

[0083] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0084] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "Of", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.

[0085] In the embodiments of the present invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.

[0086] To make the technical problems to be solved, technical solutions and advantages of the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0087] The embodiments of the present invention provide an image fusion classification method based on segmentation-guided data augmentation. This method can be implemented by an image fusion classification device, which can be a terminal or a server. As Figure 1 shown in the flowchart of the image fusion classification method based on segmentation-guided data augmentation, the processing flow of this method can include the following steps:

[0088] S1. Obtain a training data set; wherein, the training data set includes multiple images of multiple categories.

[0089] In a feasible implementation, the image classification data set adopted by the present invention is the Mini-ImageNet data set, which has a total of 100 categories, 48 images in each category, for a total of 48,000 images. The image size is 224×224. This data set is widely used for the evaluation of image classification tasks.

[0090] The present invention aims to provide a segmentation-based image fusion classification data augmentation strategy SegAug as Figure 2 shown, to solve the problems of data sparsity and inconsistent distribution.

[0091] S2. For each image in the training data set, use the FastSAM network for segmentation to obtain the segmentation map corresponding to each image; the segmentation map contains the key semantic regions and structural features of the image.

[0092] In a feasible implementation, Figure 3 It is a schematic flowchart of the method according to the embodiment of the present invention. The whole process can be divided into four steps: image segmentation, optimal data selection, model training, and model inference.

[0093] Specifically, use the FastSAM model to segment the input image to generate a segmentation map , as shown in the following formula (1). This process ensures the high quality and structural consistency of the image data. The segmented dataset is obtained, providing the basic data for the subsequent enhancement process.

[0094] (1)

[0095] FastSAM (Fast Segment Anything Model) is a lightweight segmentation network that can effectively extract the structural features of the key regions of the image, provide semantic guidance for subsequent enhancement, avoid the destruction of semantic structures by traditional enhancement methods, and distinguish the foreground from the background. Compared with other traditional SAM (Segment Anything Model) frameworks, FastSAM is more efficient and can greatly reduce the data preprocessing time. The FastSAM segmentation process is as Figure 4 shown. By generating high-quality segmentation feature maps, it can help the model identify the key structures in the image, ensure that the structural changes in the data enhancement process can be reasonably retained, and maintain the semantic consistency of the original data.

[0096] In the embodiment of the present invention, in the data preprocessing stage, a segmentation model is used for feature extraction to extract the key semantic regions and structural features in the image. These segmentation information can help identify the important objects and background regions in the image, providing guidance for subsequent data enhancement operations. Using FastSAM to replace the traditional segmentation model reduces the preprocessing calculation overhead and adapts to resource-constrained devices. Generate enhanced samples preferentially in data-sparse regions, reduce redundancy in dense regions, and improve the data enhancement efficiency.

[0097] S3. Calculate the cosine similarity between each image and the corresponding segmentation map, select a balanced segmentation map according to the cosine similarity, and construct a segmentation dataset according to the balanced segmentation map.

[0098] In a feasible implementation, filter the segmentation maps based on the cosine similarity, and use the Information Gain Optimization (IGO) method for adaptive data enhancement. Select a balanced segmentation map according to the similarity value, and eliminate over-segmented and under-segmented images.

[0099] Optionally, the above step S3 may include the following steps S31 - S32:

[0100] S31. Determine the similarity between the image and the segmentation map by calculating the cosine similarity.

[0101] Specifically, calculate the cosine similarity between each image and the corresponding segmentation map by the following formula (2):

[0102] (2)

[0103] In the formula, represents the cosine similarity, represents the th image in the training dataset, represents the th segmentation map corresponding to the image.

[0104] S32. Divide the segmentation maps into under - segmented maps, balanced segmentation maps, and over - segmented maps according to a preset cosine similarity threshold, and select the balanced segmentation maps to construct a segmentation dataset.

[0105] In a feasible implementation manner, according to the calculation results, select the balanced segmentation maps, that is, those segmentation maps that are highly consistent with the structure of the original image, to ensure that the structural changes of the enhanced data conform to the semantic features of the original data. For images with a high cosine similarity, mark them as under - segmented and eliminate them; for images with a low similarity, mark them as over - segmented and also eliminate them. The cosine similarity filtering process is as Figure 5 shown.

[0106] By calculating the cosine similarity, the effect of adaptive data augmentation is ensured, ensuring that the enhanced data can both increase the diversity of samples and maintain consistency with the original data, and avoiding the negative impact of over - or under - enhanced data on the training process. Dynamically adjust the augmentation parameters through the context information of the segmentation map and the original image to achieve scene adaptation of the augmentation method.

[0107] In the embodiment of the present invention, a dynamic interpolation strategy is proposed to adjust the parameters of the data augmentation operation according to the segmentation features, and the cosine similarity screening is used to ensure a balance between the diversity and distribution consistency of the enhanced data.

[0108] S4. Construct a distance norm metric and information gain, determine the number of enhanced segmentation maps according to the distance norm metric and information gain, select enhanced images from the segmentation dataset according to the number of enhanced segmentation maps, and fuse the enhanced images with the training dataset to obtain an enhanced dataset.

[0109] In a feasible implementation, the number of interpolation samples is dynamically adjusted through an information gain optimization strategy, and enhanced data is preferentially generated in sparse regions of the data to ensure the uniformity of the data distribution and avoid generating too much redundant data in dense regions. The processed dataset is used for subsequent steps.

[0110] Optionally, constructing the distance norm metric and information gain in S4, and determining the number of segmentation map enhancements according to the distance norm metric and information gain may include the following steps S41 - S44:

[0111] S41. Construct the distance norm metric.

[0112] In a feasible implementation, a distribution consistency metric is used to first define a distance norm metric to measure the distribution consistency between the original dataset and the enhanced dataset. This metric combines the mean and variance between the input image and the segmentation image, as shown in the following formula (3):

[0113] (3)

[0114] In the formula, represents the distance norm metric, represents the number of enhanced images, represents the average feature vector of the enhanced dataset, represents the average feature vector of the training dataset, represents the standard deviation vector of the enhanced dataset, represents the standard deviation vector of the training dataset. The distance norm evaluates the semantic similarity between datasets by considering the mean and standard deviation differences. The smaller it is, the higher the consistency between the enhanced dataset and the original dataset, which is crucial for maintaining model performance.

[0115] S42. Construct the information gain.

[0116] In a feasible implementation, an information gain optimization is adopted to adjust the number of enhanced data to ensure that the added data can effectively improve the model performance. The information gain function is as follows:

[0117] (4)

[0118] In the formula, represents the information gain function, represents the number of enhanced images, represents the first hyperparameter for adjustment, represents the size of the training dataset, Represents the second hyperparameter for adjustment. Information gain optimization ensures the optimization of the amount of data augmentation, avoids introducing noise with excessive samples, and further optimizes the training process by balancing performance improvement and overfitting risk.

[0119] S43. Construct a loss metric based on the distance norm metric and information gain.

[0120] Optionally, the distance norm metric and information gain are combined into a new loss metric:

[0121] (5)

[0122] wherein, represents the loss metric, represents the number of augmented images, represents the third hyperparameter, represents the fourth hyperparameter, which respectively controls the weights of distribution consistency and information gain in the final loss. represents the distance norm metric, represents the information gain function.

[0123] S44. Minimize the loss metric to determine the number of augmented segmentation maps.

[0124] In a feasible implementation, by minimizing this combined loss, it is ensured that the augmented data not only maintains the distribution consistency of the original data but also maximizes its contribution to the model performance.

[0125] The present invention designs a distance norm metric to quantify the distribution difference between the augmented data and the original data from the mean and variance dimensions, and constraints the rationality of the generated data. An information gain optimization function is proposed to dynamically balance the number of augmented samples and the model performance gain, and avoid the overfitting risk caused by redundant data. By jointly optimizing the distribution consistency and information gain through a combined loss function, the optimal selection of augmented samples is achieved.

[0126] Optionally, in S4, the augmented images are selected from the segmentation dataset according to the number of augmented segmentation maps, and the augmented images are fused with the training dataset to obtain an augmented dataset, including:

[0127] The balanced segmentation maps in the segmentation dataset are sorted in descending order according to the cosine similarity, with a step size of 50. The segmentation maps are traversed in turn and fused with the original image, and the optimal number of augmented segmentation map samples is determined according to the loss metric to obtain the augmented dataset.

[0128] In a feasible implementation, the optimal data augmentation samples are selected and fused with the original image data to ensure that the augmented dataset is both diverse and consistent with the original data distribution, thereby improving the performance and efficiency of the model. The best data selection process is as Figure 6 shown.

[0129] In the embodiment of the present invention, an optimal data selection strategy is proposed. First, a distribution consistency evaluation index is used to maintain data distribution consistency. Second, an information gain optimization index is designed to dynamically adjust the number of augmented samples. Finally, a comprehensive loss function is used to select the optimal number of augmented samples while maintaining data distribution consistency and avoid overfitting caused by excessive augmentation.

[0130] S5. Use ResNet50 as the classification model, and train the classification model according to the augmented dataset to obtain a trained classification model.

[0131] In a feasible implementation, the data processed in steps S2, S3, and S4 is input into the classification model, and cross-validation and hyperparameter tuning are used to ensure that the model performance reaches the best. The best model parameters are saved during the training process for subsequent inference.

[0132] Specifically, a classic classification model and the latest classification models are selected for experiments, namely ResNet50, RepViT, MobileViT, EfficientFormerV2, and FastViT. In the present invention, ResNet50 is used as the basic classification model. This model consists of 50 layers of residual networks and can effectively perform deep-level image feature extraction. During the training process, the model performance is evaluated through cross-validation, and the Adam optimizer is used for gradient descent optimization.

[0133] Furthermore, first, the dataset is divided into a training set and a validation set; second, after the training set data is processed through steps S2, S3, and S4, an augmented dataset is obtained and used as the input of the classification model; finally, hyperparameter tuning is performed, and hyperparameters such as the learning rate, batch size, and number of training epochs are adjusted through a random search method to ensure that the model can achieve the best performance on the training data. The model is trained using the training set data and the model is saved.

[0134] Furthermore, inference is performed using the best model obtained from training. The trained model is loaded from the specified path to ensure that the model parameters are correct. The preprocessed test data is input into the trained model for classification. By comparing the prediction results with the true labels, the classification accuracy and other evaluation metrics are calculated to evaluate the inference performance of the model.

[0135] For the details of model parameters, when using the ResNet-50 model, the batch size is set to 160, the number of epochs is set to 100, and the learning rate is 0.1; when using the RepViT model, the batch size is set to 256, the number of epochs is set to 300, and the learning rate is 0.001. All other parameters are fixed and remain consistent in the experiment.

[0136] This invention adopts the Pytorch framework, and the specific experiment is carried out on NVIDIA GeForce RTX 3090 for model training and evaluation.

[0137] S6. Obtain the image to be classified, input the image to be classified into the trained classification model, and obtain the image classification result.

[0138] The purpose of this invention is to use the image segmentation method and information gain optimization strategy to increase the sample size and diversity of image classification data. The cosine similarity between the segmented image and the original image is used as the constraint condition for the enhancement operation to ensure the semantic consistency between the generated data and the original image, and the distribution consistency and information gain optimization are adopted to dynamically adjust the enhanced data, so as to improve the generalization ability and prediction accuracy of the model.

[0139] The existing patent (CN116740482A) discloses an underwater small sample augmentation method based on image segmentation fusion. After the samples to be fused are segmented to extract single sample information, the background color hue is converted to approximately calculate the hue of the underwater background to be fused, and Poisson fusion is used to generate the image after adding samples, which can achieve the effect of enriching the recognition target information and balancing the training data. The main differences between this invention and this patent are as follows: 1) The FastSAM network is used to segment the image to extract the key semantic regions and structural features in the image; 2) A dynamic enhancement sample and fusion strategy is proposed, and the parameters of the data enhancement operation are adjusted according to the segmentation features, and cosine similarity screening is adopted to ensure a balance between diversity and distribution consistency of the enhanced data; 3) The accuracy is significantly improved on classification models such as RepViT.

[0140] In the embodiments of the present invention, in the image classification problem, the classification performance depends to a large extent on the quality and diversity of the training data. In the case of scarce data, how to improve the generalization ability of the model and avoid overfitting without destroying the original data distribution is the core issue. Currently, for the problem of small data volume, traditional data augmentation methods are mainly used, such as flipping, random cropping, changing color contrast, etc. These methods can achieve good results when the data is sufficient and can increase data diversity to a certain extent. However, there are still limitations in the classification problem of less data sets. The present invention believes that the segmentation map generated by the segmentation network can further extract the structural and semantic features of the image, generate diverse augmented data, help the model learn richer features, reduce the ineffective augmentation that may be caused by traditional augmentation methods, effectively expand the training data set, and alleviate the problem of insufficient data.

[0141] Since the data generated by the segmentation network also has the risk of introducing ineffective augmented data and noise, in order to further improve the quality of the augmented data, the present invention measures the consistency between the augmented data and the original data distribution through the distance norm, ensures that the augmented data does not deviate from the semantic distribution of the original data, and further reduces the overfitting risk. At the same time, an information gain optimization strategy is adopted to evaluate the value of the augmented data. By dynamically adjusting the number of augmented data, it is ensured that the generated augmented data has high information content. Only the segmentation maps that are most helpful for improving the model performance are added to the training set, avoiding the introduction of redundant data, thereby optimizing the model performance in the case of scarce data.

[0142] Through cosine similarity and information gain optimization, the present invention can intelligently select high-quality augmented data and reduce the workload of manual screening.

[0143] The performance of the present invention is significantly improved on multiple classic image classification models. The Top-1 accuracy of ResNet-50 is increased from 77.02% to 81.72%, and the Top-1 accuracy of RepViT is increased from 85.61% to 87.03%.

[0144] Figure 7 It is a block diagram of an image fusion classification device based on segmentation-guided data augmentation shown according to an exemplary embodiment. This device is used for the image fusion classification method based on segmentation-guided data augmentation. Referring to Figure 7 As shown in the figure, this device includes an acquisition module 310, a segmentation module 320, a segmentation data set construction module 330, an augmented data set construction module 340, a training module 350, and a classification module 360. Among them:

[0145] The acquisition module 310 is used to acquire the training data set; among them, the training data set includes multiple images of multiple categories.

[0146] The segmentation module 320 is used to segment each image in the training dataset using the FastSAM network to obtain a segmentation map corresponding to each image; the segmentation map contains the key semantic regions and structural features of the image.

[0147] The segmentation dataset construction module 330 is used to calculate the cosine similarity between each image and the corresponding segmentation map, select balanced segmentation maps according to the cosine similarity, and construct a segmentation dataset based on the balanced segmentation maps.

[0148] The enhanced dataset construction module 340 is used to construct a distance norm metric and information gain, determine the number of enhanced segmentation maps according to the distance norm metric and information gain, select enhanced images from the segmentation dataset according to the number of enhanced segmentation maps, and fuse the enhanced images with the training dataset to obtain an enhanced dataset.

[0149] The training module 350 is used to adopt ResNet50 as a classification model and train the classification model according to the enhanced dataset to obtain a trained classification model.

[0150] The classification module 360 is used to obtain the image to be classified, input the image to be classified into the trained classification model, and obtain the image classification result.

[0151] In the embodiments of the present invention, in the image classification problem, the classification performance largely depends on the quality and diversity of the training data. In the case of scarce data, how to improve the generalization ability of the model and avoid overfitting without destroying the original data distribution is the core issue. Currently, traditional data augmentation methods, such as flipping, random cropping, and changing color contrast, are mainly used to address the problem of small data volume. Such methods can achieve good results when the data is sufficient and can increase data diversity to a certain extent. However, there are still limitations in the classification problem of less datasets. The present invention believes that the segmentation maps generated by the segmentation network can further extract the structural and semantic features of the image, generate diverse enhanced data, help the model learn richer features, can reduce the ineffective augmentation that may be caused by traditional augmentation methods, effectively expand the training dataset, and alleviate the problem of insufficient data.

[0152] Since the data generated by the segmentation network also has the risk of introducing invalid augmented data and noise, in order to further improve the quality of the data augmentation in the present invention, the consistency between the augmented data and the original data distribution is measured by the distance norm to ensure that the augmented data does not deviate from the semantic distribution of the original data, further reducing the risk of overfitting. At the same time, an information gain optimization strategy is adopted to evaluate the value of the augmented data. By dynamically adjusting the quantity of the augmented data, it is ensured that the generated augmented data has high information content, and only the segmentation maps that are most helpful for improving the model performance are added to the training set, avoiding the introduction of redundant data, thereby optimizing the model performance in the case of data scarcity.

[0153] Through cosine similarity and information gain optimization, the present invention can intelligently select high-quality augmented data, reducing the workload of manual screening.

[0154] The performance of the present invention has been significantly improved on multiple classical image classification models. The Top-1 accuracy of ResNet-50 has increased from 77.02% to 81.72%, and the Top-1 accuracy of RepViT has increased from 85.61% to 87.03%.

[0155] Figure 8 It is a schematic structural diagram of an image fusion classification device provided by an embodiment of the present invention. As Figure 8 shown, the image fusion classification device may include the above-mentioned Figure 7 image fusion classification device based on segmentation-guided data augmentation as shown. Optionally, the image fusion classification device 410 may include a first processor 2001.

[0156] Optionally, the image fusion classification device 410 may further include a memory 2002 and a transceiver 2003.

[0157] Among them, the first processor 2001 is connected to the memory 2002 and the transceiver 2003, such as through a communication bus.

[0158] Next, in combination with Figure 8 a specific introduction to each component of the image fusion classification device 410 will be given:

[0159] Among them, the first processor 2001 is the control center of the image fusion classification device 410, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0160] Optionally, the first processor 2001 can execute various functions of the image fusion classification device 410 by running or executing software programs stored in the memory 2002 and invoking data stored in the memory 2002.

[0161] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 8 the CPU0 and CPU1 shown in

[0162] In a specific implementation, as an embodiment, the image fusion classification device 410 may also include multiple processors, such as Figure 8 the first processor 2001 and the second processor 2004 shown in

[0163] Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0164] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 8 not shown) of the image fusion classification device 410. The embodiments of the present invention do not make specific limitations on this.

[0165] The transceiver 2003 is used to communicate with a network device or with a terminal device.

[0166] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 8 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.

[0167] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 8 not shown) of the image fusion classification device 410. The embodiments of the present invention do not make specific limitations on this.

[0168] It should be noted that Figure 8 the structure of the image fusion classification device 410 shown in does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0169] In addition, the technical effects of the image fusion classification device 410 may refer to the technical effects of the image fusion classification method based on segmentation-guided data augmentation described in the above method embodiments, and will not be elaborated here.

[0170] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0171] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0172] The above-described embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above-described embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0173] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.

[0174] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0175] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0176] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0177] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0178] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0179] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0180] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0181] When the above-described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0182] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An image fusion classification method based on segmentation-guided data augmentation, characterized in that The method includes: S1. Obtain a training data set; wherein, the training data set includes multiple images of multiple categories; S2. For each image in the training data set, use the FastSAM network for segmentation to obtain a segmentation map corresponding to each image; the segmentation map contains the key semantic regions and structural features of the image; S3. Calculate the cosine similarity between each image and the corresponding segmentation map, select a balanced segmentation map according to the cosine similarity, and construct a segmentation data set according to the balanced segmentation map; S4. Construct a distance norm metric and an information gain, determine the number of segmentation map enhancements according to the distance norm metric and the information gain, select enhanced images from the segmentation data set according to the number of segmentation map enhancements, and fuse the enhanced images with the training data set to obtain an enhanced data set; S5. Use ResNet50 as a classification model, and train the classification model according to the enhanced data set to obtain a trained classification model; S6. Obtain an image to be classified, input the image to be classified into the trained classification model, and obtain an image classification result.

2. The image fusion classification method based on segmentation-guided data augmentation according to claim 1, wherein, The calculating the cosine similarity between each image and the corresponding segmentation map in S3, selecting a balanced segmentation map according to the cosine similarity, and constructing a segmentation data set according to the balanced segmentation map includes: S31. Calculate the cosine similarity between each image and the corresponding segmentation map through the following formula (1): (1) In the formula, represents the cosine similarity, represents the th image in the training dataset, represents the th segmentation map corresponding to the image; S32. Divide the segmentation maps into under-segmented maps, balanced segmentation maps, and over-segmented maps according to a preset cosine similarity threshold, and select the balanced segmentation maps to construct a segmentation data set.

3. The image fusion classification method based on segmentation-guided data augmentation according to claim 1, wherein, The constructing a distance norm metric and an information gain in S4, and determining the number of segmentation map enhancements according to the distance norm metric and the information gain includes: S41. Construct a distance norm metric; S42. Construct an information gain; S43. Construct a loss metric according to the distance norm metric and the information gain; S44. Minimize the loss metric to determine the number of segmentation map enhancements.

4. The image fusion classification method based on segmentation-guided data augmentation according to claim 3, wherein, The distance norm metric in S41 is shown in the following formula (2): (2) wherein, represents the distance norm metric, represents the number of enhanced images, represents the average feature vector of the enhanced dataset, represents the average feature vector of the training dataset, represents the standard deviation vector of the enhanced dataset, represents the standard deviation vector of the training dataset.

5. The image fusion classification method based on segmentation-guided data augmentation according to claim 3, wherein The information gain in S42 is shown in the following formula (3): (3) In the formula, represents the information gain function, represents the number of enhanced images, represents the first hyperparameter for adjustment, represents the size of the training dataset, represents the second hyperparameter for adjustment.

6. The image fusion classification method based on segmentation-guided data augmentation according to claim 3, wherein The loss metric in S43 is shown in the following formula (4): (4) Wherein, represents the loss metric, represents the number of enhanced images, represents the third hyperparameter, represents the distance norm metric, represents the fourth hyperparameter, represents the information gain function.

7. The image fusion classification method based on segmentation-guided data augmentation according to claim 1, wherein The selecting enhanced images from the segmentation data set according to the number of segmentation map enhancements in S4, and fusing the enhanced images with the training data set to obtain an enhanced data set includes: Arrange the balanced segmentation maps in the segmentation data set in descending order according to the cosine similarity, select enhanced images from the arrangement according to the number of segmentation map enhancements, and fuse the enhanced images with the corresponding images in the training data set to obtain an enhanced data set.

8. An image fusion classification device based on segmentation-guided data augmentation, which is used to implement the image fusion classification method based on segmentation-guided data augmentation according to any one of claims 1-7, characterized in that The device includes: An acquisition module, configured to acquire a training data set; wherein, the training data set includes multiple images of multiple categories; A segmentation module, configured to use the FastSAM network to segment each image in the training data set to obtain a segmentation map corresponding to each image; the segmentation map contains the key semantic regions and structural features of the image; The segmentation dataset construction module is used to calculate the cosine similarity between each image and the corresponding segmentation map, select a balanced segmentation map according to the cosine similarity, and construct a segmentation dataset according to the balanced segmentation map; The enhanced dataset construction module is used to construct a distance norm metric and an information gain, determine the number of enhanced segmentation maps according to the distance norm metric and the information gain, select enhanced images from the segmentation dataset according to the number of enhanced segmentation maps, and fuse the enhanced images with the training dataset to obtain an enhanced dataset; The training module is used to adopt ResNet50 as a classification model and train the classification model according to the enhanced dataset to obtain a trained classification model; The classification module is used to obtain an image to be classified, input the image to be classified into the trained classification model, and obtain an image classification result.

9. An image fusion classification device, characterized in that, The image fusion classification device includes: A processor; A memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, the method described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image semantic segmentation method and device and computer equipment

    CN110163862A

  • Image segmentation method based on large model guidance

    CN118298169A

  • Image semantic segmentation method and system based on few samples

    CN119339381A

  • Multimodal intelligent dysphagia classification system and method based on ultrasonic image and electromyographic signal

    CN119357783A

  • Crop growth prediction method based on multi-source data fusion analysis

    CN119398284A