A Threshold Loss Optimization Method for Data Augmentation and Learnable Parameters Based on Small-Sample Learning
Through the data augmentation of small samples and the threshold loss optimization method of learnable parameters, the problems of high dependence on large pre-trained data sets in the prior art, high computational complexity, insufficient tasks to adapt to specific fields, single threshold optimization mechanism and high error detection rate are solved, and the model is efficiently trained and accurately predicted under small samples.
Patent Information
- Application Number
- CN202411013608.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-07-26
AI Technical Summary
The existing small sample learning methods have problems such as high dependence on large pre-trained data sets, high computational complexity, insufficient tasks to adapt to specific fields, insufficient support for small sample learning, relatively single threshold optimization mechanism and high error detection rate.
Using data augmentation based on small sample learning and threshold loss optimization methods for learnable parameters, the second-stage data augmentation technology and threshold loss optimization for learnable parameters is combined with deep learning neural network model to perform offline and online data augmentation of sample graphs, and the threshold loss function is adjusted through learning parameters to optimize model performance.
It significantly improves the robustness and generalization ability of the model under small sample conditions, reduces the false detection rate, simplifies the model training process, and improves prediction accuracy and engineering practicality.
Smart Images

Figure CN118799681B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sample enhancement, and specifically relates to a data enhancement method based on few-shot learning and a threshold loss optimization method for learnable parameters. Background Art
[0002] In today's era of rapid technological development, machine learning and deep learning technologies have become core tools in many fields. Especially in the fields of pattern recognition, image processing, and natural language processing, these technologies have made great progress. However, in some cases, due to the scarcity of data samples or class imbalance, traditional machine learning methods may face challenges, such as in the fields of medical diagnosis, financial risk control, and security detection.
[0003] In this context, few-shot learning has become a research direction that has attracted much attention. Few-shot learning aims to learn from very limited samples and generalize to new samples to address the problem of data scarcity. The research in this field focuses on how to build a robust model using a small number of labeled samples. In the field of few-shot learning, data enhancement and loss function optimization are two key technical points. The following is an introduction to these two aspects respectively:
[0004] (I) Data Enhancement
[0005] Commonly used data enhancement methods adopt few-shot learning methods, including transfer learning, meta-learning, etc.
[0006] (1) Transfer Learning
[0007] Transfer learning pre-trains a model on a large dataset and then fine-tunes it on the target small dataset. This method utilizes the features learned on the large dataset to enable the model to better adapt to new tasks. For example, when using a pre-trained model for few-shot object detection, the model can be fine-tuned to adapt to a specific small sample dataset.
[0008] The specific implementation method is as follows: Pre-train a deep neural network model (such as ResNet or VGG) on a large-scale image dataset (such as ImageNet), and then apply the pre-trained model to a specific small sample dataset. By fine-tuning the parameters of the model, it can be made to adapt to new tasks.
[0009] This method has the following problems: This method still requires a large pre-training dataset, and overfitting may occur during the fine-tuning process, especially when the distribution of the target small sample dataset is significantly different from that of the pre-training dataset.
[0010] (2) Meta-Learning
[0011] Meta-learning enables a model to quickly adapt to new tasks by training it. Typical meta-learning methods include MAML (Model-Agnostic Meta-Learning) and ProtoNet (Prototypical Networks). These methods train on multiple few-shot tasks, enabling the model to better generalize to new few-shot tasks.
[0012] The specific implementation methods are as follows: For example, the MAML method trains on multiple tasks to learn an initial set of parameters that can quickly adapt to new tasks with only minor updates. ProtoNet classifies by calculating the distance between samples and class prototypes and is suitable for few-shot scenarios.
[0013] This approach has the following problems: Meta-learning methods typically require large datasets of diverse tasks for training and have high requirements in terms of computational complexity and training time. Additionally, for few-shot tasks in specific domains, existing meta-learning methods may not be able to fully utilize prior knowledge in those domains.
[0014] It can be seen that existing few-shot learning methods suffer from problems such as high dependence on large pre-trained datasets, high computational complexity, and insufficient adaptation to specific domain tasks.
[0015] (2) Optimization of Loss Functions
[0016] In machine learning and deep learning, threshold loss optimization methods are widely used to improve the accuracy of classification and detection tasks. Traditional threshold loss optimization methods usually set fixed or simple dynamic thresholds and optimize the model by minimizing the loss function during training. The main traditional threshold loss optimization methods include the following:
[0017] (1) Fixed Threshold Method:
[0018] The fixed threshold method uses a fixed threshold to determine the classification or detection results of samples. This method lacks flexibility and is difficult to adapt to changes in different classes or samples.
[0019] (2) Dynamic Threshold Method:
[0020] The dynamic threshold method dynamically adjusts the threshold based on the characteristics of samples and the confidence of the model during training. Although more flexible than the fixed threshold, the adjustment mechanism of this method is usually relatively simple and cannot fully utilize the diversity and complexity of the training data.
[0021] (3) Existing Adaptive Threshold Optimization Methods:
[0022] Some advanced technologies have introduced an adaptive threshold adjustment mechanism to improve the accuracy of classification or detection by optimizing threshold parameters during model training. However, these methods mainly focus on optimizing the performance of a single task and still have deficiencies in dealing with multi-task, multi-scale problems, and high false detection rates.
[0023] Therefore, the existing threshold loss optimization methods have the following problems:
[0024] (1) Insufficient support for small sample learning: Most of the existing adaptive threshold optimization methods rely on a large amount of data for threshold adjustment, lacking support for small sample learning and being unable to generalize effectively in the case of scarce data. (2) Limitations of threshold optimization: The threshold optimization mechanism of the existing methods is relatively single, usually just adding a simple threshold term and lacking fine-tuning for different samples and tasks. (3) High false detection rate problem: In practical applications, the existing methods often face a high false detection rate, especially when dealing with complex scenarios and multi-scale object detection, this problem is particularly prominent.
[0025] In summary, the existing threshold loss optimization methods have significant deficiencies in supporting small sample learning, applying data augmentation techniques, the fineness of threshold optimization, and the false detection rate when dealing with complex scenarios. Summary of the Invention
[0026] Aiming at the defects existing in the prior art, the present invention provides a threshold loss optimization method based on data augmentation and learnable parameters for small sample learning, which can effectively solve the above problems.
[0027] The technical solution adopted by the present invention is as follows:
[0028] The present invention provides a threshold loss optimization method based on data augmentation and learnable parameters for small sample learning, including the following steps:
[0029] Step S1, establish an actual background image sample library G and a small sample image sample library S; multiple actual background images are stored in the actual background image sample library G, and each actual background image is represented as G i ; multiple small sample images are stored in the small sample image sample library S, and each small sample image is represented as S j ;
[0030] Step S2, traverse the small sample image sample library S, and for each small sample image S j traversed, based on the actual background images in the actual background image sample library G, perform one-stage actual background image offline augmentation on the small sample image S j , expand the small sample image S j , to obtain a sample image F after one-stage actual background image offline augmentation, and its label is lable(F);
[0031] Step S3: Traverse the small sample image library S, and for each small sample image S traversed j , perform one-stage offline enhancement of the solid-color background image to expand the small sample image S j , and obtain the sample image I after one-stage offline enhancement of the solid-color background image, with its label being lable(I);
[0032] Step S4: For the sample image F and the sample image I, perform two-stage fusion enhancement using formula (1) to obtain the fused and enhanced sample image Mix, with its label being lable(Mix):
[0033] Mix = λ·F + (1 - λ)·I (1)
[0034] lable(Mix) = λ·lable(F) + (1 - λ)·lable(I)
[0035] where λ is the mixing coefficient randomly sampled from the Beta distribution;
[0036] Step S5: Through steps S2 to S4, obtain a training sample set formed by multiple sample images F, sample images I, and sample images Mix;
[0037] Step S6: Establish a deep learning neural network model; the deep learning neural network model has learnable parameters p;
[0038] Step S7: Use the training sample set to train the deep learning neural network model to obtain a trained deep learning neural network model. The specific training method is as follows:
[0039] Step S7.1: Initialize the learnable parameter p as the identity matrix; the sample images F, sample images I, and sample images Mix each have an initial intensity;
[0040] Step S7.2: Initialize the initial value of the iteration number t as 1;
[0041] Step S7.3: Use the online data augmentation module to perform online data augmentation intensity processing on the sample images F, sample images I, and sample images Mix according to formula (2) to obtain the processed sample images F, sample images I, and sample images Mix. The method is as follows:
[0042]
[0043] where:
[0044] A t is the online data augmentation intensity of the sample images in the t-th round of training;
[0045] A0 is the initial intensity of the sample images;
[0046] t is the current iteration number;
[0047] T is the total number of iterations;
[0048] η t is the learning rate for the t-th round of training;
[0049] η0 is the initial learning rate;
[0050] p t is the value of the learnable parameter for the t-th round of training;
[0051] Step S7.4, input the sample graph F, sample graph I, and sample graph Mix processed in step S7.3 into the deep learning neural network model for training;
[0052] Step S7.5, let t = t + 1, and return to step S7.3, and continuously iterate and train like this until the training termination condition is reached.
[0053] Preferably, in step S2, perform one-stage offline enhancement of the actual background graph on the small sample graph S j Specifically:
[0054] Step S2.1, randomly select an actual background graph G i ;
[0055] Step S2.2, convert the actual background graph G i into the RGBA form to obtain the actual background graph G i (1), and read the size of the actual background graph G i (1);
[0056] Step S2.3, the small sample graph S j has a label lable(S j ), convert the small sample graph S j into the RGBA form to obtain the small sample graph S j (1), and read the size of the small sample graph S j (1);
[0057] Compare the size of the small sample graph S j (1) and the size of the actual background graph G i (1), so as to determine the transformable area in the actual background graph G i (1);
[0058] Step S2.4, convert the small sample graph S j (1) into an array form, and perform image scale change and image enhancement processing using an image transformation method to obtain the processed small sample graph S in array formj (1);
[0059] Step S2.5, convert the processed small sample image S in array form j (1) into a small sample image in the form of a picture object in RGBA form, called small sample image S j (2);
[0060] Step S2.6, use the channel parameter of the alpha channel of the actual background image G i (1) as the fusion parameter, and fuse the transformable region of the small sample image S j (2) and the actual background image G i (1), so as to integrate the small sample image S into the transformable region of the actual background image G i (1), obtain the fused sample image F, and update the label of the sample image F according to the label lable(S j ) of the small sample image S, so that the label lable(F) of the sample image F is the label lable(S j ); j ) j ;
[0061] Step S2.7, return to Step S2.1, and randomly select an actual background image again to perform offline enhancement processing on the small sample image S j , and repeat this process multiple times to achieve offline enhancement of the small sample image S j .
[0062] Preferably, in Step S2.4, the image transformation method is used for image scale change and image enhancement processing, specifically:
[0063] Use a random function to randomly determine the image scaling scale; perform scale transformation on the small sample image S in array form j (1), and then perform random enhancement processing, including flipping, row-column swapping, grid distortion, elastic transformation, rotation, random gamma, image mean filtering, random brightness, and adding Gaussian noise processing.
[0064] Preferably, in Step S3, perform one-stage solid-color background image offline enhancement on the small sample image S j , specifically:
[0065] Step S3.1, randomly select an actual background image G in the actual background image sample library G i , and read the size of the actual background image G i (1);
[0066] Step S3.2, initialize and generate a solid-color background canvas of the same size according to the size of the actual background image G i (1);
[0067] Step S3.3: Pre-define multiple solid background colors, including mainstream solid background colors and random solid background colors; randomly select one solid background color and fill it into the solid background canvas to generate a solid background image B0 in RGBA format.
[0068] Step S3.4: The small sample image S j with the label lable(S j ), convert the small sample image S j into RGBA format to obtain the small sample image S j (1), and read the size of the small sample image S j (1).
[0069] Step S3.5: Compare the size of the small sample image S j (1) and the size of the solid background image B0, so as to determine the transformable area in the solid background image B0.
[0070] Step S3.6: Convert the small sample image S j (1) into an array form, and perform image scale change and image enhancement processing using an image transformation method to obtain the processed small sample image S j (1) in array form.
[0071] Step S3.7: Convert the processed small sample image S j (1) in array form back to a small sample image in the form of an RGBA-formatted picture object, called the small sample image S j (2).
[0072] Step S3.8: Use the channel parameter of the alpha channel of the solid background image B0 as the fusion parameter to fuse the small sample image S j (2) and the transformable area of the solid background image B0, so as to incorporate the small sample image S j (2) into the transformable area of the solid background image B0 to obtain the fused sample image I, and update the label of the sample image I according to the label lable(S j ) of the small sample image S, so that the label lable(I) of the sample image I is the label lable(S j ). j )
[0073] Step S3.9: Return to Step S3.1, and repeat this process multiple times to achieve offline enhancement of the small sample image S j .
[0074] Preferably, in Step S3.6, the specific operations of performing image scale change and image enhancement processing using an image transformation method are as follows:
[0075] Randomly determine the image scaling scale using a random function; for the small sample image S in the form of an array j (1) Perform scale transformation and then random enhancement processing. The specific method is as follows:
[0076] If the small sample image S in the form of an array j (1) is a contour-sensitive sample, perform color perturbation and channel shuffling on the color data;
[0077] If the small sample image S in the form of an array j (1) is a color-sensitive sample, perform basic brightness and contrast enhancement, and at the same time introduce rectangular area occlusion, central cropping, high-intensity grid distortion, and elastic transformation processing.
[0078] Preferably, in step S7.4, the method for training the deep learning neural network model is as follows:
[0079] Step S7.4.1, in each iteration, input the training sample into the deep learning neural network model, and the deep learning neural network model outputs the model prediction result ModelOutput;
[0080] Step S7.4.2, according to the model prediction result ModelOutput, calculate the confidence simulation value C of each training sample i ;
[0081] Step S7.4.3, calculate the proportion n0 / n of the number of training samples n0 with a confidence simulation value greater than the set value k in the total number of training samples n, that is, (TH>k) / (TH>=0);
[0082] Step S7.4.4, calculate the initial penalty ratio PEN = 1 - n0 / n;
[0083] Step S7.4.5, calculate the final penalty value PEN loss = σ(obj_k·PEN·e p )
[0084] Where: obj_k is a constant coefficient of the scaling order of magnitude; p is a learnable parameter; σ represents an activation function;
[0085] Step S7.4.6, use the following formula to obtain the corrected confidence loss value OBJ loss :
[0086] OBJ loss = obj loss + PEN loss
[0087] Where: obj loss is the original confidence loss value; obtained through the target existence score;
[0088] Step S7.4.7, according to the corrected confidence loss value OBJ loss , adjust the parameters of the deep learning neural network model, including the learnable parameter p, so that the learnable parameter p is updated with the gradient, and the deep learning neural network model is optimized towards the set expected threshold.
[0089] The method for optimizing the threshold loss of data augmentation and learnable parameters based on few-shot learning provided by the present invention has the following advantages:
[0090] The present invention provides a method for optimizing the threshold loss of data augmentation and learnable parameters based on few-shot learning. By combining the few-shot two-stage data augmentation technology and the threshold loss optimization of learnable parameters, the present invention realizes a practical engineering innovation solution in the background of extremely few samples, greatly improves the model robustness and generalization ability in related few-shot fields, and at the same time improves the prediction accuracy of various categories and reduces the occurrence of false detections. Brief Description of the Drawings
[0091] Figure 1 It is a schematic flow chart of a method for optimizing the threshold loss of data augmentation and learnable parameters based on few-shot learning provided by the present invention. Detailed Embodiment
[0092] In order to make the technical problems, technical solutions and beneficial effects solved by the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0093] In the field of few-shot learning, data augmentation and loss function optimization are two key technical points, which are crucial for improving the performance of the model on a limited data set. The present invention provides a method for optimizing the threshold loss of data augmentation and learnable parameters based on few-shot learning, which solves the shortcomings of the existing technology such as strong data dependence, insufficient generalization ability, high false detection rate, single data augmentation means and high model training complexity:
[0094] In the two aspects of data augmentation and threshold loss optimization of learnable parameters in few-shot learning, compared with the traditional method, the present invention has the following characteristics:
[0095] (1) Data dependence
[0096] (1) Data dependence
[0097] Existing transfer learning methods strongly rely on the quality of pre-trained models and the diversity of source datasets. If the feature expression ability of the pre-trained model is insufficient or the data distribution difference between the source dataset and the target task is large, the effect of transfer learning will be greatly reduced. The present invention weakens the dependence on the source dataset by introducing data augmentation techniques and increasing the diversity of training data.
[0098] Existing technologies rely on a large amount of labeled data. Existing technologies usually require large-scale labeled datasets. In special fields such as medical diagnosis and security detection, the collection of data becomes difficult, and samples may have various differences and a particularly wide distribution. The present invention aims to greatly reduce the demand for labeled data and make better use of extremely few labeled data through special data augmentation methods to reduce the dependence on a large amount of labeled data in special fields.
[0099] (2) Generalization ability
[0100] Traditional few-shot learning methods often perform poorly when generalized to new tasks. Although the performance of the model can be improved to a certain extent through fine-tuning, when the target dataset has a large distribution difference from the training dataset, the generalization ability of the model is still limited. The present invention enables the model to better adapt to different data distributions during training through a threshold loss optimization method for learnable parameters, thereby improving the generalization ability.
[0101] (3) False detection rate
[0102] When existing technologies handle few-shot object detection, they often face the problem of a high false detection rate. This is because in the case of scarce data, it is difficult for the model to fully learn the features for distinguishing objects from the background. The present invention specially designs a new threshold loss function by introducing a threshold loss optimization method for learnable parameters, and effectively reduces the false detection rate through a threshold penalty mechanism.
[0103] (4) Data augmentation means
[0104] Traditional data augmentation methods have limited means and are usually limited to geometric transformations of images such as rotation, cropping, and scaling, and cannot fully capture the diversity of data. The present invention designs a more effective two-stage data augmentation method, which more comprehensively expands the diversity of training data and enhances the robustness of the model.
[0105] (5) Model training complexity
[0106] When existing technologies perform few-shot learning, they often require a complex model fine-tuning process, consuming a large amount of time and computing resources. The present invention simplifies the model training process and complexity while improving the model performance by introducing a new threshold loss function and data augmentation methods.
[0107] In summary, the existing technologies have the disadvantages of strong data dependence, insufficient generalization ability, high false detection rate, single data augmentation means, and high model training complexity. By proposing a data augmentation and threshold loss optimization method for learnable parameters based on few-shot learning, the present invention provides an effective solution to these disadvantages, reduces the dependence on a large amount of data in a specific field, enhances the generalization ability of the model, and reduces the false detection rate, significantly improving the performance and robustness of the model, thereby making it more practical and competitive in engineering in special fields.
[0108] (2) Threshold loss optimization of learnable parameters
[0109] (1) Support for few-shot learning:
[0110] Existing adaptive threshold optimization methods usually rely on a large amount of data for threshold adjustment and model training. When data is scarce, these methods are difficult to fully exert their effectiveness, resulting in poor generalization ability of the model. The present invention effectively expands the diversity of training data by combining a data augmentation method for few-shot learning, significantly improving the learning effect under few-shot conditions.
[0111] (2) Data augmentation means:
[0112] Traditional adaptive threshold optimization methods have relatively limited means in data augmentation, mainly relying on simple image transformations (such as rotation, cropping, scaling, etc.), and cannot make full use of data augmentation techniques to increase the diversity and quantity of training data. The present invention designs a dedicated data augmentation technique, making the generated data more diverse and rich, thereby further improving the robustness of the model.
[0113] (3) Refinement of threshold optimization:
[0114] The threshold optimization mechanism of the existing technology is relatively single, usually simply adding a threshold term, lacking fine adjustment for different samples and tasks. This method has limited effectiveness in dealing with complex and diverse data. The present invention makes the threshold adjustment more flexible and accurate by introducing a threshold loss optimization method for learnable parameters, and can better adapt to the characteristics of different samples. A non-zero parameter is added to the model layer of the present invention, and this parameter participates in both loss function adjustment and model training, enabling the model to learn the data distribution under the new threshold loss.
[0115] (4) False detection rate:
[0116] When dealing with object detection tasks, especially in multi-scale scenarios and complex backgrounds, the prior art often faces the problem of a high false detection rate. This is because traditional methods lack a targeted penalty mechanism when optimizing thresholds and it is difficult to effectively distinguish objects from non-objects. The present invention significantly reduces the false detection rate and improves the detection accuracy by designing a new threshold loss function and introducing a threshold penalty mechanism.
[0117] (5) Model training complexity
[0118] Adaptive threshold optimization methods usually involve complex calculation processes and require a large amount of computing resources and time for parameter optimization, which may become a bottleneck in practical applications. In the design of the present invention, the efficiency of training is fully considered. By combining an optimization algorithm and data augmentation techniques, only a learnable parameter is added to the final output of the model to dynamically adjust the threshold loss function and the weight distribution of the model, without adding complex calculations, and the impact on the speed and efficiency of training and inference can be almost ignored.
[0119] In summary, the prior art has significant deficiencies in aspects such as support for few-shot learning, data augmentation means, fineness of threshold optimization, false detection rate, and training complexity. The present invention effectively overcomes these shortcomings by proposing a data augmentation method combined with few-shot learning and a threshold loss optimization method with learnable parameters, thereby improving the performance and practicality of the model. The present invention can effectively reduce the dependence on a large amount of data in a specific field, enhance the generalization ability of the model, and reduce the false detection rate, making it more practical and competitive in engineering in special fields.
[0120] The key technical solutions of the present invention include a few-shot data augmentation technique and a threshold loss optimization method with learnable parameters. These technical solutions will cooperate with each other in a multi-scale object detection framework to achieve better performance and generalization ability.
[0121] Refer to Figure 1 , the data augmentation method combined with few-shot learning and the threshold loss optimization method with learnable parameters provided by the present invention include the following steps:
[0122] Step S1, establish an actual background image sample library G and a few-shot image sample library S; multiple actual background images are stored in the actual background image sample library G, and each actual background image is represented as G i ; multiple few-shot images are stored in the few-shot image sample library S, and each few-shot image is represented as S j ; wherein, the number of few-shot images stored in the few-shot image sample library S is much less than the number of actual background images stored in the actual background image sample library G.
[0123] Step S2, traverse the few-shot image sample library S, and for each few-shot image S traversed j, based on the actual background images in the actual background image sample library G, for the small sample image S j perform one-stage offline enhancement of the actual background image to expand the small sample image S j , obtaining the sample image F after one-stage offline enhancement of the actual background image, with its label being lable(F);
[0124] In step S2, for the small sample image S j perform one-stage offline enhancement of the actual background image, specifically:
[0125] Step S2.1, randomly select an actual background image G in the actual background image sample library G i ;
[0126] Step S2.2, convert the actual background image G i into the RGBA format to obtain the actual background image G i (1), and read the size of the actual background image G i (1);
[0127] Step S2.3, the small sample image S j has the label lable(S j ), convert the small sample image S j into the RGBA format to obtain the small sample image S j (1), and read the size of the small sample image S j (1);
[0128] Compare the size of the small sample image S j (1) and the size of the actual background image G i (1), so as to determine the transformable area in the actual background image G i (1);
[0129] Step S2.4, convert the small sample image S j (1) into an array form, and perform image scale change and image enhancement processing using an image transformation method to obtain the processed small sample image S in array form j (1);
[0130] Step S2.4, the specific operations of performing image scale change and image enhancement processing using an image transformation method are as follows:
[0131] Use a random function to randomly determine the image scaling scale; perform scale transformation on the small sample image S j (1) in array form, and then perform random enhancement processing, including flipping, row-column swapping, grid distortion, elastic transformation, rotation, random gamma, image mean filtering, random brightness, and adding Gaussian noise processing.
[0132] Step S2.5, the processed small sample graph S in array form j (1), and then convert it into a small sample graph in the form of a picture object in RGBA form, called small sample graph S j (2);
[0133] Step S2.6, use the channel parameter of the alpha channel of the actual background graph G i (1) as the fusion parameter, and fuse the transformable area of the small sample graph S j (2) and the actual background graph G i (1), so as to integrate the small sample graph S i (1) into the transformable area of the actual background graph G j (2), obtain the fused sample graph F, and update the label of the sample graph F according to the label lable(S j ) of the small sample graph S, so that the label lable(F) of the sample graph F is the label lable(S j ), and the label can include information such as scale, position, category, etc. j )
[0134] Step S2.7, return to Step S2.1, and randomly select an actual background graph again to perform offline enhancement processing on the small sample graph S j , and repeat this many times continuously to achieve the offline enhancement of the small sample graph S j .
[0135] Therefore, since the number of actual background graphs is much larger than the number of small sample graphs, each small sample graph will undergo multiple enhancements and fusions.
[0136] Step S3, traverse the small sample graph sample library S, and for each small sample graph S j traversed, perform offline enhancement of the pure color background graph in the first stage to expand the small sample graph S j , and obtain the sample graph I after offline enhancement of the pure color background graph in the first stage, and its label is lable(I);
[0137] In Step S3, for the small sample graph S j , perform offline enhancement of the pure color background graph in the first stage, specifically:
[0138] Step S3.1, randomly select an actual background graph G i in the actual background graph sample library G, and read the size of the actual background graph G i (1);
[0139] Step S3.2, according to the size of the actual background graph G i (1), initialize and generate a pure color background canvas of the same size;
[0140] Step S3.3: Pre-define multiple solid background colors, including mainstream solid background colors and random solid background colors; randomly select one solid background color and fill it into the solid background canvas to generate a solid background image B0 in RGBA format.
[0141] Step S3.4: The small sample image S j has a label lable(S j ), convert the small sample image S j into RGBA format to obtain the small sample image S j (1), and read the size of the small sample image S j (1).
[0142] Step S3.5: Compare the size of the small sample image S j (1) with the size of the solid background image B0, so as to determine the transformable area in the solid background image B0.
[0143] Step S3.6: Convert the small sample image S j (1) into an array form, and perform image scale change and image enhancement processing using an image transformation method to obtain the processed small sample image S j (1) in array form.
[0144] Step S3.6: The specific operations of performing image scale change and image enhancement processing using an image transformation method are as follows:
[0145] Use a random function to randomly determine the image scaling scale; perform scale transformation on the small sample image S j (1) in array form, and then perform random enhancement processing. The specific method is as follows:
[0146] If the small sample image S j (1) in array form is a contour-sensitive sample, perform color perturbation and channel shuffling on the color data; for contour-sensitive samples, the one-stage actual background image offline enhancement method in step S2 of the present invention can be additionally used to perform one-stage actual background image offline enhancement on it. Channel shuffling processing means performing fusion between channels.
[0147] If the small sample image S j (1) in array form is a color-sensitive sample, perform basic brightness and contrast enhancement, and at the same time introduce rectangular area occlusion, central cropping, high-intensity grid distortion, and elastic transformation processing.
[0148] Therefore, in the present invention, different data enhancement processing methods are performed on contour-sensitive small sample images and color-sensitive small sample images respectively.
[0149] Step S3.7: The processed small sample image S j(1), and then convert it into a small sample image in the form of an RGBA - formatted image object, called small sample image S j ;(2)
[0150] Step S3.8, use the channel parameters of the alpha channel of the solid - color background image B0 as the fusion parameters, and fuse the small sample image S j (2) and the transformable area of the solid - color background image B0, so as to incorporate the small sample image S j (2) into the transformable area of the solid - color background image B0, obtaining the fused sample image I, and update the label of the sample image I according to the label lable(S j ) of the small sample image S j ), so that the label lable(I) of the sample image I is the label lable(S j );
[0151] Step S3.9, return to Step S3.1, and repeat this process many times to achieve the offline enhancement of the small sample image S j .
[0152] Through Step S3, the actual background is replaced by a solid - color background, so the fused sample image has sample characteristics different from those of the offline enhancement of the first - stage actual background image.
[0153] Step S4, randomly perform second - stage enhancement on the sample image F and the sample image I with a probability of 0.5. Specifically, use formula (1) for second - stage fusion enhancement to obtain the fused and enhanced sample image Mix, and its label is lable(Mix):
[0154] Mix = λ·F+(1 - λ)·I (1)
[0155] lable(Mix)=λ·lable(F)+(1 - λ)·lable(I)
[0156] Among them, λ is the mixing coefficient randomly sampled from the Beta distribution. For example, λ can be 0.2.
[0157] Step S5, through Steps S2 to S4, obtain a training sample set formed by multiple sample images F, sample images I, and sample images Mix; as a specific application, the sample images F and the sample images I account for 50% of the total number of samples, and the sample images Mix account for 50% of the total number of samples.
[0158] Step S6, establish a deep - learning neural network model; the deep - learning neural network model has learnable parameters p
[0159] As a specific application, the deep - learning neural network model constructed in the present invention is expressed as:
[0160] F(x)=M0·Mout ·e p
[0161] Where: M0 is the main structure of the model; M out is the output layer of the model; p is the learnable parameter; in the present invention, the form of e p can prevent the situation that the learnable parameter p becomes negative during automatic update, thereby causing an abnormal penalty value of the additional loss function.
[0162] Step S7, training the deep learning neural network model with a training sample set to obtain a trained deep learning neural network model. The specific training method is as follows:
[0163] Step S7.1, initializing the learnable parameter p as an identity matrix; the sample graphs F, I, and Mix respectively have initial intensities;
[0164] Step S7.2, initializing the initial value of the iteration number t as 1;
[0165] Step S7.3, using an online data augmentation module to perform online data augmentation intensity processing on the sample graphs F, I, and Mix according to formula (2) to obtain the processed sample graphs F, I, and Mix. The method is as follows:
[0166]
[0167] Where:
[0168] A t is the online data augmentation intensity of the sample graph in the t-th round of training;
[0169] A0 is the initial intensity of the sample graph;
[0170] t is the current iteration number;
[0171] T is the total number of iterations;
[0172] η t is the learning rate in the t-th round of training;
[0173] η0 is the initial learning rate;
[0174] p y is the value of the learnable parameter in the t-th round of training;
[0175] It can be seen that the present invention provides an online data augmentation method. During the model training process, online data augmentation attenuation is introduced. The sample graphs F, I, and Mix enter the same online data augmentation module, and moreover, the intensity of data augmentation changes with the iteration of the model and the change of relevant parameters.
[0176] Step S7.4, input the sample graphs F, I, and Mix processed in step S7.3 into a deep learning neural network model for training;
[0177] Step S7.4, the method for training the deep learning neural network model is as follows:
[0178] Step S7.4.1, in each iteration, input the training samples into the deep learning neural network model, and the deep learning neural network model outputs the model prediction result ModelOutput;
[0179] Step S7.4.2, according to the model prediction result ModelOutput, calculate the confidence simulation value C of each training sample i , and the specific calculation method is:
[0180] Define the sigmoid function:
[0181]
[0182] Therefore, the confidence simulation value where c i is the original confidence value of each training sample;
[0183] Step S7.4.3, calculate the proportion n0 / n of the number of training samples n0 with confidence simulation value greater than the set value k in the total number of training samples n, that is, (TH>k) / (TH>=0);
[0184] Step S7.4.4, calculate the initial penalty ratio PEN = 1 - n0 / n;
[0185] Specifically,
[0186] t mask = C i > k
[0187] F mask = C i
[0188]
[0189] where: T mask is the mask where the confidence simulation value of the sample is greater than the set value k, F mask is the mask of all samples, T count is the number of masks of T mask and F count is the number of masks of F mask . Finally, obtain the initial penalty ratio PEN.
[0190] Step S7.4.5, calculate the final penalty value PEN loss= σ(obj_k·PEN·e p )
[0191] Where: obj_k is a constant coefficient of the scaling order of magnitude; p is a learnable parameter; σ represents an activation function;
[0192] In this step, since PEN·e p is still relatively large compared to the original confidence loss value, which is not conducive to the fast and stable convergence of the model, so the present invention multiplies by a scaling coefficient obj_k to scale this auxiliary confidence penalty loss value to the same order of magnitude as the original confidence loss value.
[0193] Step S7.4.6, using the following formula, obtain the corrected confidence loss value OBJ loss :
[0194] OBJ loss = obj loss + PEN loss
[0195] Where: obj loss is the original confidence loss value; obtained through the target existence score;
[0196] Therefore, through this step, add the original confidence loss value obj loss to the penalty value PEN loss , obtain the corrected confidence loss value OBJ loss , and adjust the parameters of the deep learning neural network model through the corrected confidence loss value OBJ loss . Since the learnable parameter p will be updated along with the gradient, the model is optimized towards the set expected threshold, and finally a new model with a very low false detection rate under the limited threshold is obtained.
[0197] Step S7.4.7, according to the corrected confidence loss value OBJ loss , adjust the parameters of the deep learning neural network model, including the learnable parameter p, so that the learnable parameter p is updated along with the gradient, and the deep learning neural network model is optimized towards the set expected threshold.
[0198] Step S7.5, let t = t + 1, return to step S7.3, and continuously iterate and train like this until the training termination condition is reached.
[0199] In the present invention, during model training, update the learnable parameter p, thereby affecting the output of the model, and automatically generalize the model to the expected range. In the training optimizer, add the learnable parameter p to the gradient update, so that the model will automatically update this parameter during the training process.
[0200] The present invention provides a method for optimizing the threshold loss of data augmentation and learnable parameters based on few-shot learning, with the following innovations:
[0201] 1. Few-shot data augmentation technique:
[0202] Through a two-stage data augmentation method, the limited sample data is diversified, including the augmentation of the few-shot data itself and the augmentation after background fusion, etc., significantly reducing the dependence on a large amount of data and enhancing the generalization ability of the model.
[0203] 2. Method for optimizing the threshold loss of learnable parameters:
[0204] Introduce a learnable penalty_factor parameter, which is automatically adjusted during the model training process. By penalizing and adjusting the confidence loss of the model output, the performance of the model under the set threshold is optimized, reducing false positives. At the same time, this parameter can also act on the actual output of the model to dynamically adjust the output representation of the model, further improving the reliability, practicality and engineering value of the model.
[0205] The present invention provides a method for optimizing the threshold loss of data augmentation and learnable parameters based on few-shot learning. By combining the few-shot two-stage data augmentation technique and the threshold loss optimization of learnable parameters, the present invention realizes a practical engineering innovation solution under the background of extremely few samples, greatly enhancing the model robustness and generalization ability in related few-shot fields. At the same time, it also improves the prediction accuracy of various categories and reduces the occurrence of false positives. Compared with traditional methods, the present invention not only achieves a significant improvement in model performance, but also achieves outstanding results in reducing the false positive rate, bringing important technological innovation and practical application prospects to the field of few-shot data processing.
[0206] The technical solution of the present invention brings beneficial effects in many aspects, solves the problem of dependence on a large amount of data in specific fields, enhances the generalization ability of the model and reduces the false positive rate, bringing important technological innovation and practical application prospects to the field of few-shot data processing. The following is a detailed description of the beneficial effects brought by the present invention:
[0207] 1. Reducing data dependence:
[0208] Traditional methods have a high dependence on a large amount of data, while the few-shot data augmentation technique adopted by the present invention significantly reduces the demand for a large amount of data by diversifying the limited sample data. This enables the construction of an effective model even in the case of scarce data, reducing the cost of data collection and annotation.
[0209] 2. Reducing the false positive rate:
[0210] Adopting a threshold loss optimization method with learnable parameters, the present invention enables the model to optimize the model according to a set desired threshold during the training process, restricting the occurrence of false detections. By penalizing and adjusting the confidence loss of the model output, false alarms in real scenarios are effectively reduced, improving the reliability and practicality of the model.
[0211] 3. Reduced the training difficulty of the model:
[0212] The technical solution of the present invention has brought remarkable effects in reducing the training difficulty of the model. In traditional methods, training for small-sample data usually requires a more complex model architecture or more training resources to obtain satisfactory performance. The small-sample data augmentation technique and the threshold loss optimization method with learnable parameters adopted by the present invention greatly simplify the model training process, providing a more convenient and efficient solution for engineering practice.
[0213] 4. Enhanced the robustness of the model:
[0214] The technical solution of the present invention is not only applicable to small-sample data processing, but also enhances the adaptability of the model to various complex situations and scenarios. This enhanced robustness enables the model to better cope with data noise, interference, and changes, maintaining stable performance.
[0215] 5. Improved the prediction accuracy:
[0216] By effectively augmenting small-sample data and optimizing model parameters, the present invention significantly improves the prediction accuracy of the model. This is crucial for the engineering deployment of tasks such as object detection and image classification, and can provide more reliable results for practical applications.
[0217] In summary, the technical solution of the present invention not only plays a significant role in solving technical problems in specific fields, but also brings important engineering application value in terms of improving model performance, reducing costs, and enhancing reliability.
[0218] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for optimizing the threshold loss of data augmentation and learnable parameters based on few-shot learning, characterized in that It includes the following steps: Step S1, establish an actual background image sample library G and a small sample image sample library S; multiple actual background images are stored in the actual background image sample library G, and each actual background image is represented as G i ; multiple small sample images are stored in the small sample image sample library S, and each small sample image is represented as S j ; Step S2, traverse the small sample graph sample library S, and for each small sample graph S traversed j , based on the actual background graphs in the actual background graph sample library G, for the small sample graph S j perform one-stage actual background graph offline enhancement to expand the small sample graph S j , and obtain the sample graph F after one-stage actual background graph offline enhancement, whose label is lable(F); Step S3, traverse the small sample image library S, and for each small sample image S traversed j , perform one-stage solid background image offline enhancement to expand the small sample image S j , to obtain the sample image I after one-stage solid background image offline enhancement, whose label is lable(I); Step S4: For the sample graph F and the sample graph I, perform two-stage fusion enhancement using formula (1) to obtain the fused and enhanced sample graph Mix, whose label is lable(Mix): Mix = λ·F+(1 - λ)·I (1) lable(Mix) = λ·lable(F)+(1 - λ)·lable(I) where λ is the mixing coefficient randomly sampled from the Beta distribution; Step S5: Through steps S2 to S4, obtain a training sample set formed by multiple sample graphs F, sample graphs I, and sample graphs Mix; Step S6: Establish a deep learning neural network model; the deep learning neural network model has learnable parameters p; Step S7: Use the training sample set to train the deep learning neural network model to obtain a trained deep learning neural network model. The specific training method is as follows: Step S7.1: Initialize the learnable parameter p as the identity matrix; the sample graphs F, I, and Mix each have an initial intensity; Step S7.2: Initialize the initial value of the iteration number t as 1; Step S7.3: Use the online data augmentation module to perform online data augmentation intensity processing on the sample graphs F, I, and Mix according to formula (2) to obtain the processed sample graphs F, I, and Mix. The method is as follows: where: A t is the online data augmentation intensity of the sample graph in the t-th round of training; A0 is the initial intensity of the sample graph; t is the current iteration number; T is the total number of iterations; η t is the learning rate for the t-th round of training; η0 is the initial learning rate; p y is the value of the learnable parameter for the t-th round of training; Step S7.4: Input the sample graphs F, I, and Mix processed in step S7.3 into the deep learning neural network model for training; Step S7.5: Let t = t + 1, and return to step S7.
3. Keep iterating and training like this until the training termination condition is reached.
2. The threshold loss optimization method for data augmentation and learnable parameters based on few-shot learning according to claim 1, characterized in that In step S2, perform one-stage offline enhancement on the small sample graph S j Specifically, it is as follows: Step S2.1, randomly select an actual background image G from the actual background image sample library G i ; Step S2.2, convert the actual background image G i into the RGBA format to obtain the actual background image G i (1), and read the size of the actual background image G i (1); Step S2.3, small sample graph S j with label lable(S j ), convert the small sample graph S j into RGBA format to obtain the small sample graph S j (1), and read the size of the small sample graph S j (1); Compare the small sample graph S j with the size of the actual background graph G i and the size of (1), so as to determine i the transformable region in the actual background graph G Step S2.4, convert the small sample graph S j (1) into an array form, and perform image scale change and image enhancement processing using an image transformation method to obtain the processed small sample graph S in array form j (1); Step S2.5, convert the processed small sample graph S in the form of an array j (1) into a small sample graph in the form of a picture object in RGBA format, which is called small sample graph S j (2); Step S2.6, actual background image G i Use the channel parameters of the alpha channel of (1) as the fusion parameter to fuse the small sample image S j (2) and the actual background image G i Fuse the transformable region of (1) so as to incorporate the small sample image S i Into the transformable region of the actual background image G(1) j (2) to obtain the fused sample image F, and update the label of the sample image F according to the label lable(S j ) of the small sample image S j ) such that the label lable(F) of the sample image F is the label lable(S j ); Step S2.7, return to Step S2.1, randomly select an actual background image again, and perform offline enhancement processing on the small sample image S j Repeat this process multiple times to achieve offline enhancement of the small sample image S j for offline enhancement.
3. A method for optimizing the threshold loss of data augmentation and learnable parameters based on few-shot learning according to claim 2, characterized in that Step S2.4: Use the image transformation method to perform image scale change and image enhancement processing. Specifically: Randomly determine the image scaling scale using a random function; for the small sample image S in array form j (1) Perform scale transformation and then perform random augmentation processing, including flipping, row-column swapping, grid distortion, elastic transformation, rotation, random gamma, image mean filtering, random brightness, and adding Gaussian noise processing.
4. A method for optimizing threshold loss of data augmentation and learnable parameters based on few-shot learning according to claim 1, characterized in that In step S3, for the small sample graph S j , perform one-stage offline enhancement of the solid-color background graph, specifically: Step S3.1, randomly select an actual background image G from the actual background image sample library G i , and read the actual background image G i (1)'s size; Step S3.2, according to the size of the actual background image G i (1), initialize and generate a solid-color background canvas of the same size; Step S3.3: Pre-define multiple solid background colors, including mainstream solid background colors and random solid background colors; randomly select a solid background color and fill it into the solid background canvas to generate a solid background graph B0 in RGBA format; Step S3.4, small sample graph S j with label lable(S j ), convert the small sample graph S j into RGBA form to obtain the small sample graph S j (1), and read the size of the small sample graph S j (1); Step S3.5, compare the size of the small sample image S j (1) with the size of the solid color background image B0, so as to determine the transformable area in the solid color background image B0; Step S3.6, convert the small sample graph S j (1) into an array form, and perform image scale change and image enhancement processing using an image transformation method to obtain the processed small sample graph S in array form j (1); Step S3.7, convert the processed small sample graph S in array form j (1) into a small sample graph in the form of a picture object in RGBA form, called small sample graph S j (2); Step S3.8, use the channel parameters of the alpha channel of the solid-color background image B0 as the fusion parameter, and fuse the small sample image S j (2) with the transformable region of the solid-color background image B0, so as to integrate the small sample image S j (2) into the transformable region of the solid-color background image B0, obtaining the fused sample image I, and update the label of the sample image I according to the label lable(S j ) of the small sample image S j , so that the label lable(I) of the sample image I is the label lable(S j ); Step S3.9, return to Step S3.1, and repeat this process multiple times to achieve the offline enhancement of the small sample graph S j .
5. A method for optimizing the threshold loss of data augmentation and learnable parameters based on few-shot learning according to claim 4, characterized in that, Step S3.6: Use the image transformation method to perform image scale change and image enhancement processing. Specifically: Randomly determine the image scaling scale using a random function; for the small sample image S in the form of an array j (1) Perform scale transformation and then perform random enhancement processing. The specific method is as follows: If the small sample graph S in the form of an array j (1) is a contour-sensitive sample, perform color perturbation and channel shuffling on the color data; If the small sample graph S in the form of an array j (1) is a color-sensitive sample, perform basic brightness and contrast enhancement, and at the same time introduce rectangular area occlusion, central cropping, high-intensity grid distortion, and elastic transformation processing.
6. The threshold loss optimization method for data augmentation and learnable parameters based on few-shot learning according to claim 1, characterized in that Step S7.4: The method for training the deep learning neural network model is: Step S7.4.1: In each iteration, input the training sample into the deep learning neural network model, and the deep learning neural network model outputs the model prediction result ModelOutput; Step S7.4.2, calculate the confidence simulation value C of each training sample according to the model prediction result ModelOutput i ; Step S7.4.3: Calculate the proportion n0 / n of the number of training samples n0 with a confidence simulation value greater than the set value k in the total number of training samples n, that is, (TH > k) / (TH >= 0); Step S7.4.4: Calculate the initial penalty ratio PEN = 1 - n0 / n; Step S7.4.5, calculate the final penalty value PEN loss = σ(obj_k·PEN·e p ) where: obj_k is the constant coefficient of the scaling order of magnitude; p is the learnable parameter; σ represents the activation function; Step S7.4.6, use the following formula to obtain the corrected confidence loss value OBJ loss : OBJ loss = obj loss + PEN loss where: obj loss is the original confidence loss value; obtained from the target existence score; Step S7.4.7, according to the corrected confidence loss value OBJ loss , adjust the parameters of the deep learning neural network model, including the learnable parameter p, so that the learnable parameter p is updated along with the gradient, and the deep learning neural network model is optimized towards the set expected threshold.
Citation Information
Patent Citations
Medical image classification method based on small sample meta learning
CN116503668A
Small sample image classification optimization method based on supervised comparative learning and multi-task setting
CN116524242A