Method, device and medium for small-sample incremental learning by manipulating a small number of network parameters
By adding parameter coding modules and a small number of parameter control submodules to the convolutional neural network, the problem of high training cost and large sample demand for convolutional neural networks when adding new identification categories is solved, efficient small sample incremental learning is achieved, and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202210435521.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-04-24
AI Technical Summary
Existing convolutional neural networks need to be retrained when adding new identification categories and require a large number of samples, resulting in high training costs and inconsistent with cognitive learning.
A small sample incremental learning method for manipulating parameters of small samples is designed. By adding parameter encoding modules to the convolutional neural network, the base class sample set is used to train the base class identification network, and incremental training is performed through a small number of new class sample sets, combining a small number of parameter control submodules and base parameter control submodules to generate coding features to solve the problems of forgetting and underfitting.
It is realized that only a small number of samples are needed to train effectively when adding new identification categories, which significantly improves accuracy, reduces training costs, and improves prediction performance.
Smart Images

Figure CN114819078B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for small-sample incremental learning of a small number of network parameter manipulation parameters, and belongs to the technical field of computer vision. Background Art
[0002] Convolutional neural networks (CNNs) have precise annotations and can absorb annotation information, and then increase their usability through large-scale data sets, making them widely used in visual image recognition.
[0003] However, during the training of convolutional neural networks, a large number of samples need to be provided, that is, a large number of objects need to be annotated. Annotating a large number of objects is both expensive and laborious, and this is also inconsistent with cognitive learning. Moreover, when new recognition categories need to be added after training, a large number of new category pictures need to be re-obtained to retrain the convolutional neural network.
[0004] Therefore, it is necessary to provide a method that can effectively recognize new class images by training a convolutional neural network with only a small number of samples when adding new recognition types. Summary of the Invention
[0005] In order to overcome the above problems, the inventors of the present invention have conducted in-depth research and designed a method for small-sample incremental learning of a small number of network parameter manipulation parameters, including the following steps:
[0006] S1. Set up a small-sample incremental learning neural network, which is obtained by adding a parameter encoding module to a convolutional neural network;
[0007] S2. Train the small-sample incremental learning neural network through a base class sample set, that is, a picture sample set with base class labels, to obtain a base class recognition network.
[0008] S3. Input the pictures with base class labels into the base class recognition network, and the base class recognition network outputs the labels of the pictures to obtain the information in the pictures.
[0009] When pictures with new type labels need to be identified, that is, new class label pictures, there are also steps:
[0010] S4. Train the base class recognition network through a picture sample set with new class labels to obtain an augmented class recognition network.
[0011] S5. Input the pictures with base class or new class labels into the augmented class recognition network, and the augmented class recognition network outputs the labels of the pictures to obtain the information of the pictures.
[0012] Further, in S4, the number of samples of any class in the picture sample set with new class labels is less than the number of samples of any class in the base class sample set.
[0013] Further, in S1, the parameter encoding module is arranged between the feature extractor and the classifier of the convolutional neural network;
[0014] The feature extractor of the convolutional neural network transfers the feature map to the parameter encoding module, encodes the feature map through the parameter encoding module, obtains the encoded feature and transfers it to the classifier, and the classifier outputs a prediction result according to the encoded feature;
[0015] The parameter encoding module includes a small-parameter control sub-module and a base-parameter control sub-module,
[0016] The small-parameter control sub-module is used to generate the control coefficient α of the picture, and the base-parameter control sub-module is used to generate the base parameter θ of the picture γ , and the encoded feature is a linear combination of the control coefficient α and the base parameter θ γ of.
[0017] Preferably, the small-parameter control sub-module includes a pooling layer, a fully connected layer, a non-linear layer, a fully connected layer, and a softmax layer connected in sequence, and its output is expressed as:
[0018]
[0019] α is a set of control coefficients, x is the feature map transferred by the feature extractor, and θ α represents a set of network parameters of different layers, where, represents the parameters of the second fully connected layer, represents the parameters of the first fully connected layer, represents the bias of the first fully connected layer, represents the bias of the second fully connected layer;
[0020] The base-parameter control sub-module is formed by connecting one convolutional kernel or multiple convolutional kernels in series, and the set of trainable parameters of the convolutional kernel is called the base parameter θ γ .
[0021] Preferably, the encoded feature θ e can be expressed as
[0022]
[0023] where, α n represents different parameters in the set of control coefficients, represents different parameters in the set of base parameters, and N is a hyperparameter representing the number of parameters in the set.
[0024] Preferably, in S2, during the training process, the parameters in the feature extractor, the base parameter θ of the base-parameter control sub-module γ , the network parameter θ of the small-parameter control sub-moduleα All are trainable parameters.
[0025] Preferably, in S4, during the training process, the parameters in the feature extractor and the base parameters θ of the base parameter control sub-module γ remain unchanged, and the network parameters θ of the small number of parameter control sub-module α are trainable parameters.
[0026] Preferably, in S4, a distillation loss function is adopted during the training process.
[0027] In addition, the present invention also provides an electronic device, including:
[0028] at least one processor; and
[0029] a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.
[0030] In addition, the present invention also provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above method.
[0031] The beneficial effects of the present invention include:
[0032] (1) The present invention provides a simple and effective method to solve the fundamental contradiction between a small number of training samples and a large number of training parameters in incremental categories;
[0033] (2) The present invention discloses a learnable parameter encoding mechanism, making it possible to generate richer feature representations on a fixed set of base parameters;
[0034] (3) The present invention significantly improves the accuracy rate and obtains better prediction performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 Shows the small-sample incremental learning neural network structure diagram in the method of small-sample incremental learning with a small number of network parameter manipulation parameters according to a preferred embodiment of the present invention;
[0036] Figure 2 Shows the structure diagram of the small number of parameter control sub-module in the method of small-sample incremental learning with a small number of network parameter manipulation parameters according to a preferred embodiment of the present invention;
[0037] Figure 3 Shows the visualization results of feature dimensionality reduction after classification of Example 1 and Comparative Example 7 in the experimental example. DETAILED DESCRIPTION
[0038] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Through these descriptions, the features and advantages of the present invention will become more clearly defined.
[0039] The specific term "exemplary" herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.
[0040] A method for small-sample incremental learning of a small number of network parameter manipulation parameters according to the present invention includes the following steps:
[0041] S1. Set up a small-sample incremental learning neural network, which is obtained by adding a parameter encoding module to a convolutional neural network;
[0042] S2. Train the small-sample incremental learning neural network through a base class sample set, that is, a picture sample set with base class labels, to obtain a base class recognition network.
[0043] S3. Input the pictures with base class labels into the base class recognition network, and the base class recognition network outputs the labels of the pictures, so as to obtain the information in the pictures.
[0044] When it is necessary to identify pictures with new type labels, that is, pictures with new class labels, there are also steps:
[0045] S4. Train the base class recognition network through a picture sample set with new class labels to obtain an augmented class recognition network, that is, perform incremental training on the convolutional neural network;
[0046] S5. Input the pictures with base class or new class labels into the augmented class recognition network, and the augmented class recognition network outputs the labels of the pictures, so as to obtain the information of the pictures.
[0047] According to the present invention, in S4, in the picture sample set with new class labels, the number of samples of any class is less than the number of samples of any class in the base class sample set, that is, when performing incremental training on the convolutional neural network, only a small number of samples are required to complete effective training. Preferably, during the incremental training process, the number of samples of any class for training only needs to be 0.5-1% of the number of samples of any class in the base class sample set. For example, there are 60 classes of samples in the base class sample set, and the number of samples of each class is 600. During the incremental training process, the number of samples of each class in the picture sample set with new class labels can be 5.
[0048] In the present invention, the convolutional neural network can be any known convolutional network, such as Resnet18, Resnet20, etc. Those skilled in the art can select a suitable convolutional neural network according to actual needs and add a parameter encoding module to obtain a small-sample incremental learning neural network.
[0049] Further, in the convolutional neural network, there are an input layer, a feature extractor, and a classifier. Among them, the feature extractor is used to extract a feature map from the input picture and pass the value to the classifier, and the classifier predicts and outputs the picture label type according to the feature map.
[0050] In the present invention, the input layer, feature extractor, and classifier of the convolutional neural network are not improved, and only a parameter encoding module is added thereto. The parameter encoding module can be added in the feature extractor, for example, in different blocks, or can be set between the feature extractor and the classifier.
[0051] Preferably, the parameter encoding module is set between the feature extractor and the classifier of the convolutional neural network, as Figure 1 shown;
[0052] The feature extractor of the convolutional neural network passes the feature map to the parameter encoding module, encodes the feature map through the parameter encoding module to obtain an encoded feature, and passes it to the classifier, and the classifier outputs a prediction result according to the encoded feature;
[0053] In the present invention, a small-sample class incremental learning method is used to improve the convolutional neural network. Although the present invention is the same as the existing small-sample class incremental learning method in that a model is trained using the base classes and then the model is continuously generalized to new classes through new class samples, the existing small-sample class incremental learning ignores the internal contradiction between a large number of network parameters and a small number of training samples, resulting in problems of forgetting and underfitting in the neural network. How to solve the above problems is the difficulty of the present invention.
[0054] In the present invention, the parameter encoding module includes a small-parameter control sub-module and a base-parameter control sub-module,
[0055] The small-parameter control sub-module is used to generate a control coefficient α for the picture, and the base-parameter control sub-module is used to generate a base parameter θ for the picture γ , and the encoded feature is a linear combination of the control coefficient α and the base parameter θ γ . The feature representation ability on the incremental new class is expanded through the control coefficient α.
[0056] Preferably, the small-parameter control sub-module includes a pooling layer, a fully connected layer, a non-linear layer, a fully connected layer, and a softmax layer connected in sequence, as Figure 2 shown, and its output is expressed as:
[0057]
[0058] α is a set of control coefficients, x is the feature map passed by the feature extractor, and θ α represents a set of network parameters of different layers, where, represents the parameters of the second fully connected layer, represents the parameters of the first fully connected layer, represents the bias of the first fully connected layer, represents the bias of the second fully connected layer;
[0059] Furthermore, two fully connected layers in the small number of parameter control sub-module are used to reduce the dimensionality of the image features to the dimension of the control coefficients, and the final softmax layer ensures that the sum of each group of output control coefficients is 1.
[0060] The base parameter control sub-module is formed by connecting one convolutional kernel or multiple convolutional kernels in series, and the set of trainable parameters of the convolutional kernel is called the base parameter θ γ .
[0061] In a preferred embodiment, the encoded feature θ e can be expressed as
[0062]
[0063] where, α n represents different parameters in the set of control coefficients, represents different parameters in the set of base parameters, and N is a hyperparameter representing the number of parameters in the set.
[0064] In the present invention, by introducing control coefficients to generate new encoded features for different samples, using the control coefficients output by the base parameter control sub-module as combination weights, and combining the base parameters in a linearly weighted manner, a unique set of parameters is generated for each sample to simultaneously solve the problems of forgetting and underfitting. Specifically,
[0065] In S2, during the training process, the parameters in the feature extractor, the base parameter θ of the base parameter control sub-module γ , and the network parameter θ of the small number of parameter control sub-module α are all trainable parameters. In S4, during the training process, the parameters in the feature extractor, the base parameter θ of the base parameter control sub-module γ remain unchanged, and the network parameter θ of the small number of parameter control sub-module α is a trainable parameter.
[0066] In the present invention, by freezing the parameters and base parameters in the feature extractor, the feature drift of the base class is alleviated to solve the forgetting problem. At the same time, a small number of parameters in the control sub-module are fine-tuned to gradually generalize the feature space to new classes, alleviating the underfitting on the new classes.
[0067] According to the present invention, during the training processes of steps S2 and S4, similar to the training of a traditional convolutional neural network, by calculating the loss between the predicted output and the annotation, preferably the cross-entropy loss, and calculating the gradient of the loss function, the error gradient is backpropagated through the network to update the network parameters.
[0068] Further, in S4, similar to the traditional few-shot class incremental learning method, after adding pictures with new class labels, the head of the classifier is correspondingly increased so that the classifier adapts to the number of picture types after the addition.
[0069] According to a preferred embodiment of the present invention, in S4, during the training process, a distillation loss function is adopted to further prevent the forgetting of base class features, thereby realizing updating only a small number of control parameters and classifier parameters with a small number of samples of new classes to control the extension of the feature representation to new classes.
[0070] The various embodiments of the methods described above in the present invention can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, the one or more computer programs being executable and / or interpretable on a programmable system including at least one programmable processor, the programmable processor can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0071] The program code for implementing the methods of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0072] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0073] In order to provide interaction with a user, the methods and apparatuses described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0074] The methods and apparatuses described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of the communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0075] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system or a server combined with a blockchain.
[0076] It should be understood that various forms of processes shown above can be used, reordering, adding or deleting steps. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved, and no limitations are imposed herein.
[0077] Embodiment
[0078] Embodiment 1
[0079] A small-sample incremental learning experiment with a small number of network parameter manipulation parameters is carried out using a publicly available dataset, where the dataset is the CUB200 dataset.
[0080] CUB200 is a fine-grained dataset and also a benchmark image dataset for current fine-grained classification and recognition research. This dataset has a total of 11,788 bird images, including 200 bird subclasses. Among them, 5,994 images are used for training and 5,794 images are used for testing. In the experiment, the first 100 classes are used as the base classes and the last 100 classes are used as the new classes.
[0081] During the training process of the new classes, a 5-way-5-shot setting is adopted, that is, each task contains 5 new classes, and each class has 5 images.
[0082] In this experiment, the accuracy metric is used for performance evaluation. This metric is the probability of correct prediction among all samples. For each class, the accuracy is calculated as accuracy = (TP + TN) / (TP + FN + FP + TN), where TP, FP, and FN represent true positives, false positives, and false negatives respectively.
[0083] Specifically, the small-sample incremental learning with a small number of network parameter manipulation parameters is carried out using the following steps:
[0084] S1. Set up a small-sample incremental learning neural network, which is obtained by adding a parameter encoding module to a convolutional neural network;
[0085] S2. Train the small-sample incremental learning neural network with a base class sample set, that is, a set of picture samples with base class labels, to obtain a base class recognition network.
[0086] S3. Input pictures with base class labels into the base class recognition network, and let the base class recognition network output the labels of the pictures, so as to obtain the information in the pictures.
[0087] S4. Train the base class recognition network with a set of picture samples with new class labels to obtain an incremental class recognition network.
[0088] S5. Input pictures with base class or new class labels into the incremental class recognition network, and let the incremental class recognition network output the labels of the pictures, so as to obtain the information of the pictures.
[0089] Further, in S4, the number of samples of any class in the set of picture samples with new class labels is less than the number of samples of any class in the base class sample set.
[0090] Further, in S1, the convolutional neural network is a Resnet18 network, and the parameter encoding module is set between the feature extractor and the classifier of the convolutional neural network;
[0091] The feature extractor of the convolutional neural network transfers the feature map to the parameter encoding module, encodes the feature map through the parameter encoding module to obtain the encoded feature, and transfers the encoded feature to the classifier. The classifier outputs the prediction result according to the encoded feature;
[0092] The parameter encoding module includes a small parameter control sub-module and a base parameter control sub-module.
[0093] The small parameter control sub-module is used to generate the control coefficient α of the picture, and the base parameter control sub-module is used to generate the base parameter θ of the picture. γ , and the encoded feature is a linear combination of the control coefficient α and the base parameter θ. γ
[0094] Further, the small parameter control sub-module includes a pooling layer, a fully connected layer, a non-linear layer, a fully connected layer, and a softmax layer connected in sequence, expressed as:
[0095]
[0096] Among them, α is a set of control coefficients, x is the feature map transferred by the feature extractor, represents the parameters of the second fully connected layer, represents the parameters of the first fully connected layer. represents the bias of the first fully connected layer, represents the bias of the second fully connected layer;
[0097] The base parameter control sub-module is a convolutional kernel, and the trainable parameters of the convolutional kernel are called base parameters θ γ .
[0098] Furthermore, the encoded feature θ e can be expressed as
[0099]
[0100] Furthermore, in S2, during the training process, the parameters in the feature extractor, the base parameters θ of the base parameter control sub-module γ , and the network parameters θ of the small number of parameter control sub-modules α are all trainable parameters. Furthermore, in S4, during the training process, the parameters in the feature extractor, the base parameters θ of the base parameter control sub-module γ remain unchanged, and the network parameters θ of the small number of parameter control sub-modules α are trainable parameters.
[0101] Comparative example
[0102] Comparative example 1
[0103] Small-sample class incremental image classification is performed using the same dataset as in Example 1, except that the iCaRL method is used, where iCaRL is proposed in the literature "icarl: Incremental classifier and representation learning. In: IEEE CVPR. (2017)".
[0104] Comparative example 2
[0105] Small-sample class incremental image classification is performed using the same dataset as in Example 1, except that the EEIL method is used, where EEIL is proposed in the literature "End-to-end incremental learning. In: ECCV. (2018)".
[0106] Comparative example 3
[0107] Small-sample class incremental image classification is performed using the same dataset as in Example 1, except that the NCM method is used, and NCM is proposed in the literature "Learning a unified classifier incrementally via rebalancing. In: IEEE CVPR. (2019)".
[0108] Comparative Example 4
[0109] Small-sample class-incremental image classification is performed using the same data set as in Example 1, except that the TOPIC method is used, where TOPIC is proposed in the literature "Few-shot class incremental learning. In: IEEE CVPR. (2020)".
[0110] Comparative Example 5
[0111] Small-sample class-incremental image classification is performed using the same data set as in Example 1, except that the SKW method is used, where SKW is proposed in the literature "Semantic-aware knowledge distillation for few-shot class-incremental learning. In: IEEE CVPR(2021)".
[0112] Comparative Example 6
[0113] Small-sample class-incremental image classification is performed using the same data set as in Example 1, except that the FSLL method is used, where FSLL is proposed in the literature "Few-shot lifelong learning. In: AAAI(2021)".
[0114] Comparative Example 7
[0115] Small-sample class-incremental image classification is performed using the same data set as in Example 1, except that the SPPR method is used, where SPPR is proposed in the literature "Self-promoted prototype refinement for few-shot class-incremental learning. In: CVPR.(2021)".
[0116] Comparative Example 8
[0117] Small-sample class-incremental image classification is performed using the same data set as in Example 1, except that the CEC method is used, where CEC is proposed in the literature "Few-Shot Incremental Learning with Continually Evolved Classifiers. In: CVPR.(2021)".
[0118] Experimental Example
[0119] Comparing the results of Example 1 with those of Comparative Examples 1-7, as shown in the table.
[0120] Table 1 Test Performance on CUB200 Dataset
[0121] Method Accuracy rate Comparative example 7 56.43 Example 1 57.52
[0122] Table 1 shows the results of the test performance of Example 1 and Comparative Example 7 on the CUB200 dataset. It can be seen from Table 1 that Example 1 has improved the accuracy by 1.09% (57.52% compared to 56.43%).
[0123] Table 2 Performance of Each Method on CUB200 Dataset
[0124]
[0125] In Table 2, Example 1 was compared with Comparative Examples 1-7 on the CUB200 dataset. It can be seen that Example 1 has improved by 5.58% compared to the highest-performance Comparative Example 8 (57.86% compared to 52.28%), that is, the method in Example 1 is significantly superior to other technologies.
[0126] Figure 3 The visualization results of the feature dimensionality reduction after classification of the methods in Example 1 and Comparative Example 7 are shown. By introducing a small number of parameter control sub-modules and base parameter control sub-modules in Example 1, it effectively reduces the feature drift of the base class during the training process of small-sample new classes, thereby reducing the forgetting of the base class.
[0127] Compare the number of parameters in S2 training and S4 training in Example 1; among them, the parameters of the feature extractor, classifier, base parameters, and control coefficients in step S2 are 11.2M, 0.78M, 9.4M, and 66.6K respectively; in S4, the parameters to be updated are 0.84M, which only accounts for 1 / 25 of the total number of parameters, greatly reducing the overfitting of the model on new classes.
[0128] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", "front", "rear", etc. is the orientation or positional relationship based on the working state of the present invention. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", "third", "fourth" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0129] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0130] The present invention has been described in conjunction with preferred embodiments above, but these embodiments are merely exemplary and only serve an illustrative purpose. On this basis, various substitutions and improvements can be made to the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A method for incremental learning of a small sample of manipulation parameters with a small number of network parameters, characterized in that, It includes the following steps: S1. Set up a few-shot incremental learning neural network, which is obtained by adding a parameter encoding module to a convolutional neural network; S2. Train the few-shot incremental learning neural network with a base class sample set, that is, a set of picture samples with base class labels, to obtain a base class recognition network; S3. Input the picture with the base class label into the base class recognition network, and the base class recognition network outputs the label of the picture, so as to obtain the information in the picture; When it is necessary to identify pictures with new type labels, that is, pictures with new class labels, there are also steps: S4. Train the base class recognition network with a picture sample set with new class labels to obtain an incremental class recognition network; S5. Input the picture with the base class or new class label into the incremental class recognition network, and the incremental class recognition network outputs the label of the picture, so as to obtain the information of the picture; In S1, the parameter encoding module is set between the feature extractor and the classifier of the convolutional neural network; The feature extractor of the convolutional neural network transfers the feature map to the parameter encoding module, encodes the feature map through the parameter encoding module to obtain an encoded feature and transfers it to the classifier, and the classifier outputs a prediction result according to the encoded feature; The parameter encoding module includes a small number of parameter control sub-modules and base parameter control sub-modules; The small number of parameter control sub-module is used to generate the control coefficient α of the picture, and the base parameter control sub-module is used to generate the base parameter θ of the picture γ , and the encoded feature is a linear combination of the control coefficient α and the base parameter θ γ ; The small number of parameter control sub-module includes a pooling layer, a fully connected layer, a non-linear layer, a fully connected layer and a softmax layer connected in sequence, and its output is expressed as: α is a set of control coefficients, x is the feature map passed by the feature extractor, θ α represents a set of network parameters for different layers, where, represents the parameters of the second fully connected layer, represents the parameters of the first fully connected layer, θ α3 represents the bias of the first fully connected layer, represents the bias of the second fully connected layer; The base parameter control sub-module is formed by a single convolution kernel or multiple convolution kernels connected in series, and the set of trainable parameters of the convolution kernel is called the base parameter θ γ .
2. The method for few-shot incremental learning of manipulating few network parameters according to claim 1, characterized in that In S4, the number of samples of any class in the picture sample set with new class labels is less than the number of samples of any class in the base class sample set.
3. The method for few-shot incremental learning of manipulating few network parameters according to claim 1, characterized in that The encoding feature θ e is represented as Among them, α n represents different parameters in the control coefficient set, represents different parameters in the base parameter set, and N is a hyperparameter representing the number of parameters in the set.
4. The method for few-shot incremental learning of manipulating few network parameters according to claim 1, characterized in that In S2, during the training process, the parameters in the feature extractor, the base parameter θ of the base parameter control sub-module γ , and the network parameter θ of the small number of parameter control sub-module α are all trainable parameters.
5. The method for few-shot incremental learning of manipulating few network parameters according to claim 1, characterized in that In S4, during the training process, the parameters in the feature extractor and the base parameter θ of the base parameter control sub-module γ remain unchanged, and the network parameter θ of the small number of parameter control sub-modules α is a trainable parameter.
6. The method for few-shot incremental learning of manipulating few network parameters according to claim 1, characterized in that In S4, a distillation loss function is adopted during the training process.
7. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-6.
8. A computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to make the computer execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Small sample class incremental learning method based on feature space combination
CN111931807A
X-ray chest radiography image classification method based on small sample learning and self-supervised learning
CN112348792A