A method for fast image data classification that can continuously learn

By using the pre-trained Resnet152 network and incremental classifier in the robot system, the problem that the robot system is difficult to quickly learn new data and maintain old data in the continuous learning mode in the dynamic environment, achieving the effect of rapid learning and classification.

CN113239974BActive Publication Date: 2025-06-24COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110427115.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-21
Publication Date
2025-06-24
Estimated Expiration
2041-04-21

AI Technical Summary

Technical Problem

In a dynamic environment, it is difficult for robot systems to quickly learn new data and complete classification decisions in continuous learning mode, and catastrophic forgetting of old data will also occur.

Method used

The feature vectors output from the adaptive average pooling layer of image data are extracted using the Resnet152 network pre-trained on the Imagenet dataset, and a new encoding is obtained through nonlinear transformation and binarization processing. Then, an incremental classifier is built, and its structure and weights are updated dynamically to enable rapid learning of new class data and retention of old class data.

Benefits of technology

It realizes rapid learning and classification of image data in dynamic environments, avoids catastrophic forgetting of old data, and significantly improves training speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113239974B_ABST
    Figure CN113239974B_ABST
Patent Text Reader

Abstract

The present invention discloses a fast image data classification method capable of continuous learning, including: for the image data to be classified, learning the sample binary features through a deep neural network; dynamically determining the binary coding information of the newly added categories according to the number of categories; establishing a classification neural network between the feature vectors and the categories, dynamically adjusting the connection weights, and realizing the pattern classification of the newly added categories on the premise of minimizing the impact on the existing classification neural network. According to the class incremental learning method of the present invention, the robot system can achieve fast continuous learning and classification of incremental image data in a dynamic environment, avoid catastrophic forgetting in the incremental learning process, and greatly shorten the time for training the incremental classifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of robots, and particularly to a fast image data classification method that can continuously learn in an open environment. Background Art

[0002] Deep learning technologies based on the backpropagation algorithm have achieved a recognition level comparable to that of humans in perception classification tasks in closed scenarios, but their continuous learning performance in dynamic environments is far inferior to that of humans. Continuous learning, or even lifelong learning, in dynamic environments is the key to the development of robot technologies.

[0003] Continuous learning is a special machine learning paradigm, and its special feature lies in that it does not satisfy the independent and identically distributed assumption of training data and test data in machine learning. The purpose of continuous learning is to enable the system to perform more flexible online learning and avoid catastrophic forgetting of the learned knowledge.

[0004] The catastrophic forgetting problem of neural networks can be traced back to the late 1980s. Specifically, it refers to the fact that when the network learns new knowledge in a serialized manner, it will cause serious interference to the old knowledge, resulting in catastrophic forgetting of the old knowledge. The essence of the catastrophic forgetting problem is the synaptic plasticity (weight learning / updating) problem of neural networks. The backpropagation algorithm commonly used in current deep learning is a global weight update strategy, which will make the network weights too plastic when learning new knowledge, thus bringing catastrophic forgetting of the old knowledge. How to avoid the catastrophic forgetting brought by the global update of the backpropagation algorithm in the learning process of neural networks is a major challenge faced by the continuous learning paradigm.

[0005] To solve the catastrophic forgetting problem of neural networks and better achieve the continuous learning of the system, researchers in the industry have proposed solutions from many aspects (such as network regularization, dynamic network structure, memory-based replay, brain-inspired complementary learning theory, etc.). However, the above solutions still have problems such as low accuracy and slow training speed in large-scale class incremental learning tasks with more practical application value.

[0006] In the mode of continuous learning, how to enable the robot system to efficiently and quickly learn new data to complete classification decisions and not have catastrophic forgetting of the old data is an urgent problem to be solved. Summary of the Invention

[0007] In view of the problems existing in the prior art, an embodiment of the present invention provides a fast image data classification processing method capable of continuous learning to solve the problem of fast image data classification in a robot system in a continuous learning mode. Compared with the classical backpropagation algorithm, the present invention does not rely on the independent and identically distributed assumption of data and does not need to update the calculation weights in an iterative manner. Therefore, the present invention has a huge advantage in training speed.

[0008] The present invention provides a fast image data classification method capable of continuous learning, and the method includes:

[0009] 1. Extract the feature vector (Feature Vector, abbreviated as FV) output by the adaptive average pooling layer before the classification layer of the data to be classified by using the Resnet152 network pre-trained on the Imagenet dataset, denoted as FV ∈ R N×1 , N = 2048.

[0010] For the image data to be classified, the process of obtaining the FV of the image is as follows: First, convert the input image into an image with 3 RGB channels, and linearly normalize the pixel values in the range [0, 255] to the interval [0, 1]; standardize the three channels of the image respectively, specifically: first subtract the mean of each channel, and then divide by the mean square deviation of each channel; for the natural image I obtained by the above processing, input I into the Resnet152 network for forward propagation calculation, and obtain the output vector of the last adaptive average pooling layer of the network as the FV of the image I.

[0011] 2. Perform a non-linear transformation on the elements in the FV of the image I in step 1 to obtain a new code, denoted as SFV ∈ R N×1 ; and binarize SFV to obtain a binarized code, denoted as BFV ∈ R N×1 . BFV(i) represents the activation value of the i-th neuron. BFV(i) = 1 indicates that the i-th neuron participates in encoding the information of the image I; BFV(i) = 0 indicates that the i-th neuron does not participate in encoding the information of the image I.

[0012] The non-linear transformation described here includes: left-multiplying FV ∈ R N×1 by a matrix W ∈ R N×N for linear transformation to obtain a new N-dimensional vector, denoted as: NFV ∈ R N×1 ; then substitute each element in the NFV vector into the sigmoid function to obtain a new element with a value range in the interval (0, 1), and obtain: SFV ∈ R N×1 . The operation of binarizing SFV is: sequentially judge all elements in SFV. If the element value is greater than or equal to 0.5, set it to 1, otherwise set it to 0. The vector composed of binarized elements is called: BFV ∈ RN×1 .

[0013] 3. Remove the end-to-end classifier of Resnet152 and build an incremental classifier.

[0014] The structure of the classifier is an artificial neural network with two fully connected layers, where the first layer is the input layer and the second layer is the classification layer, and the input layer and the classification layer are fully connected.

[0015] An incremental classifier refers to a classifier that can automatically add output neurons based on the number of newly learned categories on the basis of the original classifier structure, that is, the structure of the incremental classifier has dynamic scalability.

[0016] 4. Dynamically update the structure of the classifier when learning new categories of data.

[0017] Dynamic means: for each new category of data, a new output neuron corresponding to that category is added;

[0018] 5. When a new category of training data appears (referred to as category C), the binary encoding BFV of all training data of category C is obtained through the above 1 and 2, and the diversity and specificity of each neuron in the binary encoding layer are updated. The vector composed of the diversity of all neurons in the binary encoding layer is recorded as: DV∈R N×1 ; The vector composed of the specificity of all neurons in the encoding layer is recorded as: SP∈R N×1 .

[0019] Here, neuron diversity refers to the number of different categories that the neuron participates in encoding. Here, "participating in encoding" specifically means: if there is a sample I' with BFV(i)=1 in the training set of category C, then the neuron i participates in encoding category C. The specificity of a neuron is the inverse of neuron diversity, satisfying: It is easy to see that the greater the diversity of neurons in the binary coding layer, the smaller its specificity.

[0020] 6. Obtain the binary encoding BFV of all training data of category C through the above 1 and 2, and calculate the average encoding vector of all binary encodings of category C, recorded as: ACV C ∈R N×1 ;

[0021] Remember ACV C (i) is the activation value of the i-th (1≤i≤N) neuron in class C. Its physical meaning is: the frequency of the i-th neuron in the encoding vector participating in encoding class C, ACV C (i) can also be understood as the weight of the i-th neuron's specificity to class C.

[0022] 7. Complete the online update of the weights of all neurons on the incremental classifier, including the weights of the positive receptive fields and the weights of the negative receptive fields.

[0023] Here, denote the positive receptive field of class C as The negative receptive field of class C as The positive receptive field of class C refers to: the union of the positive receptive fields of all binarized encodings under class C. The above-mentioned positive receptive field of a single binarized encoding is the union of the elements with an activation value of 1 in this binarized encoding. The negative receptive field of class C refers to: the union of the parts that do not belong to in the positive receptive fields of all classes K (K≠C) that have a union with the positive receptive field of class C . Satisfy:

[0024] The above method can achieve the rapid learning and classification of the system for incremental image data. Brief Description of the Drawings

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0026] Figure 1 It is a schematic diagram of the weight update of the incremental classifier in the fast image data classification method according to the embodiment of the present invention.

[0027] Figure 2 Summary chart of the final correct rates of continuous learning of the method (NAL) of the present invention and the OWM method based on the backpropagation algorithm on a large-scale dataset

[0028] Figure 3 Statistics of the continuous learning time of the method (NAL) of the present invention and the OWM method based on the backpropagation algorithm on a large-scale dataset

[0029] Figure 4 It is a schematic diagram of the brief process of the present invention Detailed Embodiments

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] The present invention proposes a fast image data classification method capable of continuous learning, and the method includes the following steps:

[0032] S1: Read all training samples of the new class of the image to be learned, and obtain the high-level encoding of the sample to be classified through a deep neural network

[0033] Specifically, a deep neural network model pre-trained on a large-scale image dataset ImageNet (for example: Resnet152) is used as the feature extractor of the sample, and the input vector of the classification layer of the deep neural network model is recorded as the high-level encoding of the sample. The high-level encoding is recorded as FV ∈ R N×1 , N = 2048.

[0034] S2: Perform a non-linear transformation on the high-level encoding to convert it into a binary encoding

[0035] Since the elements FV i (1 ≤ i ≤ N) in FV obtained in S1 are generally not binary. To obtain a binary encoding, here the binary encoding is denoted as BFV ∈ R N×1 , and the following method is used to obtain it:

[0036] 1. Multiply FV on the left by a matrix W ∈ R N×N , and obtain a new encoding NFV ∈ R N×1 , NFV = W · FV,;

[0037] 2. Perform a non-linear transformation on each element in NFV to obtain a new encoding SFV ∈ R N×1 , and the transformation formula is as follows:

[0038] where SFV i refers to the i-th element in the SFV vector, and NFV i refers to the i-th element in the NFV vector;

[0039] 3. Binarize the elements in NFV to obtain BFV ∈ R N×1 , and the binarization formula is as follows:

[0040] where BFV i refers to the i-th element in the BFV vector, and SFVi Refers to the i-th element in the SFV vector.

[0041] S3: Structural update of the incremental classifier

[0042] Specifically: When learning a new class of image samples, the structure of the classifier is updated, which is reflected in adding a corresponding output neuron for each new class.

[0043] S3: Weight update of the incremental classifier

[0044] To better describe the following weight update method of the incremental classifier, the definitions of the basic concepts involved in this method are given as follows:

[0045] Definition 1: The diversity of neurons in the binary coding layer is the number of different classes it participates in coding. The vector composed of the diversities of neurons in the binary coding layer is called the diversity vector, denoted as DV ∈ R N×1 .

[0046] Definition 2: The specificity of neurons in the binary coding layer is the reciprocal of the neuron diversity. The vector composed of the specificities of neurons in the binary coding layer is called the specificity vector, denoted as SP ∈ R N×1 ; where:

[0047] It is easy to know that in the binary coding layer, the greater the diversity of neurons, the smaller its specificity. Definition 3: The specificity weight of class C is the average coding vector ACV of all samples within this class C ∈ R N×1 , where N is the dimension of the coding vector. Denote: ACV C (i) represents the specificity weight of the i-th (1 ≤ i ≤ N) neuron in subclass C; for example: if the i-th neuron appears 6 times in 10 C-class samples, then

[0048] Definition 4: The positive receptive field of the label neuron of class C is: the union of the positive receptive fields of all binary codings under class C. The size of the positive receptive field of class C is denoted as CS C .

[0049] The steps of the fast weight update algorithm for the incremental classifier are as follows:

[0050] 1. Obtain all training samples of class K, and calculate the positive receptive field of class K according to Definition 4

[0051] 2. Calculate the average coding vector ACV of class K according to Definition 3 K , and let ACV KStored in the ACV_List;

[0052] 3. Update according to Definition 2 The specificity of neurons in;

[0053] 4. According to the SP value in 3, calculate / update the positive receptive field weights of classes L = 1, 2, …, K in sequence.

[0054] 5. Update and calculate the negative receptive fields of classes L (L = 1, 2, …, K - 1) and the weights of the negative receptive fields.

[0055] 6. Update and calculate the negative receptive fields of class K and the weights of the negative receptive fields.

[0056] 7. Update completed

[0057] In the above method, the incremental classifier can add output units according to new class data and quickly fine-tune the weights of the classification layer, realizing fast learning and decision-making for new classes, and will not cause catastrophic forgetting of the decisions of old classes.

[0058] The embodiment of the present invention provides a fast image data classification method capable of continuous learning. The method includes: for the image data to be classified, acquiring sample binary features through a deep neural network; dynamically determining the binary coding information of newly added classes according to the number of classes; establishing a classification neural network between the feature vectors and classes, dynamically adjusting the connection weights, and realizing pattern classification of newly added classes on the premise of not affecting the existing classification neural network as much as possible; the device executes the above method. According to the class incremental learning method and device of the present invention, fast continuous learning of class incremental data in a dynamic environment can be realized, avoiding the catastrophic forgetting bottleneck in the class incremental learning process, and greatly shortening the training time of the incremental classifier.

[0059] Furthermore, the present invention can also be applied to the following aspects:

[0060] (1) Access control that requires fast face recognition. For example: in relevant access controls that adopt human recognition technology, the system needs to learn the face features of a large number of new people online in real time and complete the verification of the faces of new people. At this time, the fast pattern classification algorithm and device of the present invention can be used to implement this function online, without having to summarize the samples of all face data and then use the traditional backpropagation algorithm to optimize a specific cost function to obtain the classifier weights, greatly saving the training time.

[0061] (2) E-commerce applications that require rapid identification and classification of newly added products. In this case, the present invention can be used to perform online learning on newly added products without having to aggregate samples of all products and then use the traditional backpropagation algorithm to optimize a specific cost function to obtain classifier weights, thus greatly saving training time.

[0062] (3) Multimedia data that requires online rapid classification, including image data, speech data, video data, etc. Using the proposed continuous learning-based fast data classification method of the present invention, rapid learning and classification of incremental multimedia data can be performed without having to aggregate samples of all multimedia data to be classified and then use the traditional backpropagation algorithm to optimize a specific cost function to obtain classifier weights, thus greatly saving training time.

[0063] Figure 2 Shows that: After the proposed method of the present invention (NAL; orange bars) has continuously learned samples of all categories on the large-scale natural image dataset ImageNet1000 and the large-scale Chinese handwritten character dataset CASIA3755, the classification recognition accuracy of the network for the data test set (the closer the accuracy is to 1, the better the algorithm performance). Another method, the OWM method, is a continuous learning method based on the backpropagation algorithm.

[0064] Figure 3 Shows that: A comparison of the training time taken for continuous learning of the proposed method of the present invention and the OWM method on the large-scale natural image dataset ImageNet1000 and the large-scale Chinese handwritten character dataset CASIA3755. It can be seen from the figure that the proposed method of the present invention has a significant advantage in learning rate compared to the OWM method.

Claims

1. A fast image data classification method capable of continuous learning, characterized by: 1) Extract the feature vector FV output by the adaptive average pooling layer before the Resnet152 classification layer for the data to be classified using the Resnet152 network pre-trained on the Imagenet dataset. FV ∈ R N×1 , N = 2048; For the image data to be classified, the process of obtaining the FV of the image is as follows: first, the input image is converted into an image with RGB three channels, and the pixel values ​​in the range of [0,255] are linearly normalized and converted to the interval of [0,1]; the three channels of the image are standardized respectively, specifically: first subtract the mean of each channel, and then divide by the mean square error of each channel; for the obtained natural image I, I is input into the Resnet152 network for forward propagation calculation, and the output vector of the last adaptive average pooling layer of the network is obtained as the FV of image I; 2) Non-linearly transform the elements within the FV of the image I to obtain a new code, denoted as SFV ∈ R N×1 ; and binarize SFV to obtain a binarized code, denoted as BFV ∈ R N×1 ; where BFV(i) represents the activation value of the i-th neuron, BFV(i) = 1 indicates that the i-th neuron participates in encoding the information of the image I; BFV(i) = 0 indicates that the i-th neuron does not participate in encoding the information of the image I; Perform a non - linear transformation: For FV ∈ R N×1 Pre - multiply by a matrix W ∈ R N×N Perform a linear transformation to obtain a new N - dimensional vector, denoted as: NFV ∈ R N×1 ; Then substitute each element in the NFV vector into the sigmoid function to obtain new elements in the range (0, 1), getting: SFV ∈ R N×1 ; The operation of binarizing SFV is: successively judge all elements in SFV. If the element value is greater than or equal to 0.5, set it to 1, otherwise set it to 0. The vector composed of binarized elements is called: BFV ∈ R N×1 ; 3) Remove the end-to-end classifier of Resnet152 and build an incremental classifier; The structure of the classifier is an artificial neural network with two fully connected layers, where: the first layer is the input layer, the second layer is the classification layer, and the input layer and the classification layer are fully connected; An incremental classifier is a classifier that can automatically add output neurons based on the number of newly learned categories on the basis of the original classifier structure, that is, the structure of the incremental classifier has dynamic scalability; 4) Dynamically update the structure of the classifier when learning new categories of data; here, Dynamic means: for each new category, a new output neuron corresponding to that category is added; 5) When new categories of training data appear, denoted as category C, obtain the binary encoding BFV of all training data of category C through steps 1) and 2), update the diversity of each neuron in the binary encoding layer and the specificity of each neuron, and the vector composed of the diversity of all neurons in the binary encoding layer is denoted as: DV ∈ R N×1 ; the vector composed of the specificity of all neurons in the encoding layer is denoted as: SP ∈ R N×1 ; Here, neuron diversity refers to the number of different categories encoded by the neuron. Here, "participating in encoding" specifically means that if BFV(i) = 1 for a sample I' in the category C training set, then the neuron i participates in encoding the category C; the specificity of a neuron is the reciprocal of the neuron diversity and satisfies: It is easy to know that the greater the diversity of the neurons in the binary encoding layer, the smaller its specificity; 6) By obtaining the binary encoding BFV of all training data of class C and calculating the average encoding vector of all binary encodings of class C, denoted as: ACV C ∈R N×1 ; Denote ACV C (i) is the activation value of the i-th neuron in class C, where 1 ≤ i ≤ N; its physical meaning is: the frequency at which the i-th neuron in the encoding vector participates in encoding class C, ACV C (i) is understood as the weight of the i-th neuron for the specificity of class C; 7) Complete the online update of the weights of all neurons on the incremental classifier, including the weights of the positive receptive field and the weights of the negative receptive field; Here, denote the positive receptive field of class C as and the negative receptive field of class C as The positive receptive field of class C refers to: the union of the positive receptive fields of all the binarized encodings under class C, where the positive receptive field of a single binarized encoding is the union of the elements with activation value 1 in this binarized encoding; the negative receptive field of class C refers to: the union of the parts that do not belong to in the positive receptive fields of all classes K that have a union with the positive receptive field of class C and K≠C ; satisfying:

Citation Information

Patent Citations

  • Incremental learning image classification training method under big data scene

    CN107358257A

  • Lung cancer image pathological classification method and device

    CN110472694A