An image continuous learning and recognition method, device and equipment

Through the continuous image learning method, feature extraction sharing and knowledge transfer learning are used to solve the problems of waste and catastrophic forgetting of traditional image recognition resources, and efficient multi-task image recognition and continuous learning are achieved.

CN115630694BActive Publication Date: 2025-07-08WUHAN TEXTILE UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211287865.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-20
Publication Date
2025-07-08
Estimated Expiration
2042-10-20

AI Technical Summary

Technical Problem

In traditional image recognition methods, each task needs to train a separate model, resulting in tight computing and storage resources, and multi-task learning fails to imitate human continuous learning methods, which easily leads to catastrophic forgetting and task drag.

Method used

The image continuous learning method is adopted, through feature extraction sharing and knowledge transfer learning, and unsupervised pre-training is used to use self-coding networks to perform unsupervised pre-training, combining BP backpropagation algorithms and lifelong machine learning to build a knowledge warehouse to realize knowledge transfer and feedback, and avoid catastrophic forgetting.

Benefits of technology

It realizes the use of a model to recognize any kind of image, reduce resource waste, improve learning efficiency and recognition accuracy, and keep the performance of old tasks from being affected by new tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115630694B_ABST
    Figure CN115630694B_ABST
Patent Text Reader

Abstract

The present invention provides an image continuous learning and recognition method, apparatus and device. By using this continuous learning method, an infinite number of image recognition learning tasks can be achieved, mainly including two major steps: feature extraction sharing and knowledge transfer learning. Among them, feature extraction sharing consists of unsupervised pre-training and supervised tuning, and knowledge transfer learning consists of knowledge warehouse construction, knowledge transfer, knowledge base update, model learning and receiving feedback. The present invention can use the knowledge transfer technology to continuously train and learn image recognition using one model, and finally can achieve the recognition task of any type of image using one model, and the effect of the present invention is verified through experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machine learning, and particularly relates to the fields of image processing and pattern recognition. Background Art

[0002] When performing image recognition tasks, the traditional approach is to train a deep neural network model for each task to break them down one by one. The drawbacks of this approach are: ① Each task is isolated, and it is impossible to achieve learning by analogy (cooperative learning and transfer learning). ② It will cause problems of tight computing and storage resources. If there are 1,000 tasks, 1,000 models need to be trained and stored. To address the above problems, some people have proposed a multi-task deep learning model to use one model to perform multi-task cooperative learning simultaneously. However, since the difficulty levels of different task recognitions are different, once the training samples are insufficient or the loss function design is unreasonable, it is easy for the more difficult classification tasks to drag down the easier classification tasks. In addition, since it is easier to collect single-task training samples than multi-task training samples, it has hindered the application of multi-task deep learning. On the other hand, multi-task learning does not imitate the way humans continuously learn (using the knowledge of old tasks for transfer learning when learning new tasks), resulting in the obstruction of its learning ability and application scenarios.

[0003] In order to apply the knowledge or patterns learned in a certain field or task to different but related fields or problems, scholars have developed transfer learning techniques. That is, transfer the labeled data or knowledge structure from related fields to complete or improve the learning effect of the target field or task. As knowledge is continuously accumulated in learning new tasks, traditional transfer learning transfers the knowledge of old tasks to new tasks only to better recognize new tasks without caring about old tasks. Therefore, traditional transfer learning will lead to the problem of "catastrophic forgetting", that is, learning new tasks and forgetting old tasks. Summary of the Invention

[0004] To solve the above problems, the present invention proposes an image continuous learning and recognition method, device and equipment. Using this continuous learning method, an infinite number of image recognition learning tasks can be achieved, such as Figure 1 As shown, it mainly includes two major steps: feature extraction sharing and knowledge transfer learning. Among them, feature extraction sharing consists of unsupervised pre-training and supervised tuning, and knowledge transfer learning consists of knowledge warehouse construction, knowledge transfer, knowledge base update, model learning, and receiving feedback.

[0005] To achieve the above object, the technical solution provided by the present invention is: an image continuous learning and recognition method, including two major steps: feature extraction sharing and knowledge transfer learning. Among them, feature extraction sharing includes the following sub-steps:

[0006] S11, unsupervised pre-training, using the unsupervised reconstruction pre-training method in the autoencoder network to train and obtain a feature extraction sharing model;

[0007] S12. Maintain and update the feature extraction shared model, and use the BP backpropagation algorithm to fine-tune and update the feature extraction shared model generated in step S11;

[0008] S13. Use the updated feature extraction shared model to perform layer-by-layer feature extraction on the image X to be recognized (t) to obtain the image feature H (t) , where t represents the t-th recognition task;

[0009] S14. Decompose and solve the model parameters θ (t) for the feature H of the current task (t) and the knowledge base matrix L;

[0010] Knowledge transfer learning includes knowledge repository construction, knowledge transfer, model learning, updating the knowledge base, and receiving feedback;

[0011] Among them, in knowledge transfer, the feature H extracted by using the feature extraction shared module (t) is used to construct a new task Z (t) =(H (t) , y (t) ), where y (t) is the label. Subsequently, transfer learning is performed on this set of training data using the lifelong machine learning algorithm to obtain the final image recognition learning model f (t) (θ) of each task, and image recognition is performed using the learning model.

[0012] Furthermore, in S1, first randomly initialize the autoencoder network model parameters W l , b l , and use the unlabeled training data set where T c is the number of tasks to be recognized in the task pool, and use the contrastive divergence algorithm to optimize and update the parameters W l , b l one by one from bottom to top; the contrastive divergence algorithm first uses the Gibbs sampling method to obtain the hidden layer features H l-1,0 , H l,0 , H l-1,1 , H l,1 after three samplings, and then updates the matrix W l ←W l +α(H l-1,0 T H l,0 -H l-1,1 T H l,1 ), and l update b update b l, where \(T\) represents the transpose of the matrix, \(\alpha\) is the learning rate for network pre-training, and \(loss(\cdot)\) is the loss function. denotes the partial derivative.

[0013] Furthermore, in step S12, after pre-training is completed, continuous image recognition tasks are started. Set the current task to be recognized as task \(t\), and the representation of task \(t\) is \(Z\) (t) =(X (t) , y (t) ), where \(X\) (t) is the image sample in task \(t\), and \(y\) (t) is its corresponding label; use this data to fine-tune and update the feature extraction shared model generated in step S11 to avoid affecting the representativeness of features due to distribution drift, and use the BP backpropagation algorithm to update the feature model;

[0014] First, calculate the features of each layer from bottom to top for the input \(X\) (t) to obtain \(H\) l , \(l = 1, 2, 3, \cdots, n\) L , \(n\) L represents the maximum number of layers. Use the single-task learner to learn \((H\) l , y (t) ) to obtain the parameter \(W0\), and apply the BP algorithm to optimize the network parameters; for each layer \(l\), calculate \(\cup\) l =H l *(1 - H l ) \(\cup\) l+1 W l+1 T . If it is the last layer, then \(\cup\) l =H l *(1 - H l ) y (t) . Through \(W\) l ←W l - \(\alpha\) t H l-1 T \(\cup\) l update the matrix \(W\) l ;

[0015] Considering avoiding the occurrence of task negative transfer, the learning rate \(\alpha\) t of the model changes with task relevance, that is, for each task \(t\), calculate its relevance \(\gamma\) t with the previous task:

[0016] \(\gamma\) t =cos(\(\theta\) (t) , \(\theta\) (t-1) ) = cos(Ls (t) , Ls (t-1) ) = cos(s(t) , s (t-1) )

[0017] Among them, is the knowledge base matrix, s (t) is the linear combination coefficient; the parameters θ of each task model (t) are linearly combined by some column vectors in the L matrix, that is, θ (t) = Ls (t) , cos() represents the cosine angle between two vectors. The smaller the angle between the partitioning hyperplanes of the task models, the higher the task correlation; further, the task learning rate α t is determined by the following formula:

[0018] α t = α c (γ t + 1) / 2

[0019] where α c is the benchmark learning rate for BP tuning and is a hyperparameter that needs to be tuned manually.

[0020] Further, in step S14, for the input (H (t) , y (t) ), the model parameters θ (t) and the Hessian matrix D (t) can be obtained in a closed form by the following formula:

[0021] θ (t) = (H (t) H (t)T ) -1 H (t) y (t)

[0022]

[0023] Subsequently, optimize s from (t) , where μ is a balance factor and is calculated using the stochastic gradient descent method Then update the shared task base L, that is, the knowledge base matrix, through , where T represents the maximum number of tasks, and n t represents the number of samples corresponding to the t-th task.

[0024] Further, assume there are M image recognition tasks T (1) , T (2) , … T (M) and a large amount of labeled data that conforms to the learned task distribution. Each task T (m) = (f (m) (θ), H (m) , y (m)), where H (m) is the feature of the training sample of task m, and y (m) is the corresponding sample label, and f (m) (θ) is the recognition function of task m, and θ is the network parameter of the learning model; that is, the label of each task is determined by a true implicit function f (m) (θ), which is continuously updated and optimized as the recognition tasks increase; each task is given n m training samples and the corresponding label y (m) , d is the dimensionality of the sample features extracted by the feature extraction sharing module; at any point in time, when the image recognition learning model receives a batch of labeled samples from task m, they may be from a new task or a task that has been learned. After receiving the training data, the goal of the learning model is to learn the learning model of each task through the given training samples as close as possible to the true objective function f (m) ; the objective function of the learning model is:

[0025]

[0026] where σ(·) is the activation function; represents the feature representation of the i-th input of the m-th task at the l-th layer, is the label of the i-th sample in task m, and n m represents the number of training samples of task m, and W l represents the weight connection matrix of the l-th layer, is the knowledge base matrix, and s (m) is the linear combination coefficient; the parameters θ (m) of each task model are linearly combined by some column vectors in the L matrix, that is, θ (m) = Ls (m) ; both μ and λ are balance parameters.

[0027] Furthermore, in order to perform knowledge transfer and avoid catastrophic forgetting, for each image recognition task, after training the task, calculate the importance Ω ij of each parameter in the network for this task, that is, the proportion of the parameter value in the i-th row and j-th column to the entire parameter value, and apply it to the subsequent tasks during training. Ω ij is added to the loss function in the form of a regularization term. Whenever a new task is trained: for the parameters with a large Ω ij , try to reduce the change amplitude of it during gradient descent because this parameter is very important for a certain past task and its value needs to be retained to avoid catastrophic forgetting; while for the Ω ijSmaller parameters can be updated with a larger gradient amplitude to achieve better performance on new tasks. Therefore, the loss function for the m-th task is as follows:

[0028]

[0029] where L (m) (θ) is the loss function of the current task, λ is the balance factor, and θ ij is the parameter in the i-th row and j-th column of the current model parameter matrix. is the model parameter obtained after training with the previous n - 1 tasks. Whenever a task is trained, Ω ij will be updated. When performing the first task, it takes 0.

[0030] The present invention also provides an image continuous learning and recognition device, including a feature extraction sharing module and a knowledge transfer learning module. The feature extraction sharing module includes the following units:

[0031] An unsupervised training unit for performing unsupervised pre-training and training a feature extraction sharing model using the unsupervised reconstruction pre-training method in an autoencoder network.

[0032] An update unit for maintaining and updating the feature extraction sharing model and fine-tuning and updating the feature extraction sharing model generated in step S11 using the BP backpropagation algorithm.

[0033] A feature extraction unit for using the updated feature extraction sharing model to perform layer-by-layer feature extraction on the image X to be recognized (t) to obtain image features H (t) , where t represents the t-th recognition task.

[0034] A model parameter solving unit for decomposing and solving the model parameter θ (t) and the knowledge base matrix L for the features H of the current task. (t) of the current task.

[0035] Among them, in the knowledge transfer learning module, the features H extracted using the feature extraction sharing module are used (t) to construct a new task Z (t) =(H (t) , y (t) ), where y (t) is the label. Subsequently, transfer learning is performed on this set of training data using the lifelong machine learning algorithm to obtain the final image recognition learning model f (t) (θ) for each task, and image recognition is performed using the learning model.

[0036] The present invention also provides an electronic device, including a distributed memory, a processor, and a computer program that can run in the processor in the memory. When the processor executes the computer program, the steps of an image continuous learning and recognition method described in the above solution are implemented.

[0037] The present invention also provides a computer-readable storage medium storing computer software instructions, and the computer software instructions implement the steps of an image continuous learning and recognition method described in the above solution.

[0038] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:

[0039] In the traditional image recognition method, one model can only recognize one or several types of images. If there are thousands of types of images to be recognized, thousands of models need to be developed, resulting in serious waste of resources. Obviously, the traditional image recognition method does not conform to the pattern of human image recognition. Humans can easily recognize various types of images with the same brain. The image continuous recognition and learning method and device proposed by the present invention can use the knowledge transfer technology to continuously train and learn image recognition using one model, and finally can use one model to achieve the recognition task of any type of image. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a framework diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] The method of the present invention includes two major steps: feature extraction sharing and knowledge transfer learning. Among them, S1, feature extraction sharing mainly includes four steps:

[0043] S11. Unsupervised pre-training

[0044] Before receiving continuous image recognition tasks, a large number of unlabeled data related to the recognition tasks are sampled from the input data space distribution, and then an unsupervised reconstruction pre-training method in an autoencoder network is used to train a feature extraction sharing model, which will be applied to step S13 for feature extraction.

[0045] First, randomly initialize the autoencoder network model parameters W l , b l , and use the unlabeled training data set where T c(where is the number of task pools to be recognized), use the contrastive divergence algorithm to optimize and update the parameter W layer by layer from bottom to top l , b l . The contrastive divergence algorithm first uses the Gibbs sampling method to obtain the hidden layer features H after three samplings l-1,0 , H l,0 , H l-1,1 , H l,1 , and then update the matrix W through W l ← W l + α(H l-1,0 T H l,0 - H l-1,1 T H l,1 ), where T represents the transpose of the matrix, α is the learning rate of network pre-training, loss(·) is the loss function, l , Update b l , where, T represents the transpose of the matrix, α is the learning rate of network pre-training, loss(·) is the loss function, represents the partial derivative.

[0046] S12. Maintain and update the feature extraction shared model

[0047] After pre-training, the learning machine starts to continuously perform image recognition tasks. Assume that the current task to be recognized is task t, and the representation of task t is Z (t) = (X (t) , y (t) ), where, X (t) is the image sample in task t, and y (t) is its corresponding label. Use these data to fine-tune and update the feature extraction shared model generated in step S11 to avoid affecting the representativeness of features due to distribution drift. Use the BP backpropagation algorithm to update the feature model:

[0048] First, calculate the features of each layer from bottom to top for the input X (t) to obtain H l (l = 1, 2, 3,..., n L ), n L represents the maximum number of layers. Use the single-task learner to learn (H l , y (t) ) to obtain W0, and use the BP algorithm to optimize the network parameters. For each layer l, calculate ∪ l = H l * (1 - H l ) ∪ l+1 W l+1 T . If it is the last layer, then ∪ l = H l * (1 - H l ) y(t) , through W l ←W l -α t H l-1 T ∪ l Update matrix W l .

[0049] Considering avoiding the occurrence of task negative transfer, the learning rate α of the model t varies with the task correlation. That is, for each task t, calculate its correlation γ with the previous task t :

[0050] γ t =cos(θ (t) ,θ (t-1) ) = cos(Ls (t) ,Ls (t-1) ) = cos(s (t) ,s (t-1) )

[0051] where is the knowledge base matrix, s (t) is the linear combination coefficient; the parameters θ of each task model (t) are linearly combined by some column vectors in the L matrix, that is, θ (t) =Ls (t) , and cos() represents the cosine angle between two vectors. The smaller the angle between the partition hyperplanes of the task models, the higher the task correlation. Further, the task learning rate α t is determined by the following formula:

[0052] α t =α c (γ t +1) / 2

[0053] where α c is the benchmark learning rate for BP tuning and is one of the hyperparameters that need to be tuned manually. The main function of the above formula is to adjust the task correlation parameter γ t to the interval [0,1] to meet the requirements of the learning rate of the BP algorithm. After introducing the task-related learning rate, the network can effectively avoid overlearning for a small number of unrelated tasks and prevent the shared feature network from deviating from the general direction of the data.

[0054] S13. Feature extraction

[0055] For each learning task image X about to enter the knowledge transfer learning module (t) , use the updated feature extraction shared model to perform layer-by-layer feature extraction on X (t) to obtain the image feature H (t) .

[0056] After the feature network is tuned using the BP algorithm, re-input the current task X (t) Calculate the features of each layer bottom-up once to obtain H l (l = 1, 2, 3, …, n L ) and define the following feature extraction function:

[0057] For l = 1:n L

[0058] Sample H l from the distribution P(H l-1 ) = σ(b l + W l H l-1 ), where σ(·) is the activation function; l

[0059] End

[0060] The finally obtained H nL is the feature H (t) of task t.

[0061] S14. Update the top-level parameter sharing model

[0062] First, solve the model parameters θ (t) and the Hessian matrix D (t) for the feature H (t) of the current task. θ (t) is the parameter of the top-level parameter sharing model. For the input (H (t) , y (t) ), θ (t) , D (t) can be obtained by the following closed-form formula:

[0063] θ (t) = (H (t) H (t)T ) -1 H (t) y (t)

[0064]

[0065] Subsequently, optimize s (t) from , where μ is a balancing factor, and calculate it using the stochastic gradient descent method Then update the shared task base L, that is, the knowledge base matrix, through . T represents the maximum number of tasks, and n t represents the number of samples corresponding to the t-th task.

[0066] ​Since a large amount of unlabeled data X is effectively utilized before the start of the learning process u The information of is used to construct a network feature extraction model through unsupervised training, so that the labeled training data (X (t) , y (t) ) after feature extraction conforms to the same distribution more, making the initial value of the knowledge base better represent and fit the distribution of task parameters; at the same time, it also makes the subsequent training process easier to converge. Experiments show that even when reducing the amount of labels in each task, due to the features obtained through feature extraction carrying more structured information than the original data and conforming to the same distribution more, the algorithm still maintains good performance.

[0067] S2. Transfer learning adopts a supervised lifelong machine learning system, and the features H obtained by using the feature extraction shared module (t) are used to construct a new task Z (t) =(H (t) , y (t) ). Subsequently, transfer learning is performed on this set of training data using the lifelong machine learning algorithm, including knowledge warehouse construction, knowledge transfer, model learning, updating the knowledge warehouse, and receiving feedback.

[0068] Suppose there are M image recognition tasks T (1) , T (2) , … T (M) and a large amount of labeled data that conforms to the distribution of the learned tasks. Each task T (m) =(f (m) (θ), H (m) , y (m) ), where H (m) is the feature of the training samples of task m (obtained using the feature extraction shared module), y (m) is the corresponding sample label, f (m) (θ) is the recognition function of task m, and θ is the network parameter of the learning model. That is, the label of each task is determined by a true implicit function f (m) (θ), which is continuously updated and optimized as the recognition tasks increase; each task is given n m training samples (d is the dimensionality of the sample features extracted by the feature extraction shared module) and the corresponding label y (m) . At any point in time, when the image recognition learning machine receives a batch of labeled samples from task m, they may be from a new task or a task that has been learned. After receiving the training data, the goal of the learning model is to learn the learning model of each task through the given training samples as close as possible to the true target function f (m) , and the evaluation of the learning performance is determined by calculating the error on all test samples. The target function of the learning model is:

[0069]

[0070] Among them, σ(·) is the activation function; represents the feature representation of the i-th input of the m-th task at the l-th layer, is the label of the i-th sample in the m-th task, n m represents the number of training samples of task m, W l represents the weight connection matrix of the l-th layer, is the knowledge base matrix, s (m) is the linear combination coefficient; the parameters θ of each task model (m) are linearly combined by some column vectors in the L matrix, that is, θ (m) = Ls (m) ; both μ and λ are balance parameters.

[0071] S21. Knowledge Warehouse Construction

[0072] It is mainly used to store the knowledge that the learning machine has learned and is necessary to save for later use. For the storage and representation methods of knowledge, there are the following two approaches. One is to directly store the learned data samples. Although this approach retains the most complete information, it will cause too high a space complexity and it is also difficult to ensure the efficiency of storage and reading during the learning process. Another approach is to save the feature representation of the knowledge, which is equivalent to performing feature extraction and compression on the original information.

[0073] S22. Knowledge Transfer

[0074] It refers to selecting the knowledge useful for the learning of the new model from the knowledge warehouse for transfer to help the learning of the new model, which is the theoretical basis of the learning machine. The present invention constructs a knowledge transfer model using a lifelong learning-based method. To perform knowledge transfer and avoid catastrophic forgetting, for each image recognition task, after training the task, calculate the importance Ω ij of each parameter in the network for this task, that is, the proportion of the parameter value in the i-th row and j-th column to the entire parameter value, and apply it to the training of subsequent tasks. Ω ij is added to the loss function in the form of a regularization term. Whenever a new task is trained: for the parameter with a larger Ω ij , try to reduce the change amplitude of it during gradient descent because this parameter is important for a certain past task and its value needs to be retained to avoid catastrophic forgetting; while for the parameter with a smaller Ω ij , its gradient can be updated with a larger amplitude to obtain better performance on the new task. Therefore, the loss function of the m-th task is:

[0075]

[0076] Among them, L (m) (θ) is the loss function of the current task, λ is the balance factor, and θ ij is the parameter at the i-th row and j-th column in the current model parameter matrix. is the model parameter obtained after training by the previous m - 1 tasks. Whenever a task is completed, Ω ij will be updated. When performing the first task, take 0.

[0077] S23. Model learning. In the image recognition continual learning method, it is responsible for the rapid learning of new tasks and interacts with the knowledge repository through two processes: knowledge transfer and knowledge update, effectively utilizing past knowledge to improve the speed of learning new tasks, and at the same time integrating the new knowledge learned in the new task into the original knowledge. This process is independent of other parts of the system. Additionally, not all knowledge in the knowledge repository will be used in the process of learning a new model.

[0078] S24. Knowledge update

[0079] Knowledge update is a crucial link in the image recognition continual learning method. It is the process of converting working memory into long - term memory, ensuring that the knowledge repository can be continuously updated so that this knowledge can be effectively transferred when learning new tasks. During the update process, the system should screen the knowledge accordingly, integrate as much correct knowledge as possible without integrating useless knowledge, and at the same time ensure that there is no loss of the original knowledge after integration.

[0080] S25. Receiving feedback

[0081] For an image recognition continual learning method, in addition to interacting with the environment, relevant supplementary information can be obtained during the interaction with people or other machines. On the one hand, this information increases the expert knowledge of the system; on the other hand, the system can identify its own errors from the comparison with the supplementary information and correct these errors. This process may greatly improve the system performance. And this process can be added to each of the above - mentioned components.

[0082] The present invention also provides an image continual learning recognition device, including a feature extraction sharing module and a knowledge transfer learning module. The feature extraction sharing module includes the following units:

[0083] An unsupervised training unit, used for performing unsupervised pre - training, and training a feature extraction sharing model by using the unsupervised reconstruction pre - training method in an auto - encoding network;

[0084] An update unit, which is used to maintain an updated feature extraction shared model, and uses the BP backpropagation algorithm to finely tune and update the feature extraction shared model generated in step S11;

[0085] A feature extraction unit, which is used to use the updated feature extraction shared model to perform layer-by-layer feature extraction on the image X to be recognized (t) to obtain image features H (t) , where t represents the t-th recognition task;

[0086] A model parameter solving unit, which is used to decompose and solve the model parameters θ (t) for the features H of the current task (t) and the knowledge base matrix L;

[0087] Among them, in the knowledge transfer learning module, the features H extracted by using the feature extraction shared module (t) are used to construct a new task Z (t) =(H (t) , y (t) ), where y (t) is the label. Subsequently, transfer learning is performed on this set of training data using the lifelong machine learning algorithm to obtain the final image recognition learning model f (t) (θ) of each task, and the learning model is used for image recognition.

[0088] The present invention also provides an electronic device, including a distributed memory, a processor, and a computer program that can run in the processor in the memory. When the processor executes the computer program, the steps of the above-described image continuous learning and recognition method are implemented.

[0089] The present invention also provides a computer-readable storage medium, storing computer software instructions, and the computer software instructions implement the steps of the above-described image continuous learning and recognition method.

[0090] The following uses a specific experiment to illustrate the process and effect of image implementation of the present invention. The datasets used in the experiment are as follows:

[0091] ImageNet dataset:

[0092] ImageNet is a large visual database for visual object recognition software research. More than 14 million image URLs are manually annotated by ImageNet to indicate the objects in the pictures; bounding boxes are also provided for at least one million images. ImageNet contains more than 20,000 categories; a typical category, such as "balloon" or "strawberry", contains hundreds of images.

[0093] MNIST handwritten digit recognition dataset:

[0094] The MNIST dataset is a very comprehensive handwritten digit dataset, containing 60,000 training samples and 10,000 test samples. The MNIST handwritten digit database contains data of 10 classes from character 0 - 9. We randomly divide a total of 2,000 data into 10 groups, and each group is regarded as a regression learning task.

[0095] Animal Image Classification (Cat and Dog) Dataset:

[0096] The Animal Image Database (Animals Dataset) contains 30,475 images of 50 classes in 6 views, and each image has only one class information. We select 4 types of dogs (Chihuahua, Collie, Dalmatian, and German Shepherd) and 2 types of cats (Bobcat and Persian cat) as positive classes, and randomly select the remaining animal classes as parent class samples to form a total of 6 binary classification multi - task databases.

[0097] CIFAR10 Dataset:

[0098] The CIFAR10 dataset is a relatively comprehensive object image dataset, with a total of 50,000 training samples and 10,000 test samples, which are divided into 10 classes: ship, airplane, dog, truck, deer, horse, cat, bird, car, frog.

[0099] In the first step, use the ImageNet dataset to pre - train the feature extraction shared model to obtain the network model parameters W l , b l .

[0100] In the second step, first use the MNIST handwritten digit recognition dataset to fine - tune the feature extraction shared model obtained in the first step to obtain the updated network model parameters W l , b l . Then, use the fine - tuned feature extraction shared model to extract features from handwritten digit images to obtain the image features H (t) ; at the same time, decompose and solve the top - layer parameter sharing model θ (t) and the knowledge base matrix L for the image features H (t) . Finally, use the image classification function f (t) (θ) to recognize the image features H (t) . Recognize 10,000 test samples, and the recognition rate is 94.35%.

[0101] In the third step, first use the Animal Image Classification (Cat and Dog) dataset to fine - tune the feature extraction shared model to obtain the updated network model parameters W l , b l . Then, use the fine - tuned feature extraction shared model to extract features from animal images to obtain the image features H(t) . Secondly, decompose and solve the image feature H (t) to update the top-level parameter sharing model θ (t) and the knowledge base matrix L. Finally, according to the input image and the knowledge base matrix L, use the method of lifelong learning for knowledge transfer, fuse the transferred knowledge and the image feature H (t) and use the image classification function f (t) (θ) for animal recognition. Identify 5000 test samples, and the recognition rate is 90.15%; and use the updated model to identify 10000 MNIST test samples in turn, and the recognition rate is 92.17%.

[0102] In the fourth step, first fine-tune the feature extraction sharing model using the CIFAR10 dataset to obtain the updated network model parameters W l , b l . Then, use the fine-tuned feature extraction sharing model to extract features from the CIFAR10 dataset images to obtain the image feature H (t) . Secondly, decompose and solve the image feature H (t) to update the top-level parameter sharing model θ (t) and the knowledge base matrix L. Finally, according to the input image and the knowledge base matrix L, use the method of lifelong learning for knowledge transfer, fuse the transferred knowledge and the image feature H (t) and use the image classification function f (t) (θ) for image recognition. Identify 10000 test samples, and the recognition rate is 95.46%; and use the updated model to identify 10000 MNIST test samples and 5000 animal image classification test samples in turn, and the recognition rates are 91.78% and 89.36% respectively.

[0103] This continuous learning method for image recognition conducts continuous learning 3 times, and the image recognition results after each learning are shown in the following table. It can be seen from the table that this continuous learning method for image recognition can continuously learn new image recognition tasks without much affecting the recognition performance of the original tasks. In addition, when learning new tasks, it can also make full use of the knowledge learned from existing tasks for knowledge transfer learning.

[0104] Handwritten digit recognition Animal (cat and dog) recognition CIFAR10 image recognition The first-generation model 94.35% - - The second-generation model 92.17% 90.15% - The third-generation model 91.78% 89.36% 95.46% ...... ...... ...... ......

[0105] The specific embodiments described in this article are only examples to illustrate the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. An image continuous learning and recognition method, characterized in that, It includes two major steps: feature extraction sharing and knowledge transfer learning. Among them, feature extraction sharing includes the following sub-steps: S11, unsupervised pre-training, training a feature extraction sharing model using the unsupervised reconstruction pre-training method in an autoencoder network; S12, maintaining and updating the feature extraction sharing model, using the BP backpropagation algorithm to fine-tune and update the feature extraction sharing model generated in step S11; In step S12, after the pre-training is completed, the image recognition task is continuously performed. Set the currently upcoming task to be recognized as task t, and the representation of task t is Z (t) =(X (t) , y (t) ), where X (t) is the image sample in task t, and y (t) is its corresponding label; use these data to fine-tune and update the feature extraction shared model generated in step S11 to avoid affecting the representativeness of features due to distribution drift, and use the BP backpropagation algorithm to update the feature model; First, for the input X (t) Calculate the features of each layer once from bottom to top to obtain H l , l = 1, 2, 3, …, n L , n L represents the maximum number of layers. Use a single-task learner to learn (H l , y (t) ) to obtain the parameter W0, and apply the BP algorithm to optimize the network parameters; for each layer l, calculate ∪ l = H l * (1 - H l ) ∪ l+1 W l+1 T . If it is the last layer, then ∪ l = H l * (1 - H l ) y (t) . Through W l ← W l - α t H l-1 T ∪ l update the matrix W l ; Considering avoiding the occurrence of task negative transfer, the learning rate α of the model t varies with the task relevance, that is, for each task t, calculate its relevance γ with the previous task t : γ t = cos(θ (t) , θ (t-1) ) = cos(Ls (t) , Ls (t-1) ) = cos(s (t) , s (t-1) ) where \(L\in R\) d×k is the knowledge base matrix, and \(s\) (t) is the linear combination coefficient; the parameter \(\theta\) of each task model (t) is linearly combined by some column vectors in the \(L\) matrix, that is, \(\theta\) (t) = \(Ls\) (t) , \(\cos()\) represents the cosine angle between two vectors. The smaller the angle between the division hyperplanes of the task models, the higher the task correlation; further, the task learning rate \(\alpha\) t is determined by the following formula: α t = α c (γ t + 1) / 2 where α c is the baseline learning rate for BP tuning, which is a hyperparameter that needs to be tuned manually; S13. Use the updated feature extraction sharing model to perform layer-by-layer feature extraction on the image X to be recognized to obtain the image feature H. (t) (t) , where t represents the t-th recognition task.​ S14, decompose the feature H of the current task (t) to solve the model parameter θ (t) and the knowledge base matrix L; Knowledge transfer learning includes knowledge repository construction, knowledge transfer, model learning, updating the knowledge repository, and receiving feedback; Among them, the feature H extracted by using the feature extraction sharing module in knowledge transfer (t) Construct a new task Z (t) =(H (t) , y (t) ), where y (t) is the label. Subsequently, transfer learning is performed on this set of training data using the lifelong machine learning algorithm to obtain the final image recognition learning model f (t) (θ) of each task, and the learning model is used for image recognition.

2. The image continuous learning and recognition method according to claim 1, characterized in that: In step S1, first randomly initialize the parameters W l , b l of the autoencoder network model, and use the unlabeled training dataset where T c is the number of task pools to be recognized, and use the contrastive divergence algorithm to optimize and update the parameters W l , b l one by one from bottom to top; the contrastive divergence algorithm first uses the Gibbs sampling method to obtain the hidden layer features H l-1,0 , H l,0 , H l-1,1 , H l,1 after three samplings, and then update the matrix W l ← W l + α(H l-1,0 T H l,0 - H l-1,1 T H l,1 ), where T represents the transpose of the matrix, α is the learning rate of network pre-training, loss(·) is the loss function, l , update b l , where T represents the transpose of the matrix, α is the learning rate of network pre-training, loss(·) is the loss function, represents the partial derivative.

3. The image continuous learning and recognition method according to claim 1, characterized in that: In step S14, for the input (H (t) , y (t) ), the model parameters θ (t) and the Hessian matrix D (t) are obtained by closed-form through the following formula: θ (t) =(H (t) H (t)T ) -1 H (t) y (t) Subsequently, from optimize s (t) , where μ is a balance factor, calculated using the stochastic gradient descent method Then, through update the shared task base L, that is, the knowledge base matrix, T represents the maximum number of tasks, and n t represents the number of samples corresponding to the t-th task.

4. An image continuous learning and recognition method according to claim 1, characterized in that: Assume there are M image recognition tasks T (1) , T (2) , … T (M) and a large amount of labeled data that conforms to the distribution of the learned tasks. Each task T (m) = (f (m) (θ), H (m) , y (m) ), where H (m) is the feature of the training samples of task m, y (m) is the corresponding sample label, f (m) (θ) is the recognition function of task m, and θ is the network parameter of the learning model; that is, the label of each task is determined by a true implicit function f (m) (θ), which is continuously updated and optimized as the recognition tasks increase; each task is given n m training samples and the corresponding label y (m) , and d is the dimension of the sample features extracted by the feature extraction shared module; at any point in time, when the image recognition learning model receives a batch of labeled samples from task m, they may be from a new task or a task that has been learned. After receiving the training data, the goal of the learning model is to learn the learning model of each task through the given training samples as close as possible to the true objective function f (m) ; the objective function of the learning model is: Among them, σ(·) is the activation function; represents the feature representation of the i-th input of the m-th task at the l-th layer, is the label of the i-th sample in the m-th task, n m represents the number of training samples of task m, W l represents the weight connection matrix of the l-th layer, L ∈ R d×k is the knowledge base matrix, s (m) is the linear combination coefficient; the parameter θ of each task model (m) is linearly combined by some column vectors in the L matrix, that is, θ (m) = Ls (m) ; both μ and λ are balance parameters.

5. The image continuous learning and recognition method according to claim 1, wherein: For knowledge transfer and to avoid catastrophic forgetting, for each image recognition task, after training on that task is completed, the importance Ω of each parameter in the network for that task is calculated ij , that is, the proportion of the parameter value in the i-th row and j-th column to the entire parameter value, and is carried over to subsequent tasks during training. Ω ij is added to the loss function in the form of a regularization term. Whenever a new task is being trained: for Ω ij parameters with larger values, minimize the magnitude of their change during gradient descent because these parameters are important for a past task and their values need to be preserved to avoid catastrophic forgetting; while for Ω ij parameters with smaller values, update their gradients with a larger magnitude to achieve better performance on the new task; Therefore, the loss function for the m-th task is: Among them, L ( m ) (θ) is the loss function of the current task, λ is the balance factor, and θ ij is the parameter in the i-th row and j-th column of the current model parameter matrix. is the model parameter obtained after training with the previous n - 1 tasks. Whenever a task is completed, Ω ij will be updated. When performing the first task, take 0.

6. An image continuous learning and recognition device, characterized in that, It includes a feature extraction sharing module and a knowledge transfer learning module. Among them, the feature extraction sharing module includes the following units: An unsupervised training unit for performing unsupervised pre-training and training a feature extraction sharing model using the unsupervised reconstruction pre-training method in an autoencoder network; An update unit for maintaining and updating the feature extraction sharing model and using the BP backpropagation algorithm to fine-tune and update the feature extraction sharing model generated in step S11; After the pre-training is completed, start continuously performing image recognition tasks. Set the current task to be recognized as task t, and the representation of task t is Z (t) =(X (t) , y (t) ), where X (t) is the image sample in task t, and y (t) is its corresponding label; use this data to fine-tune and update the feature extraction shared model generated in step S11 to avoid affecting the representativeness of features due to distribution drift, and use the BP backpropagation algorithm to update the feature model; First input X (t) Calculate each layer feature from bottom to top once and get H l ,l=1,2,3,…,n L , n L Represents the maximum number of layers, using a single-task learner pair (H l ,y (t) ) is used to learn the parameters W0, and the BP algorithm is used to optimize the network parameters; for each layer l, ∪ l =H l *(1-H l )∪ l+1 W l+1 T , if it is the last layer, then ∪ l =H l *(1-H l )y (t) , through W l ←W l -α t H l-1 T ∪ l Update the matrix W l ; Considering avoiding the occurrence of task negative transfer, the learning rate α of the model t varies with the task relevance, that is, for each task t, calculate its relevance γ with the previous task t : γ t = cos(θ (t) , θ (t-1) ) = cos(Ls (t) , Ls (t-1) ) = cos(s (t) , s (t-1) ) where, L ∈ R d×k is the knowledge base matrix, and s (t) is the linear combination coefficient; the parameter θ (t) of each task model is linearly combined by some column vectors in the L matrix, that is, θ (t) = Ls (t) , cos() represents the cosine angle between two vectors, and the smaller the angle between the partition hyperplanes of the task models, the higher the task correlation; further, the task learning rate α t is determined by the following formula: α t = α c (γ t + 1) / 2 where α c is the baseline learning rate for BP tuning, which is a hyperparameter that needs to be tuned manually; A feature extraction unit for using the updated feature extraction shared model to perform layer-by-layer feature extraction on the image X to be recognized to obtain the image feature H (t) where t represents the t-th recognition task; (t) ​ The model parameter solving unit is used to decompose and solve the model parameter θ (t) for the feature H of the current task (t) and the knowledge base matrix L; Among them, the feature H extracted by the feature extraction sharing module is used in the knowledge transfer learning module (t) Construct a new task Z (t) =(H (t) , y (t) ), where y (t) is the label. Subsequently, transfer learning is performed on this set of training data using the lifelong machine learning algorithm to obtain the final image recognition learning model f (t) (θ) of each task, and the learning model is used for image recognition.

7. An electronic device, comprising a distributed memory, a processor, and a computer program that can run in the processor in the memory, characterized in that: When the processor executes the computer program, it implements the steps of an image continuous learning and recognition method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing computer software instructions, characterized in that: The computer software instructions implement the steps of an image continuous learning and recognition method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method applied to small sample picture classification based on transfer learning and attention mechanism element learning

    CN114492581A

  • Electronic apparatus and method for generating trained model

    US20180357538A1