Iterative enhanced dynamic expansion continuous learning model

Through iteratively enhanced dynamically extended continuous learning model, using parallel knowledge integration and knowledge distillation technology, the problem of forgetting problems and high computing resource requirements in the existing technology is solved, and more efficient model learning and performance improvement is achieved.

CN120123765APending Publication Date: 2025-06-10RES & DEV INST OF NORTHWESTERN POLYTECHNICAL UNIV IN SHENZHEN +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510179276.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

There is a problem of forgetting in the existing continuous learning methods based on complementary learning system theory, and the cost of frequent full-parameter fine-tuning of the model is too high.

Method used

It adopts an iteratively enhanced dynamic extended continuous learning model, replaces one-way knowledge consolidation through parallel knowledge integration, and uses knowledge distillation to transfer the old and new knowledge in the plastic model to the stable model, reducing the demand for computing resources.

Benefits of technology

It effectively solves the problem of catastrophic forgetting, reduces the model's demand for computing resources, and improves the reverse migration ability of model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123765A_ABST
    Figure CN120123765A_ABST
Patent Text Reader

Abstract

The invention relates to an iterative enhanced dynamic expansion continuous learning model. A training method comprises the following steps: constructing a task sequence # imgabs0 # containing T target tasks as training data; training the pre-trained VisionTransform model by using the training data, embedding a plurality of fine tuning modules into the pre-trained VisionTransform model by using a gating network, and configuring a weight for each fine tuning module; during training, a module # imgabs1 # used for learning a task Dt is configured to obtain a plastic model # imgabs2 #, in the training process, a VisionTransform model is frozen for feature extraction, and a training module # imgabs3 # is used for learning knowledge of the task Dt; after training is completed, sampling a fixed number of sample sets from training data, and adding the sample sets into the maintained sample buffer area; and constructing a stable model # imgabs4 #, and when t is not equal to 1, migrating new knowledge of a module # imgabs6 # and old knowledge of a module # imgabs7 # in the plastic model # imgabs5 # to a module # imgabs9 # of a stable model # imgabs8 # through knowledge distillation. The problem of forgetting in an existing continuous learning method can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to an iteratively enhanced dynamic expansion continuous learning model, a computer program product, and an electronic device. Background Art

[0002] In the related art, with the development of science and technology, models pre-trained in the general field cannot meet the needs of users in specific fields, and the continuous update of knowledge makes the cost of frequently fine-tuning all parameters of the model too high and unrealistic. Continuous learning refers to the process in which a model sequentially learns tasks coming one by one in a task sequence and continuously updates parameters over time to learn new knowledge, which can well solve the need to process frequently updated data streams in actual production and life. In continuous learning, the model cannot revisit previous tasks, which requires the model to retain the knowledge learned previously when learning new tasks to avoid the occurrence of "catastrophic forgetting" during the learning process. Inspired by the theory of Complementary Learning Systems (CLS), existing related continuous learning methods usually design a hippocampus-like network and a neocortex-like network to simulate the hippocampus and neocortex in the human brain respectively, learn short-term knowledge and store long-term memories respectively, and alleviate the forgetting problem by consolidating the short-term model of the former into the latter as long-term memories. Most of the existing methods based on the CLS theory have a one-way consolidation process, that is, from the hippocampus-like network to the neocortex-like network. However, the episodic memory of the hippocampus-like network may cover the structured knowledge of the neocortex network during the transfer process, and due to the lack of an effective mechanism for recalling and consolidating old knowledge, it may lead to a decline in its performance on past tasks.

[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0004] The present invention provides an iteratively enhanced dynamic expansion continuous learning model method, a computer program product, and an electronic device, which can effectively solve the catastrophic forgetting existing in the dual-memory model and reduce the demand of the model for computing resources, and thus can overcome the defects existing in the prior art to a certain extent.

[0005] Other features and advantages of the present invention will become apparent through the following detailed description, or will be partially learned through the practice of the present invention.

[0006] According to a first aspect of the present invention, there is provided an iteratively enhanced dynamic expansion continuous learning model, and the training method of the model includes:

[0007] Construct a task sequence containing T target tasks as training data;

[0008] Use the training data to train the pre-trained Vision Transformer model F θ and embed multiple fine-tuning modules into the pre-trained Vision Transformer model using a gating network, and configure weights for each fine-tuning module; among them, the set of the gating network and each fine-tuning module is denoted as module

[0009] During training, configure the module for learning task D t of and obtain a plastic model Among them, module is the same as module in structure, and the parameters are randomly initialized; during the training process, freeze the Vision Transformer model F θ for feature extraction, and train module to learn the knowledge of task D t When t = 1, freeze F θ and train module to learn the knowledge of task D t ;

[0010] After completing the training, sample a fixed number of sample sets from the training data and add them to the maintained sample buffer;

[0011] Construct a stable model When t ≠ 1, transfer the new knowledge of module in the plastic model and the old knowledge of module to the module of the stable model through knowledge distillation. in.

[0012] In some exemplary embodiments, the transfer of the new knowledge of module in the plastic model and the old knowledge of module to the module of the stable model through knowledge distillation includes: Use the old task samples in the sample buffer to train the module

[0013] of the plastic model to obtain the predicted first probability distribution;

[0014] Use the new task samples in the training data to train module ​Train to obtain the predicted second probability distribution;

[0015] Use the loss function to calculate the distance between probability distributions to achieve knowledge transfer.

[0016] In some exemplary embodiments, the loss function of knowledge distillation includes: response-based distillation loss and feature-based distillation loss

[0017] Among them, response-based distillation is used to take the distribution output by the stable and plastic model as the learning target of the stable model;

[0018] Feature-based distillation is used to align two features output by the same depth layer of the stable model and the plastic model.

[0019] In some exemplary embodiments, feature-based distillation includes: the first part of the loss, the second part of the loss;

[0020] The first part of the loss is used to measure the similarity between two features in the data and align the features in the batch dimension;

[0021] The second part of the loss is used to measure the distance between two features at the internal level of the data and align the features in the patch dimension.

[0022] In some exemplary embodiments, the fine-tuning module includes at least one of LoRA, Prefix-Tuning, and Adapter.

[0023] In some exemplary embodiments, the method further includes:

[0024] Use the LoRA fine-tuning module to design bypasses for each block's projection matrices W θ of the VisionTransformer model F q and W k respectively for learning new tasks. Each bypass uses the lower projection matrix W Down and the upper projection matrix W Up to project the input h of the block twice, and add the projection results to the outputs of the original projection matrices W q and W k ;

[0025] Use a learnable gate to scale the output of the LoRA fine-tuning module.

[0026] In some exemplary embodiments, the method further includes:

[0027] Use the Adapter fine-tuning module to fine-tune the VisionTransformer model F θ For each Transformer layer of, embed the Adapter module after the multi-head attention module and after the feed-forward layer respectively, to keep the parameters of the original pre-trained model fixed during training, and only fine-tune the newly added Adapter module and Layer Norm layer.

[0028] In some exemplary embodiments, set two sets of soft prompts with length l and dimension d in the Prefix-Tuning fine-tuning module Concatenate them to the matrix K and matrix V of the multi-head attention in the i-th block respectively; and freeze the matrix K and matrix V during training, and train the continuous soft prompts

[0029] According to a second aspect of the present invention, there is provided a computer program product having a computer program stored thereon, and when the computer program is executed by a processor, the iterative enhanced dynamic extended continuous learning model described above is implemented.

[0030] According to a third aspect of the present invention, there is provided an electronic device, including:

[0031] A processor; and

[0032] A memory for storing executable instructions of the processor;

[0033] Wherein, the processor is configured to implement the iterative enhanced dynamic extended continuous learning model described above when executing the executable instructions.

[0034] The iterative enhanced dynamic extended continuous learning model provided by the embodiments of the present invention, during training, replaces naive one-way knowledge consolidation with parallel knowledge integration, rather than simply integrating short-term memory into the neocortical network as long-term memory. While migrating new task knowledge, it consolidates old task knowledge, which can effectively solve the forgetting problem in existing continuous learning methods based on the complementary learning system theory.

[0035] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Description of the Drawings

[0036] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0037] Figure 1 A schematic diagram schematically showing a model training method of an iteratively enhanced dynamic expansion continuous learning model according to an exemplary embodiment of the present invention;

[0038] Figure 2 A schematic diagram schematically showing an iterative enhancement dynamic expansion continuous learning model training process according to an exemplary embodiment of the present invention;

[0039] Figure 3 A schematic diagram schematically showing a UniPELT module according to an exemplary embodiment of the present invention;

[0040] Figure 4 A schematic diagram schematically showing a knowledge transfer and integration according to an exemplary embodiment of the present invention;

[0041] Figure 5 A schematic diagram schematically showing the composition of an electronic device in an exemplary embodiment of the present invention. Detailed implementation manners

[0042] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this invention will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments.

[0043] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0044] In view of the disadvantages and deficiencies of the prior art, an iteratively enhanced dynamic expansion continuous learning model is provided in this example embodiment. Referring to Figure 1 as shown, the training method of the model can specifically include the following steps:

[0045] Step S11, construct a task sequence containing T target tasks as training data;

[0046] Step S12, use the training data to train the pre-trained Vision Transformer model F θ and embed multiple fine-tuning modules into the pre-trained Vision Transformer model by using a gating network, and configure weights for each fine-tuning module; wherein, the set of the gating network and each fine-tuning module is denoted as module

[0047] Step S13, during training, configure the module for learning task D t of and obtain a plastic model wherein, module and module have the same structure, and the parameters are randomly initialized; during the training process, freeze the Vision Transformer model F θ for feature extraction, and train module for learning the knowledge of task D t When t = 1, freeze F θ and train module to learn the knowledge of task D t ;

[0048] Step S14, after completing the training, sample a fixed number of sample sets from the training data and add them to the maintained sample buffer;

[0049] Step S15, construct a stable model When t≠1, transfer the new knowledge of module in the plastic model and the old knowledge of module to the module of the stable model by knowledge distillation.

[0050] Next, each step of the model training method corresponding to the iteratively enhanced dynamic expansion continuous learning model in this exemplary embodiment will be described in more detail in conjunction with the accompanying drawings and embodiments.

[0051] In step S11, construct a task sequence containing T target tasks as training data.

[0052] Exemplarily, the user can create a training task of the model on the intelligent terminal device and select a certain type of data to construct the training data.

[0053] Specifically, a sequence containing T target tasks is given where the t-th classification task where is the classification task D t is the input sample of is the corresponding class label, n t is the task D t and the corresponding number of samples. Here, T is a positive integer.

[0054] For t ≠ t ′ , when learning the task D t , the continual learning requirement is that the model cannot revisit the data of previous tasks . Therefore, the model needs to retain the knowledge learned previously while adapting to new tasks.

[0055] In step S12, the pre-trained Vision Transformer model F θ is trained using the training data. A gating network is used to embed multiple fine-tuning modules into the pre-trained Vision Transformer model and configure weights for each fine-tuning module; among them, the set of the gating network and each fine-tuning module is denoted as module

[0056] Exemplarily, the fine-tuning module includes at least one of LoRA, Prefix-Tuning, and Adapter.

[0057] Specifically, when learning the t-th task D of the task sequence t , referring to Figure 3 shown, the gating network UniPELT is used to embed fine-tuning modules such as LoRA, Prefix-Tuning, and Adapter into the corresponding positions of the pre-trained Vision Transformer model F θ , and weights are assigned to different sub-modules through a learnable gating network. The set of each parameter-efficient fine-tuning sub-module and the gating network is denoted as module

[0058] For example, in the task D t , this continual learning model consists of a plastic model and a stable model to solve the plasticity problem and stability problem in the continual learning process respectively. and respectively represent the PEFT (Parameter-Efficient Fine-Tuning) backbone module and the PEFT extension module with the same structure, which are a set of fine-tuning modules composed of LoRA, Prefix-Tuning, and Adapter, etc., and are implemented through the integrated framework UniPELT with a gating mechanism, where m ∈ {P, S}. F θ is a shared pre-trained model with parameters θ. For example, the shared pre-trained model is a pre-trained Vision Transformer model.

[0059] Exemplarily, the method further includes: using the LoRA fine-tuning module to design bypasses for the projection matrices W θ of each block of the Vision Transformer model F q and W k respectively for learning new tasks. Each bypass uses the lower projection matrix W Down and the upper projection matrix W Up to project the input h of the block twice, and add the projection results to the outputs of the original projection matrices W q and W k ; using a learnable gate to scale the output of the LoRA fine-tuning module.

[0060] Specifically, in the LoRA (Low-Rank Adaptation) sub-module, for the projection matrices W θ of each block of F q and W k respectively design bypasses for learning new tasks. Each bypass uses the lower projection matrix W Down ∈R d×r and the upper projection matrix W Up ∈R r×d to project the input h of the block twice in sequence, and add the projection results to the outputs of the original W q and W k . Among them, r is the intermediate dimension, and the dimensions of the input and output remain d unchanged before and after processing.

[0061] In addition, the sub-module contains a learnable gate to scale the output of LoRA. σ:R d →R is a fully connected layer that maps the input dimension d to 1. Multiply the gate by the output of LoRA. When approaches 0, the output of LoRA will approach hW m , reducing the contribution of LoRA to the learning of the current task. The formula corresponding to the above process can be expressed as:

[0062]

[0063] Among them, h is the input of the LoRA sub-module; q and k respectively represent the query module and the key module of the attention mechanism.

[0064] Exemplarily, the method further includes: using the Adapter fine-tuning module to fine-tune the VisionTransformer model F θ For each Transformer layer of, an Adapter module is respectively embedded after the multi-head attention module and after the feed-forward layer, which is used to fix the parameters of the original pre-trained model unchanged during training, and only fine-tune the newly added Adapter module and the Layer Norm layer.

[0065] Specifically, each layer of Adapter contains two sub-modules, which are respectively embedded after the multi-head attention module and after the feed-forward layer. Each sub-module consists of an upper projection matrix W Up ∈R r×d , a lower projection matrix W Down ∈R d×r and also contains a non-linear activation layer φ and a gating mechanism First, the W Down matrix maps the input h to a low-dimensional space with dimension r. Subsequently, after passing through a non-linear activation layer φ, the W Up matrix projects it back to the original dimension, where f is the ReLU function. In order to measure the contribution of each Adapter to learning this task, multiply the output of each Adapter by to weight it, σ:R d →R is a fully connected layer. When this Adapter has no positive effect on learning the new task, the output degenerates to h ′ = h, and at this time it is equivalent to removing this Adapter from the model, and the formula is expressed as:

[0066]

[0067] Exemplarily, the method further includes: setting two groups of soft prompts with length l and dimension d in the Prefix-Tuning fine-tuning module Concatenate them on the matrix K and matrix V of the multi-head attention in the i-th block respectively; and freeze the matrix K and matrix V during training, and train the continuous soft prompts

[0068] Specifically, in the Prefix-Tuning (prefix tuning) sub-module, set two groups of soft prompts with length l and dimension d They are respectively concatenated onto the matrix K and matrix V of the multi-head attention in the i-th block. During training, the matrix K and matrix V are frozen, and only the continuous prompts need to be trained. Similar to that in the LoRA sub-module, learnable gates are used. Multiply by the soft prompt. To evaluate its importance, the formula is expressed as:

[0069]

[0070] Where, σ:R d →R is a fully connected layer.

[0071] When and make little contribution to learning the knowledge of this task, the value of the learnable gate will tend to 0, and the input of the multi-head attention will become the original form.

[0072] For example, the ViT-B / 16 model pre-trained on ImageNet21k can be selected as F θ ; in the PEFT module, set the rank of LoRA to 2, the prefix length to 5, and the downsampling size of the adapter to 1 / 4 of the original input size. Both α and β of the loss function are set to 0.5; set the sample buffer size to 100 samples per class, and adopt the iCaRL algorithm as the sample set construction algorithm.

[0073] In step S13, during training, configure the module t for learning task D and obtain the plastic model Where, the module has the same structure as the module , and the parameters are randomly initialized; during the training process, freeze the VisionTransformer model F θ for feature extraction, and train the module for learning the knowledge of task D t ; when t = 1, freeze F θ and train the module to learn the knowledge of task D t .

[0074] Exemplarily, referring to Figure 2 shown, for the dynamic expansion process of the model, during each learning of a new task (t≠1), dynamically expand a module with the same structure as but randomly initialized parameters, focusing on learning the new task, Thus, the dynamic expansion of the model is achieved. During the training process, the Vision Transformer F is frozen θ and used for feature extraction, and training is performed to enable it to learn the new task D t knowledge. When t = 1, F is frozen θ and training is carried out to enable it to learn the knowledge of task D t .

[0075] Specifically, the module accommodates the knowledge obtained from the previous task and is initialized from the previous task to learn the current task D t ; while is responsible for storing the integrated knowledge from . This process can be expressed by the formula in the t-th task as:

[0076]

[0077] where represents the initialized model parameters. In the next task D t+1 , is regarded as and then the above process is repeated starting from step S13 until the learning on the entire task sequence is completed.

[0078] Specifically, the module focuses on learning the new task where the input image category label n t is the number of samples contained in task D t . Since there is no need to consider the preservation of old knowledge, the obtained model parameters are optimal at this time. The module is initialized from in the previous task. Given the input t ′ ∈[1, t - 1], the module can give the corresponding prediction result In the training stage, the pre-trained model F is frozen θ , used as feature extraction, while training each PEFT module to enable it to learn the task knowledge.

[0079] In step S14, after the training is completed, a fixed number of sample sets are sampled from the training data and added to the maintained sample buffer.

[0080] Exemplarily, for the iterative enhancement strategy, after completing one round of training, sample a set of samples with a quantity of n ′ <n t from the training data and add it to the maintained sample buffer.

[0081] In step S15, construct a stable model When t≠1, through knowledge distillation, transfer the new knowledge of the modules in the malleable model and the old knowledge of the modules to the modules of the stable model simultaneously in .

[0082] Exemplarily, the transfer of the new knowledge of the modules in the malleable model and the old knowledge of the modules to the modules of the stable model through knowledge distillation includes: Step S151, use the old task samples in the sample buffer to train the modules of the malleable model

[0083] to obtain the predicted first probability distribution;

[0084] Step S152, use the new task samples in the training data to train the modules to obtain the predicted second probability distribution;

[0085]

[0086] Step S153, use the loss function to calculate the distance between the probability distributions to achieve knowledge transfer. Specifically, for the transfer and integration of knowledge, when t≠1, construct a stable model and transfer the new knowledge of the modules in the malleable model and the old knowledge of the modules to the modules of the stable model simultaneously in

[0087] . The old task samples in the sample buffer are fed into to obtain the predicted probability distribution. Similarly, the new task samples are sent into

[0088] to obtain the predicted result. The formula can be expressed as: ​​

[0089] Finally, calculate through the designed loss function of the prediction and the distance between them, so as to achieve knowledge transfer.

[0090] Exemplarily, as shown in reference Figure 4 the loss function of knowledge distillation includes: response-based distillation loss and feature-based distillation loss

[0091] Among them, response-based distillation is used to take the distribution output by the stable and plastic model as the learning target of the stable model;

[0092] Feature-based distillation is used to align the two features output by the stable model and the plastic model at the same depth layer.

[0093] Specifically, the loss function contains two parts: response-based distillation loss and feature-based distillation loss The formula includes:

[0094]

[0095] Among them, β represents the weight of the loss functions of different distillations.

[0096] Specifically, response-based distillation takes the distribution output by the stable and plastic model as the learning target of the stable model, and the formula can include:

[0097]

[0098] Among them, is calculated through KL divergence, is achieved through the cross-entropy loss function, and α is a preset hyperparameter.

[0099] Specifically,

[0100]

[0101] Among them, c j is the true distribution, and represent the outputs of the plastic model and the stable model respectively; α is a preset hyperparameter representing the weight of different loss functions; T is the temperature of distillation, aiming to adjust the smoothness of the output.

[0102] Exemplarily, feature-based distillation Including: the first part of the loss, the second part of the loss;

[0103] The first part of the loss is used to measure the similarity between two features in the data and align the features in the batch dimension;

[0104] The second part of the loss is used to measure the distance between two features at the internal level of the data and align the features in the patch dimension.

[0105] Specifically, let the feature vectors output by the i-th block in the stable model and the plastic model be

[0106] Feature-based distillation Used to align the features output by the same depth layers of the stable model and the plastic model. The loss function consists of and Two parts, namely the first part of the loss and the second part of the loss. Is used to measure the similarity of features between data, while Is used to measure the distance between two features at the internal level of the data. They are used to align the batch dimension and the patch dimension of the two features respectively. The formula can be expressed as:

[0107]

[0108] Among them,

[0109]

[0110] Among them, Represents the correlation matrix, F T Is the transpose of the vector F.

[0111] Exemplarily, when t≠1, the stable model in task D t Is regarded as The plastic model of the next task D t+1 In the next task D

[0112] In the next task D t+1 In, Is regarded as Then start repeating the above process from step S13 until the learning on the entire task sequence Is completed.

[0113] Exemplarily, the above continuous learning model can be a VisionTransformer model for image classification, and a task sequence containing T image classification tasks can be constructed As training data; Given a sequence containing T image classification tasks where the t-th classification task where is the classification task D t is the input image of is the corresponding class label, n t is the number of samples corresponding to task D t Using this image training data, the pre-trained Vision Transformer model is trained using the above method, and the above process is repeated until learning on the entire task sequence is completed, and a trained image classification model is obtained.

[0114] In the method provided by the embodiments of the present invention, the optimal parameters for the new task are obtained through the dynamic expansion module, and the recollection and retention of past knowledge are achieved through the proposed iterative enhancement strategy. The present invention adopts a parameter-efficient fine-tuning mechanism during training, effectively reducing the demand for computing resources. It can effectively promote the reverse transfer of model performance and alleviate the catastrophic forgetting problem in continuous learning.

[0115] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0116] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0117] Figure 5 shows a schematic diagram of an electronic device suitable for implementing the embodiments of the present invention.

[0118] It should be noted that Figure 5 the electronic device 1000 shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.

[0119] Such as Figure 5As shown, the electronic device 1000 includes a Central Processing Unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the Read-Only Memory (ROM) 1002 or the program loaded from the storage section 1008 into the Random Access Memory (RAM) 1003. In the RAM 1003, various programs and data required for system operation are also stored. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004.

[0120] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage section 1008 as needed.

[0121] Specifically, according to an embodiment of the present invention, the process described below with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a storage medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by the Central Processing Unit (CPU) 1001, various functions defined in the system of the present application are executed.

[0122] It should be noted that the storage medium shown in the embodiments of the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any storage medium other than a computer-readable storage medium, and this storage medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the storage medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0124] The units involved in the embodiments of the present invention can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the units themselves in some cases.

[0125] It should be noted that, on the other hand, the present application also provides a storage medium, which can be included in an electronic device; or can exist alone without being assembled into the electronic device. The above storage medium stores one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device is caused to implement the methods described in the following embodiments. For example, the electronic device can implement the Figure 1 steps of the method as shown.

[0126] In one embodiment, the present application provides a computer program product, including a computer program, which when executed by a processor implements the steps in the above method embodiments.

[0127] In addition, the above drawings are only schematic illustrations of the processes included in the methods according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0128] Those skilled in the art will readily think of other embodiments of the present invention after considering the specification and practicing the invention herein. The present application is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include the common general knowledge or conventional technical means in the technical field not disclosed in the present invention. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the claims.

[0129] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only defined by the appended claims.

Claims

1. An iteratively enhanced dynamically extended continuous learning model, characterized in that: The model training methods include: Construct a task sequence containing T target tasks As training data; Use the training data to train the pre-trained VisionTransformer model F θ Training is performed by embedding multiple fine-tuning modules into the pre-trained VisionTransformer model using a gating network, and configuring weights for each fine-tuning module; the set of gating networks and each fine-tuning module is recorded as a module During training, the configuration is used to learn task D t Modules And get a plastic model Among them, the module With module The structure is the same, and the parameters are randomly initialized; during the training process, the VisionTransformer model F is frozen θ Used for feature extraction, training module For learning task D t knowledge; when t = 1, freeze F θ And train the module Learning Task D t knowledge; After completing the training, a fixed number of sample sets are sampled from the training data and added to the maintained sample buffer; Building a stable model When t≠1, the plastic model is transformed into Medium Module New knowledge and modules The old knowledge is transferred to the stable model Modules middle.

2. The method according to claim 1, characterized in that The plastic model is transformed into Medium Module New knowledge and modules The old knowledge is transferred to the stable model Modules Including: Module for using old task samples in the sample buffer to improve plasticity Conduct training to obtain the predicted first probability distribution; Use new task samples in the training data to train the module Perform training to obtain a predicted second probability distribution; The loss function is used to calculate the distance between probability distributions to achieve knowledge transfer.

3. The method according to claim 1 or 2, characterized in that: The loss functions of knowledge distillation include: response-based distillation loss and feature-based distillation loss Among them, response-based distillation Used to stabilize the distribution of plastic model outputs as the learning target of the stable model; Feature-based distillation Used to align the two features output by the same depth layer of the stable model and the plastic model.

4. The method according to claim 3, characterized in that: Feature-based distillation Includes: the first part of the loss, the second part of the loss; The first part of the loss is used to measure the similarity between two features in the data and align the features in the batch dimension; The second part of the loss is used to measure the distance between two features at the internal level of the data and align the features in the patch dimension.

5. The method according to claim 1, characterized in that The fine-tuning module includes: at least one of LoRA, Prefix-Tuning and Adapter.

6. The method according to claim 5, characterized in that The method further comprises: Use LoRA fine-tuning module to VisionTransformer model F θ The projection matrix W of each block q and W k Design bypasses for learning new tasks, and each bypass uses the projection matrix W Down and the projection matrix W Up Project the block input h twice and add the projection result to the original projection matrix W q and W k The outputs of are added; Utilizing learnable gates Scale the output of the LoRA fine-tuning module.

7. The method according to claim 5, characterized in that The method further comprises: Use the Adapter fine-tuning module to adjust the VisionTransformer model F θ For each Transformer layer, an Adapter module is embedded after the multi-head attention module and the feed-forward layer, respectively, to fix the parameters of the original pre-trained model during training and only fine-tune the newly added Adapter module and Layer Norm layer.

8. The method according to claim 5, characterized in that The method further comprises: In the Prefix-Tuning fine-tuning module, set two sets of soft prompts with length l and dimension d. Splice it to the matrix K and matrix V of the multi-head attention in the i-th block respectively; freeze the matrix K and matrix V during training and train continuous soft prompts 9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the iteratively enhanced dynamically expanded continuous learning model according to any one of claims 1 to 8 is implemented.

10. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the iteratively enhanced dynamically expanded continuous learning model of any one of claims 1 to 8 by executing the executable instructions.