Device and method for continuous learning of neural network
The neural network continuous learning device addresses the challenge of catastrophic forgetting by using a hyper network with task layer attention blocks to generate transformed parameters, enhancing learning efficiency and maintaining previous knowledge.
Patent Information
- Application Number
- PCT/KR2024/006190
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-05-08
- Publication Date
- 2025-05-22
AI Technical Summary
Existing deep learning algorithms face challenges in continuous learning, particularly in maintaining knowledge of previous tasks while learning new tasks, due to the phenomenon of catastrophic forgetting.
A neural network continuous learning device and method that utilizes a hyper network with task layer attention blocks to generate a transformed parameter set for the base network, allowing for the integration of knowledge from previous tasks with new task information.
This approach improves the efficiency of learning by transforming previously learned knowledge to absorb new information, while also strengthening previous knowledge and suppressing unnecessary parameter generation.
Smart Images

Figure KR2024006190_22052025_PF_FP_ABST
Abstract
Description
Neural network continuous learning device and method
[0001] It relates to a neural network continuous learning device and method.
[0002] Existing deep learning algorithms are optimized for learning on a single task (either single or multiple tasks) within a single timeframe. Deep learning networks trained on different tasks in each timeframe lose knowledge of tasks learned in previous timeframes and optimize for the most recently learned task. In particular, without access to previous task data, it is difficult to retain knowledge of previous tasks and incorporate new knowledge from new tasks.
[0003] Catastrophic forgetting, where a network loses its previous knowledge when learning a new task, is an obstacle that continuous learning, which learns tasks over successive time periods, seeks to mitigate.
[0004] The purpose is to provide a neural network continuous learning device and method.
[0005] According to one aspect, a neural network continuous learning device includes an interface unit for data input / output; and a control unit connected to the interface unit, wherein the control unit can generate a transformed parameter set for the base network based on an old parameter set for the base network and a generated parameter set generated through a hyper network including one or more task layer attention blocks.
[0006] The hyper network can generate a set of generated parameters by taking as input a task token containing identification information for a new task and a layer embedding containing layer identification information of a previous parameter set.
[0007] The control unit can perform an element-wise product of the generation parameter set and the previous parameter set.
[0008] The control unit can generate a transformed parameter set by performing a union operation on the element-wise multiplied parameter set and a new parameter set for the new task.
[0009] The control unit can train the neural network based on the output of input data for a new task input into the basic network to which the transformed parameters are applied and the first loss function generated based on the label of the input data.
[0010] The control unit can train the neural network based on a second loss function generated based on the difference between the generation parameters generated in the previous task and the generation parameters generated in the new task.
[0011] The control unit can train the neural network based on a third loss function to achieve sparsity of the generation parameter set.
[0012] The third loss function can be applied to the element-wise sensitivity of the generation parameter set.
[0013] According to one aspect, a method for continuous learning of a neural network performed on a computing device having one or more processors and a memory storing one or more programs executed by the one or more processors may include the steps of generating a generated parameter set through a hyper network including one or more task layer attention blocks; and generating a transformed parameter set for a base network based on an old parameter set for a base network and the generated parameter set.
[0014] According to one aspect, a computer program stored in a non-transitory computer readable storage medium, the computer program including one or more instructions, which instructions, when executed by a computing device having one or more processors, cause the computing device to perform the steps of: generating a generated parameter set through a hyper network including one or more task layer attention blocks; and generating a transformed parameter set for a base network based on an old parameter set for the base network and the generated parameter set.
[0015] In one aspect, continuous learning based on parameter isolation transforms previously learned knowledge to absorb information from new tasks, thereby increasing learning efficiency. Furthermore, prior knowledge can be augmented with a transformer-based hyper-network that generates parameters optimized for the new task while suppressing the generation of unnecessary parameters with minimal impact on loss.
[0016] Figure 1 is a configuration diagram of a neural network continuous learning device according to one embodiment.
[0017] Figure 2 is an exemplary diagram illustrating the configuration of a neural network continuous learning device according to one embodiment.
[0018] Figure 3 is a flowchart illustrating a neural network continuous learning method according to one embodiment.
[0019] FIG. 4 is a block diagram illustrating a computing environment including a computing device suitable for use in exemplary embodiments.
[0020] Hereinafter, specific embodiments of the present invention will be described with reference to the drawings. The following detailed description is provided to facilitate a comprehensive understanding of the methods, devices, and / or systems described herein. However, these are merely examples and the present invention is not limited thereto.
[0021] In describing embodiments of the present invention, if a detailed description of a known technology related to the present invention is judged to unnecessarily obscure the gist of the present invention, the detailed description will be omitted. In addition, the terms described below are terms defined in consideration of their functions in the present invention, and this may vary depending on the intention or custom of the user or operator. Therefore, the definitions should be made based on the contents throughout this specification. The terminology used in the detailed description is only for the purpose of describing embodiments of the present invention and should not be limited in any way. Unless clearly used otherwise, the singular form includes the plural form. In this description, expressions such as "comprises" or "having" are intended to indicate certain features, numbers, steps, operations, elements, parts or combinations thereof, and should not be construed to exclude the presence or possibility of one or more other features, numbers, steps, operations, elements, parts or combinations thereof other than those described.
[0022] Furthermore, while terms such as "first" and "second" may be used to describe various components, the components should not be limited by these terms. Terms may be used to distinguish one component from another. For example, without departing from the scope of the present invention, a first component may be referred to as a "second component," and similarly, a second component may also be referred to as a "first component."
[0023] FIG. 1 is a configuration diagram of a neural network continuous learning device according to one embodiment, and FIG. 2 is an exemplary diagram for explaining the configuration of a neural network continuous learning device according to one embodiment.
[0024] Continuous learning in neural networks refers to a method by which deep learning models continuously learn from new data. Typical deep learning models train on large datasets and learn generalized patterns based on those datasets. However, in real-world environments, new data is constantly generated, and this data is likely to differ from existing data. Continuous learning addresses this issue by continuously learning from new data and gradually expanding the knowledge acquired.
[0025] For example, a neural network continuous learning device can perform continuous learning on a neural network by reframing the parameter selection problem as a generation problem. To this end, the neural network continuous learning device can be composed of two networks: a base network and a hyper network. The transformer-based hyper network can generate a small set of meaningful parameters that significantly influence the loss. The neural network continuous learning device can generate parameters using information about the current task and embeddings of information about the layers of the base network as inputs to the hyper network. This allows the neural network continuous learning device to generate customized parameters for new tasks while suppressing the generation of unnecessary parameters that do not significantly affect the loss function.
[0026] Referring to FIG. 1, a neural network continuous learning device (100) may include an interface unit (110) for data input / output and a control unit (120) connected to the interface unit (110).
[0027] According to one embodiment, the control unit (120) can generate a transformed parameter set for the base network based on an old parameter set for the base network and a generated parameter set generated through a hyper network including one or more task layer attention blocks (TLAs).
[0028] According to an example, the control unit (120) operates at the t-th time zone. Task-specific parameters can be assigned to the base network F(.) through learning-pruning-relearning. At this time, is the nth input data of the tth task, is the corresponding label. The most basic approach is to assign a label to the network after task learning. It goes through a process of pruning and relearning. becomes. Finally, the base network can be parameterized as . At this time, am.
[0029] In one embodiment, the hyper network can generate a set of generated parameters by receiving a task token containing identification information for a new task and a layer embedding containing layer identification information of a previous parameter set. Referring to FIG. 2, the hyper network H(.) can receive a task token and a layer embedding as input.
[0030] According to an example, the control unit (120) sets the parameter sets assigned to previous tasks in the t-th time zone. second Hyper-network parameterized by can be converted using the sparse parameters generated in. For example, the control unit (120) can convert the parameter sets of previous tasks assigned to the lth layer of the base network to fit the new task. The hyper network can be converted into a task token representing the task. and a set of layer embeddings of the base network Generate parameter sets using as input can be created. At this time, class represents the number of layer embeddings and the layer index for the lth layer, respectively.
[0031] For example, the control unit (120) generates a set of parameters To generate a work token e t and the layer embedding set C l Hierarchical embedding of c n can be configured as input by sequentially concatenating them. That is, the input of the hyper network is am.
[0032] As an example, the mth task layer attention block (TLA) of the hyper network is You can use it as input and perform calculations as in the mathematical formula below.
[0033] [Mathematical Formula 1]
[0034]
[0035] At this time It is. N b , LN, MSA, and MLP are the number of TLA blocks, layer normalization, multi-head self-attention, and multi-layer perceptron, respectively. Input MSA is defined as follows:
[0036] [Equation 2]
[0037]
[0038] At this time, is a learnable matrix, and N h is the number of heads in the MSA layer. The i-th head is is defined as, , , and, , , are the projection matrices of the i-th head, respectively. The control unit (120) can generate a generation parameter set for the parameter set of the l-th layer of the base network using the l-th layer embedding set, and can be expressed as the following mathematical formula.
[0039] [Equation 3]
[0040]
[0041] Here, g(.) represents the sigmoid function that follows the fully connected layer. Is Indicates the first column of .
[0042] According to one embodiment, the control unit (120) can perform an element-wise product of the generated parameter set and the previous parameter set. For example, the control unit (120) can perform the previous parameter sets of the lth layer of the base network for the tth task learning. , and can be converted using the generated parameter set. The converted previous parameter sets are Defined as can be converted as follows. At this time The symbol is defined as element-wise multiplication.
[0043] According to one embodiment, the control unit (120) can generate a transformed parameter set by performing a union operation on the element-wise multiplied parameter set and the new parameter set for the new task. The control unit (120) can generate a base network based on the transformed parameter set. can be parameterized as
[0044] According to one embodiment, the control unit (120) can train the neural network based on the output of input data for a new task input into the basic network to which the transformed parameters are applied and a first loss function generated based on the label of the input data.
[0045] For example, the loss learned using the parameter sets assigned to task t and the previous parameter sets transformed to incorporate knowledge of the t-th task is as follows.
[0046] [Equation 4]
[0047]
[0048] In one embodiment, the control unit (120) can train the neural network based on a second loss function generated based on the difference between the generation parameters generated in the previous task and the generation parameters generated in the new task. For example, the loss function for the current hyper network to generate the generation parameter sets required to perform the previous task can be expressed as the following mathematical equation.
[0049] [Equation 5]
[0050]
[0051] At this time, refers to the parameter set obtained through the hyper network trained up to the t-1th task.
[0052] In one embodiment, the control unit (120) can train the neural network based on a third loss function to achieve sparsity in the generated parameter set. For example, the third loss function can be applied to element-wise sensitivities of the generated parameter set.
[0053] For example, a loss function is used to exclude unnecessary parameters from the parameter set generated by the hyper network. can be minimized. However, may become excessively rarefied, Among the parameters, sparsity should not be given to parameters that have a large impact on the loss function when not used.
[0054] For example, the control unit (120) Sensitivity to the loss function for each element Taylor expansion approximation can be used to measure . For example, is a vectorized form of The sensitivity to loss of the kth parameter can be expressed as the mathematical formula below.
[0055] [Equation 6]
[0056]
[0057] At this time, is an all-zero vector whose kth element is 1. is the loss for the nth sample.
[0058] For example, the above sensitivity can be added as an additional loss term to reduce the original score and the score of the sparse parameters.
[0059] [Equation 7]
[0060]
[0061] At this time, and represent the norm and indicator function, respectively.
[0062] As an example, the final loss when learning the tth task can be expressed as the mathematical formula below.
[0063] [Equation 8]
[0064]
[0065] Here, and are hyperparameters for balancing losses. For example, the control unit (120) The parameter set of the tth task of the current time zone of the base network through can update parameters related to the hyper network. , and Is and It can be updated through. Also, Through and may be updated.
[0066] Figure 3 is a flowchart illustrating a neural network continuous learning method according to one embodiment.
[0067] According to one embodiment, a neural network continuous learning device may be a computing device having one or more processors and a memory storing one or more programs executed by the one or more processors.
[0068] According to one embodiment, a neural network continuous learning device can generate a generated parameter set through a hyper network including one or more task layer attention blocks (310). Thereafter, the neural network continuous learning device can generate a transformed parameter set for the base network based on an old parameter set for the base network and the generated parameter set (320).
[0069] In Fig. 3, the embodiments that overlap with the contents described with reference to Figs. 1 and 2 among the embodiments are omitted.
[0070] FIG. 4 is a block diagram illustrating a computing environment (10) including a computing device suitable for use in exemplary embodiments. In the illustrated embodiment, each component may have different functions and capabilities other than those described below, and may include additional components other than those described below.
[0071] The illustrated computing environment (10) includes a computing device (12). In one embodiment, the computing device (12) may be a neural network continuous learning device.
[0072] A computing device (12) includes at least one processor (14), a computer-readable storage medium (16), and a communication bus (18). The processor (14) may cause the computing device (12) to operate according to the exemplary embodiments described above. For example, the processor (14) may execute one or more programs stored in the computer-readable storage medium (16). The one or more programs may include one or more computer-executable instructions, which, when executed by the processor (14), may be configured to cause the computing device (12) to perform operations according to the exemplary embodiments.
[0073] A computer-readable storage medium (16) is configured to store computer-executable instructions or program code, program data, and / or other suitable forms of information. A program (20) stored in the computer-readable storage medium (16) includes a set of instructions executable by the processor (14). In one embodiment, the computer-readable storage medium (16) may be a memory (volatile memory such as random access memory, non-volatile memory, or a suitable combination thereof), one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, any other form of storage medium that can be accessed by the computing device (12) and store desired information, or a suitable combination thereof.
[0074] A communication bus (18) interconnects various other components of the computing device (12), including the processor (14) and computer-readable storage media (16).
[0075] The computing device (12) may also include one or more input / output interfaces (22) that provide interfaces for one or more input / output devices (24) and one or more network communication interfaces (26). The input / output interfaces (22) and the network communication interfaces (26) are connected to the communication bus (18). The input / output devices (24) may be connected to other components of the computing device (12) via the input / output interfaces (22). Exemplary input / output devices (24) may include input devices such as pointing devices (such as a mouse or a trackpad), a keyboard, a touch input device (such as a touchpad or a touchscreen), a voice or sound input device, various types of sensor devices and / or photographing devices, and / or output devices such as display devices, printers, speakers and / or network cards. The exemplary input / output devices (24) may be included within the computing device (12) as a component constituting the computing device (12), or may be connected to the computing device (12) as a separate device distinct from the computing device (12).
[0076] While representative embodiments of the present invention have been described in detail above, those skilled in the art will appreciate that various modifications to the above-described embodiments are possible without departing from the scope of the present invention. Therefore, the scope of the present invention should not be limited to the described embodiments, but should be defined not only by the claims set forth below but also by equivalents thereof.
Claims
1. Interface section for data input / output; and Includes a control unit connected to the above interface unit, The above control unit A neural network continuous learning device that generates a transformed parameter set for the base network based on a generated parameter set generated through a hyper network including an old parameter set for the base network and one or more task layer attention blocks.
2. In paragraph 1, The above hyper network is A neural network continuous learning device that generates a set of generated parameters by inputting a task token containing identification information for a new task and a layer embedding containing layer identification information of a previous parameter set.
3. In paragraph 1, The above control unit A neural network continuous learning device that performs an element-wise product of the above-mentioned generated parameter set and the above-mentioned previous parameter set.
4. In paragraph 1, The above control unit A neural network continuous learning device that generates a transformed parameter set by performing a union operation on the above element-wise multiplied parameter set and a new parameter set for a new task.
5. In paragraph 1, The above control unit A neural network continuous learning device that learns a neural network based on a first loss function generated based on the output of input data for a new task input into a basic network to which the above-mentioned converted parameters are applied and the label of the input data.
6. In paragraph 1, The above control unit A neural network continuous learning device that learns a neural network based on a second loss function generated based on the difference between the generation parameters generated in a previous task and the generation parameters generated in a new task.
7. In paragraph 1, The above control unit A neural network continuous learning device that learns a neural network based on a third loss function to achieve sparsity of the above-mentioned generation parameter set.
8. In paragraph 7, The third loss function is a neural network continuous learning device to which element-wise sensitivity of the generation parameter set is applied.
9. One or more processors, and A method performed on a computing device having a memory storing one or more programs executed by one or more processors, A step of generating a generated parameter set via a hyper network including one or more task layer attention blocks; and A method comprising the step of generating a transformed parameter set for the base network based on an old parameter set for the base network and the generated parameter set.
10. A computer program stored in a non-transitory computer readable storage medium, The computer program comprises one or more instructions, which, when executed by a computing device having one or more processors, cause the computing device to: A step of generating a generated parameter set via a hyper network including one or more task layer attention blocks; and A computer program that causes the computer to perform the step of generating a transformed parameter set for the base network based on an old parameter set for the base network and the generated parameter set.
Citation Information
Patent Citations
Smart mirror coating glasses having muti-color mirror patterns represented on the lens
KR1020250014271A
A neural network apparatus and neural network learning method for performing continuous learning using a correlation analysis algorithm between tasks
KR102583943B1
KR20210157826A
KR20220003093A
KR20230149554A