Large model continuous learning method and device, electronic equipment and storage medium
By using new task data to train the training model in the continuous learning method of large models, and weighted processing of loss data and parameter adjustment operations, the model overfitting and forgetting problems in incremental pretraining is solved, achieving more stable model updates and performance improvements.
Patent Information
- Application Number
- CN202510140484.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-16
AI Technical Summary
During incremental pre-training, the new training data is different from the original training data, which can easily lead to model overfitting and catastrophic forgetting problems, and the mapping relationship changes drastically.
By obtaining the data to be trained and the model to be trained corresponding to the new task, some parameters are frozen, and training is performed based on the new task data, and the first loss data is weighted to obtain the target loss data, and the parameter adjustment operation is performed on the unfrozen parameters until the training termination condition is met.
It effectively alleviates the problem of drastic changes in mapping relationships caused by parameter updates, reduces the problem of overfitting and catastrophic forgetting of the model, and improves the overall performance and generalization capabilities of the model.
Smart Images

Figure CN120012868A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a large-model continuous learning method, device, electronic device and storage medium. Background Art
[0002] In recent years, large-scale pre-trained language models have achieved remarkable results in the field of natural language processing. These models are pre-trained on a large amount of text data through unsupervised learning, and can be subsequently adapted to specific downstream tasks through fine-tuning. However, training a large-scale language model from scratch is very time-consuming and resource-intensive, usually requiring a lot of computing resources and a long training cycle.
[0003] In order to reduce time and resource costs, existing open source or commercial pre-trained models are often used. However, these open source or commercial pre-trained models are generally general models. In order to adapt to specific business scenarios, incremental pre-training can be used, that is, additional data is used for further training based on the existing pre-trained model. However, in the incremental pre-training process, since the new training data is different from the original training data of the model, it is easy to cause overfitting and catastrophic forgetting problems, which makes the model's processing ability enhanced and its processing ability weakened. Summary of the invention
[0004] The embodiments of the present disclosure provide a large-model continuous learning method, device, electronic device and storage medium, which can effectively alleviate the problem of drastic changes in mapping relationships caused by parameter updates, while effectively reducing the overfitting and catastrophic forgetting problems of the model.
[0005] According to one aspect of the present disclosure, a large model continuous learning method is provided, the method comprising: obtaining data to be trained and a model to be trained corresponding to a new task, wherein some parameters of the model to be trained are frozen, and the data to be trained comprises input data and expected output data; training the model to be trained based on the training data corresponding to the new task to obtain first loss data; weighting the first loss data according to preset loss data to obtain target loss data of the model to be trained, wherein the preset loss data is the loss data generated by the frozen parameters in the model to be trained; adjusting the parameters of the unfrozen model to be trained based on the target loss data, repeating the step of training the model to be trained based on the data to be trained until the training termination condition is met, and obtaining the target task training model corresponding to the new task.
[0006] Optionally, weighted processing is performed on the first loss data according to the preset loss data to obtain target loss data of the model to be trained, including: obtaining a weighting coefficient and determining the product of the weighting coefficient and the preset loss data as the second loss data, wherein the weighting coefficient decays linearly or nonlinearly with the increase in the number of training steps; determining the difference between the first loss data and the second loss data as the loss difference; and determining the power of the loss difference as a natural number as the target loss data.
[0007] Optionally, the frozen parameters in the model to be trained are all parameters in a preset training model, and the unfrozen parameters in the model to be trained are newly added fine-tuning parameters.
[0008] Optionally, obtaining the preset loss data includes: freezing all parameters in the preset training model; and training the pre-trained model based on the to-be-trained data corresponding to the new task to obtain the preset loss data.
[0009] Optionally, training the pre-trained model based on the to-be-trained data corresponding to the new task to obtain the preset loss data includes: inputting the input data into the preset training model for forward reasoning to obtain preset output data; and determining the preset loss data based on the preset output data and the expected output data.
[0010] Optionally, the large model continuous learning method also includes: after each preset number of training steps, using the negative value of the current preset loss data as the target loss data of the model to be trained.
[0011] According to one aspect of the present disclosure, a large-model continuous learning device is provided, comprising: an acquisition module, used to acquire data to be trained and a model to be trained corresponding to a new task, wherein some parameters of the model to be trained are frozen, and the data to be trained include input data and expected output data; a training module, used to train the model to be trained based on the training data corresponding to the new task to obtain first loss data; a data processing module, used to weight the first loss data according to preset loss data to obtain target loss data of the model to be trained, wherein the preset loss data is the loss data generated by the frozen parameters in the model to be trained; the training module is also used to adjust the parameters of the unfrozen parameters in the model to be trained based on the target loss data, and repeat the step of training the model to be trained based on the data to be trained until the training termination condition is met to obtain the target task training model corresponding to the new task.
[0012] According to one aspect of the present disclosure, an electronic device is proposed, which includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory, wherein the program is executed by the processor to realize the large model continuous learning method as described above.
[0013] According to one aspect of the present disclosure, a computer-readable storage medium is proposed, wherein the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the large model continuous learning method as described above.
[0014] The present disclosure proposes a large-model continuous learning method, device, electronic device and storage medium, which use the to-be-trained data corresponding to the new task to train the to-be-trained model to obtain first loss data, and weight the first loss data according to the preset loss data corresponding to the frozen parameters in the to-be-trained model to obtain target loss data, and adjust the parameters of the unfrozen parameters in the to-be-trained model according to the target loss data, so as to balance the learning of new tasks and the retention of old knowledge during the training process, which can effectively alleviate the problem of drastic changes in mapping relationships caused by parameter updates, and effectively reduce the overfitting and catastrophic forgetting problems of the model.
[0015] Furthermore, the weighting coefficient decays linearly or nonlinearly with the increase of the number of training steps. By adjusting the weighting coefficient of the preset loss data, the adaptation speed of the model to the new data can be controlled to avoid the degradation of model performance caused by too fast parameter update. At the same time, by keeping some parameters unchanged (freezing parameters), the model can retain the key knowledge learned in the pre-training stage and reduce the occurrence of catastrophic forgetting.
[0016] Furthermore, after each preset number of training steps, the negative value of the current preset loss data is used as the target loss data of the model to be trained. During the training process, the unfrozen parameters can be updated along the negative direction of the gradient, that is, the value of the loss function is consciously expanded. This can effectively reduce the risk of the model falling into a local minimum, thereby improving the model's global optimization ability, helping to find a better solution, and improving the overall performance and generalization ability of the model.
[0017] Other features and advantages of the present disclosure will be described in the following description, and partly become apparent from the description, or understood by practicing the present disclosure. The purpose and other advantages of the present disclosure can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are used to provide further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation on the technical solution of the present disclosure.
[0019] Figure 1 is a system architecture diagram of the large model continuous learning method using the embodiment of the present disclosure;
[0020] Figure 2 It is a main flow chart of a large model continuous learning method according to an embodiment of the present disclosure;
[0021] Figure 3 yes Figure 2 A sub-flowchart of step S230;
[0022] Figure 4 is a flow chart of a large model continuous learning method according to an embodiment of the present disclosure;
[0023] Figure 5 It is a structural diagram of a large model continuous learning device proposed in an embodiment of the present disclosure;
[0024] Figure 6 It is a schematic diagram of the structure of an electronic device proposed in one embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solution and advantages of the present disclosure more clear, the present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0026] Before further describing the embodiments of the present disclosure in detail, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are subject to the following interpretations:
[0027] Pretrained Models (PM): are models trained on large datasets. These models usually have good performance on certain general tasks. Pretrained models are usually pre-trained on specific tasks and then fine-tuned in specific application scenarios. The training process of pretrained models includes a pretraining phase and a fine-tuning phase. In the pretraining phase, the model is trained on a large-scale general dataset to learn general knowledge and features; in the fine-tuning phase, the pretrained model is further trained on a dataset of a specific task to adapt to the needs of the specific task. The training data in the fine-tuning phase is usually much less than that in the pretraining phase.
[0028] Pretrained Language Model (PLM), the idea of pretraining is that the model parameters are no longer randomly initialized, but are pre-trained through some tasks to obtain a set of model parameters, and then the model is initialized with this set of parameters and then trained.
[0029] Continuous learning (CL) refers to feeding data to the network model in a non-stationary way. There is no intersection between each data set, and the system usually cannot see the trained data again. As the amount of training data increases, when a new service or task is added to the system, model training through multi-task is more time-consuming, but continuous learning can save time.
[0030] Parameter-Efficient Fine-Tuning (PEFT) is a model fine-tuning technique that allows adapting to specific downstream tasks by fine-tuning only a small number of parameters without adjusting all parameters of the pre-trained model, thereby significantly reducing computing and storage costs. This method reduces computing power expenditure, speeds up model adaptation, and avoids catastrophic forgetting while maintaining performance comparable to full parameter fine-tuning.
[0031] In related technologies, in order to adapt the pre-trained module to specific business scenarios, superimposed parameter fine-tuning methods such as PromptTuning, P-Tuning, LoRA and Adapter can be used to adjust only a small number of parameters when necessary, so as to maintain the original knowledge of the model while enhancing its processing capabilities for specific tasks; you can also choose to freeze the backbone parameters of the pre-trained model and only train additional parameters related to new tasks; you can also directly use the generation capabilities of the large model to change the input by designing clever prompts to obtain outputs in specific business scenarios. However, although the fine-tuning methods of superimposed parameters and frozen parameters can change the model mapping relationship to adapt to specific tasks, they are prone to catastrophic forgetting problems; although the method of adding prompts does not need to change the model mapping relationship, as the number of business scenarios increases, more and more prompt prefixes will be added, resulting in a reduction in the number of tokens that can be input, which will also greatly increase the encoding time. Based on this, the present disclosure proposes a large model continuous learning method, device, electronic device and storage medium. The method uses the to-be-trained data corresponding to the new task to train the to-be-trained model to obtain first loss data, and weights the first loss data according to the preset loss data corresponding to the frozen parameters in the to-be-trained model to obtain target loss data, and adjusts the unfrozen parameters in the to-be-trained model according to the target loss data. During the training process, the learning of new tasks and the retention of old knowledge are balanced, which can effectively alleviate the problem of drastic changes in the mapping relationship caused by parameter updates, and effectively reduce the overfitting and catastrophic forgetting problems of the model.
[0032] Description of the system architecture of the present disclosure embodiment
[0033] Figure 1 It is a system architecture diagram used in the large-model continuous learning method of the embodiment of the present disclosure, which includes: a terminal 110, the Internet 120, a gateway 130, and a server 140.
[0034] The terminal 110 is a device for input. It includes desktop computers, laptop computers, PDAs (personal digital assistants), mobile phones, vehicle-mounted terminals, home theater terminals, dedicated terminals, and other forms. In addition, it can be a single device or a collection of multiple devices. For example, multiple devices are connected through a local area network and work together using a common display device to form a terminal. The terminal 110 can also communicate with the Internet 120 in a wired or wireless manner to exchange data.
[0035] The gateway 130 is also called an internetwork connector or a protocol converter. The gateway 130 realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. The gateway 130 is a translator between two systems that use different communication protocols, data formats or languages, or even have completely different architectures. At the same time, the gateway 130 can also provide filtering and security functions. The message sent by the terminal 110 to the server 140 must be sent to the corresponding server 140 through the gateway 130. The message sent by the server 140 to the terminal 110 must also be sent to the corresponding terminal 110 through the gateway 130.
[0036] The server 140 refers to a computer system that can provide services to the terminal 110. Compared with the terminal 110, the server 140 has higher requirements in terms of stability, security, performance, etc. The server 140 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part of a high-performance computer (such as a virtual machine), a combination of parts of multiple high-performance computers (such as virtual machines), etc. The server 140 can also communicate with the Internet 120 in a wired or wireless manner to exchange data.
[0037] Server 140 is used to obtain data to be trained and a model to be trained corresponding to the new task, wherein some parameters of the model to be trained are frozen, and the data to be trained includes input data and expected output data; the model to be trained is trained based on the training data corresponding to the new task to obtain first loss data; the first loss data is weighted according to preset loss data to obtain target loss data of the model to be trained, wherein the preset loss data is the loss data generated by the frozen parameters in the model to be trained; the unfrozen parameters in the model to be trained are adjusted based on the target loss data, and the step of training the model to be trained based on the data to be trained is repeated until the training termination condition is met to obtain the target task training model corresponding to the new task.
[0038] Overall implementation of the large model continuous learning method of the disclosed embodiment
[0039] The present disclosure provides a large model continuous learning method, which is applied to a large model continuous learning device. Figure 2 , large model continuous learning methods include:
[0040] Step S210, obtaining data to be trained and a model to be trained corresponding to the new task, wherein some parameters of the model to be trained are frozen, and the data to be trained includes input data and expected output data;
[0041] Step S220, training the to-be-trained model based on the to-be-trained data corresponding to the new task to obtain first loss data;
[0042] Step S230, performing weighted processing on the first loss data according to preset loss data to obtain target loss data of the model to be trained, wherein the preset loss data is loss data generated by freezing parameters in the model to be trained;
[0043] Step S240, adjusting the unfrozen parameters in the model to be trained based on the target loss data, repeating the step of training the model to be trained based on the data to be trained until the training termination condition is met, and obtaining the target task training model corresponding to the new task.
[0044] In step S210, the data to be trained refers to the data set corresponding to the training new task. The new task and the old task can be tasks in different fields or tasks of different categories in the same field. The data set can be at least one of image data, video data, audio data, text data, etc., such as a large amount of dialogue and question-and-answer data, knowledge retrieval data, text classification, information extraction and summarization, etc. Natural semantic processing task data. The data to be trained includes input data and corresponding expected output data (also called "labels").
[0045] The model to be trained can be obtained from a preset training model, and the preset training model can be a pre-trained model trained on a large data set, or a model trained based on data corresponding to an old task, and this embodiment does not limit this. The frozen parameters in the model to be trained are all the parameters in the preset training model, and the unfrozen parameters in the model to be trained are the newly added fine-tuning parameters. The model to be trained can be obtained by freezing a part of the parameters in the preset training model, and the other part of the parameters are not frozen for training; it can also be obtained by freezing all the parameters in the preset training model and adding some parameters for training.
[0046] In step S220, the data to be trained is input into the model to be trained, and the first output data output1 is obtained in the forward reasoning process, and the first output data output1 and the expected output data label are calculated to obtain the first loss data. For example, the first output data output1 and the expected output data label are calculated by cross entropy to obtain the first loss data Loss1, that is, Loss1=CrossEntropy(output1,label).
[0047] In step S230, the preset loss data Loss0 is the loss data generated by freezing the parameters in the model to be trained. Each input of the data to be trained corresponds to a preset loss data Loss0. The preset loss data Loss0 is obtained by inputting the data to be trained into the preset training model with all parameters frozen for forward reasoning calculation. The step of obtaining the preset loss data Loss0 includes: freezing all parameters in the preset training model; training the pre-trained model based on the data to be trained corresponding to the new task to obtain the preset loss data. Specifically, the data to be trained is input into the preset training model with all parameters frozen for forward reasoning to obtain the preset output data output0, and the preset output data output0 and the expected output data lable are calculated to obtain the preset loss data. For example, the preset output data output0 and the expected output data lable are calculated by cross entropy to obtain the preset loss data Loss0.
[0048] That is, Loss0=CrossEntropy(output0,label).
[0049] The preset loss data Loss0 can be obtained in advance, and the input data to be trained and the preset loss data Loss0 are constructed into a hash mapping table for storage, or the preset training model is trained while the training model is trained based on the data to be trained to obtain the preset loss data Loss0.
[0050] The target loss data Loss is obtained by weighting the first loss data Loss1 according to the preset loss data Loss0, and the calculation formula is: , where n is any natural number, weight is a weighting coefficient, and 0<weight<1. In some embodiments, when n is a natural number e, Loss=exp(Loss1-weight*Loss0).
[0051] Specifically, see Figure 3 , step S230 includes the following steps.
[0052] Step S310, obtaining preset loss data and its corresponding weighting coefficient, wherein the weighting coefficient decays linearly or nonlinearly with the increase of the number of training steps;
[0053] Step S320, determining the product of the preset loss data and the weighting coefficient as the second loss data;
[0054] Step S330, determining the difference between the first loss data and the second loss data as a loss difference;
[0055] Step S340: determining the power of the loss difference of a natural number as the target loss data.
[0056] In step S310, the corresponding preset loss data Loss0 can be obtained from the hash mapping table according to the input data to be trained, or the preset training model based on the input data to be trained can be trained to obtain the preset loss data Loss0. A corresponding weighting coefficient weight is set for the preset loss data corresponding to each input data to be trained. The weighting coefficient weight decays linearly or nonlinearly with the increase of the number of training steps step.
[0057] In step S320, the product of the preset loss data Loss0 and the weighting coefficient weight is determined as the second loss data Loss2, that is, Loss2=weight*Loss0.
[0058] In step S330 , the difference between the first loss data Loss1 and the second loss data Loss2 is determined as a loss difference Loss_d=Loss1−Loss2.
[0059] In step S340, the target loss data Loss is the loss difference power of the natural number n, that is, , where n is any natural number, weight is the weighting coefficient, 0<weight<1.
[0060] In some embodiments, when n is a natural number e, the target loss data Loss=exp(Loss_d)=exp(Loss1-weight*Loss0).
[0061] In step S240, the unfrozen parameters in the training model are adjusted based on the target loss data Loss, that is: ,in, are the unfrozen parameters in the model to be trained in the i-th round of training, is the target loss data obtained from the i-th round of training, is the learning rate of the i-th round of training, is s times the learning rate of the previous round, 0<s<0.5. Repeat the step of training the model to be trained based on the training data until the training termination condition is met, and obtain the target task training model corresponding to the new task.
[0062] The large model continuous learning method provided by the embodiment of the present disclosure uses the to-be-trained data corresponding to the new task to train the to-be-trained model to obtain first loss data, and weights the first loss data according to the preset loss data corresponding to the frozen parameters in the to-be-trained model to obtain target loss data, and adjusts the unfrozen parameters in the to-be-trained model according to the target loss data, balancing the learning of new tasks and the retention of old knowledge during the training process, which can effectively alleviate the problem of drastic changes in mapping relationships caused by parameter updates, and effectively reduce the overfitting and catastrophic forgetting problems of the model.
[0063] Furthermore, the weighting coefficient decays linearly or nonlinearly with the increase of the number of training steps. By adjusting the weighting coefficient of the preset loss data, the adaptation speed of the model to the new data can be controlled to avoid the degradation of model performance caused by too fast parameter update. At the same time, by keeping some parameters unchanged (freezing parameters), the model can retain the key knowledge learned in the pre-training stage and reduce the occurrence of catastrophic forgetting.
[0064] In some embodiments, see Figure 4 , the large model continuous method includes the following steps.
[0065] Step S410, obtaining data to be trained and a model to be trained corresponding to the new task, wherein some parameters of the model to be trained are frozen, and the data to be trained includes input data and expected output data;
[0066] Step S420, training the to-be-trained model based on the to-be-trained data corresponding to the new task to obtain first loss data;
[0067] Step S430, performing weighted processing on the first loss data according to preset loss data to obtain target loss data of the model to be trained, wherein the preset loss data is loss data generated by freezing parameters in the model to be trained;
[0068] Step S450, after each preset number of training steps, taking the negative value of the current preset loss data as the target loss data of the model to be trained;
[0069] Step S440, adjusting the unfrozen parameters in the model to be trained based on the target loss data, repeating the step of training the model to be trained based on the data to be trained until the training termination condition is met, and obtaining the target task training model corresponding to the new task.
[0070] Step S410 is the same as step S210, step S420 is the same as step S220, step S430 is the same as step S230, and step S440 is the same, which will not be repeated here. In step S450, each batch of data to be trained is input into the model to be trained for forward reasoning and back propagation, which is a round of training. The preset number of training steps N is a hyperparameter and can be set in advance. After the model has been trained for N rounds, the target loss data of the Nth round can be replaced with the negative value of the preset loss data Loss0 corresponding to the data to be trained input in the Nth round, that is, Loss_N=-Loss0_N. In this way, the unfrozen parameters can be updated along the negative direction of the gradient during the training process, that is, the value of the loss function is consciously expanded, which can effectively reduce the risk of the model falling into a local minimum, thereby improving the model's ability to optimize globally, helping to find a better solution, and improving the overall performance and generalization ability of the model.
[0071] Description of the apparatus and device of the present disclosure
[0072] See also Figure 5 The embodiment of the present disclosure also provides a large model continuous learning device 500, including an acquisition module 510, a training module 520 and a data processing module 530.
[0073] The acquisition module 510 is used to acquire the data to be trained and the model to be trained corresponding to the new task, wherein some parameters of the model to be trained are frozen, and the data to be trained includes input data and expected output data;
[0074] The training module 520 is used to train the to-be-trained model based on the training data corresponding to the new task to obtain first loss data;
[0075] The data processing module 530 is used to perform weighted processing on the first loss data according to preset loss data to obtain target loss data of the model to be trained, wherein the preset loss data is loss data generated by freezing parameters in the model to be trained;
[0076] The training module 520 is also used to adjust the unfrozen parameters in the model to be trained based on the target loss data, and repeat the step of training the model to be trained based on the data to be trained until the training termination condition is met, thereby obtaining the target task training model corresponding to the new task.
[0077] The large model continuous learning device 500 disclosed in the present invention is used to execute the large model continuous learning method as in the above embodiment. Its specific processing process is the same as the large model continuous learning method in the above embodiment, and will not be repeated here.
[0078] The present disclosure also provides an electronic device 600, including:
[0079] at least one processor, and
[0080] a memory communicatively connected to at least one processor; wherein,
[0081] The memory stores instructions, and the instructions are executed by at least one processor so that when the at least one processor executes the instructions, a method as described in any one of the above embodiments of the present application is implemented.
[0082] Combine the following Figure 6 The hardware structure of the electronic device is described in detail. The electronic device includes: a processor 610 , a memory 620 , an input / output interface 630 , a communication interface 640 and a bus 650 .
[0083] The processor 610 may be implemented by a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present disclosure;
[0084] The memory 620 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 620 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 620, and the processor 610 calls and executes the large model continuous learning method of the embodiment of the present disclosure;
[0085] Input / output interface 630, used to implement information input and output;
[0086] Communication interface 640, used to realize communication interaction between the device and other devices, which can be realized through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); and
[0087] bus 650 , which transmits information between the various components of the device (e.g., processor 610 , memory 620 , input / output interface 630 , and communication interface 640 );
[0088] The processor 610 , the memory 620 , the input / output interface 630 , and the communication interface 640 are connected to each other in communication within the device via a bus 650 .
[0089] An embodiment of the present application also provides a computer-readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the large model continuous learning method of the above embodiment, which will not be repeated here.
[0090] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present disclosure described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "comprises" and "comprising" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0091] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0092] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood as not including the number, and above, below, within, etc. are understood as including the number.
[0093] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0094] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0095] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0096] It should also be understood that the various implementations provided in the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.
[0097] The above is a specific description of the implementation methods of the present disclosure, but the present disclosure is not limited to the above implementation methods. Technical personnel familiar with the art can also make various equivalent modifications or substitutions without violating the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.
Claims
1. A large model continuous learning method, characterized in that: include: Obtaining data to be trained and a model to be trained corresponding to the new task, wherein some parameters of the model to be trained are frozen, and the data to be trained includes input data and expected output data; Training the to-be-trained model based on the to-be-trained data corresponding to the new task to obtain first loss data; Performing weighted processing on the first loss data according to preset loss data to obtain target loss data of the model to be trained, wherein the preset loss data is loss data corresponding to the frozen parameters in the model to be trained; The unfrozen parameters in the model to be trained are adjusted based on the target loss data, and the step of training the model to be trained based on the data to be trained is repeated until the training termination condition is met, so as to obtain the target task training model corresponding to the new task.
2. The large model continuous learning method according to claim 1, characterized in that: The weighted processing of the first loss data according to the preset loss data to obtain the target loss data of the model to be trained includes: Obtaining preset loss data and its corresponding weighting coefficient, wherein the weighting coefficient decays linearly or nonlinearly with the increase of the number of training steps; Determining the product of the preset loss data and the weighting coefficient as second loss data; determining a difference between the first loss data and the second loss data as a loss difference; The loss difference value raised to the power of a natural number is determined as target loss data.
3. The large model continuous learning method according to claim 1, characterized in that: The training model to be trained is trained based on the training data corresponding to the new task to obtain the first loss data including: Inputting the input data into the model to be trained respectively to perform forward reasoning to obtain the first output data; First loss data is determined according to the first output data and the expected output data.
4. The large model continuous learning method according to claim 1, characterized in that: The frozen parameters in the model to be trained are all parameters in the preset training model, and the unfrozen parameters in the model to be trained are newly added fine-tuning parameters.
5. The large model continuous learning method according to claim 4 is characterized in that: Obtaining preset loss data includes: Freeze all parameters in the preset training model; The pre-trained model is trained based on the to-be-trained data corresponding to the new task to obtain the preset loss data.
6. The large model continuous learning method according to claim 5, characterized in that: Training the pre-trained model based on the to-be-trained data corresponding to the new task to obtain the preset loss data includes: Input the input data into the preset training model for forward reasoning to obtain the preset output data; Preset loss data is determined according to the preset output data and the expected output data.
7. The large model continuous learning method according to claim 1, characterized in that: Also includes: After each preset number of training steps, the negative value of the current preset loss data is used as the target loss data of the model to be trained.
8. A large model continuous learning device, characterized in that: include: An acquisition module, used for acquiring data to be trained and a model to be trained corresponding to a new task, wherein some parameters of the model to be trained are frozen, and the data to be trained includes input data and expected output data; A training module, used for training the to-be-trained model based on the training data corresponding to the new task to obtain first loss data; A data processing module, configured to perform weighted processing on the first loss data according to preset loss data to obtain target loss data of the model to be trained, wherein the preset loss data is loss data generated by freezing parameters in the model to be trained; The training module is also used to adjust the unfrozen parameters in the model to be trained based on the target loss data, repeat the step of training the model to be trained based on the data to be trained until the training termination condition is met, and obtain the target task training model corresponding to the new task.
9. An electronic device, characterized in that: The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the large model continuous learning method as described in any one of claims 1 to 7 is realized.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the large model continuous learning method as described in any one of claims 1 to 7.