Task performance control method and device, equipment, storage medium and program product
By determining key input words and neurons through causal tracing and adjusting activation levels, the problem of optimizing the performance of large language models in multi-task environments is solved, and effective control of task performance is achieved.
Patent Information
- Application Number
- CN202411278711.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-09-12
AI Technical Summary
When facing different natural language processing tasks, large language models have large differences in knowledge and capabilities, making it difficult to optimize the model's performance in a multi-task environment.
The causal tracing method is used to determine the key input words of the target model in performing contextual learning tasks, calculate the gradient changes, determine the key neurons, and control the task performance by adjusting the activation level of the neurons.
The task performance control of the target model in different tasks is achieved, and the performance of the model in a multi-task environment is improved.
Smart Images

Figure CN119415223B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a task performance control method, device, equipment, storage medium and program product. Background Art
[0002] With the rapid development of large-scale language model technology, various natural language processing tasks have demonstrated the remarkable capabilities of large-scale language models. These large-scale language models have achieved remarkable results in many fields, including translation, text generation, and question-answering systems.
[0003] In existing technologies, there are huge differences between the knowledge and capabilities required by large language models when facing different tasks. In order to optimize and improve the performance of the model in a multi-task environment, it is particularly important to control the processing performance of the same model when handling different tasks. Summary of the Invention
[0004] The present invention provides a task performance control method, device, equipment, storage medium and program product for controlling the processing performance of the same model when processing different tasks.
[0005] The present invention provides a task performance control method, comprising: determining the key input words of a target model in executing a context learning task based on a causal tracing method; determining the gradient change of the target model on the key input words during the reasoning process; determining neurons at a preset number of positions with the largest gradient changes as key neurons; and controlling the task performance of the target model by adjusting the activation levels of the key neurons.
[0006] According to a task performance control method provided by the present invention, the causal tracing method is used to determine the key input words of the target model in executing the context learning task, including: adding noise to each input word of the feedforward neural network layer to obtain a fluorescence vector; determining the predicted probability of the correct label corresponding to the fluorescence vector in executing the context learning task to obtain a perturbation matrix; and determining the key input word based on the perturbation matrix.
[0007] According to a task performance control method provided by the present invention, determining the gradient change of the target model on the key input word during the reasoning process includes: calculating the loss function when the key input word is predicted, and determining the gradient change based on the loss function.
[0008] According to a task performance control method provided by the present invention, the task performance of the target model is controlled by adjusting the activation level of the key neuron, including: amplifying the activation value of the key neuron by increasing the activation level; or suppressing the activation value of the key neuron by reducing the activation level.
[0009] The context learning task includes at least one of a question and answer task, a sentiment analysis task, a question understanding task, a text classification task, a legal text classification task, a cause and effect classification task, a sentiment classification task, and a text matching task.
[0010] The application further provides a task performance control device, comprising a determination module and an adjustment module; the determination module is used for determining a key input word of a target model in performing a context learning task based on a cause and effect tracking method; determining a gradient change of the target model on the key input word in performing an inference process; determining neurons at a preset number of positions with the largest gradient change as key neurons; and the adjustment module is used for controlling a task performance of the target model by adjusting an activation level of the key neurons.
[0011] The determination module is used for adding noise on each input word of a feedforward neural network layer to obtain a fluorescence vector; determining a prediction probability of the fluorescence vector on a corresponding correct label in performing a context learning task to obtain a perturbation matrix; and determining the key input word according to the perturbation matrix.
[0012] The determination module is used for calculating a loss function when the key input word is predicted, and determining a gradient change according to the loss function.
[0013] The adjustment module is used for amplifying an activation value of the key neurons by increasing an activation level, or inhibiting the activation value of the key neurons by reducing the activation level.
[0014] The context learning task includes at least one of a question and answer task, a sentiment analysis task, a question understanding task, a text classification task, a legal text classification task, a cause and effect classification task, a sentiment classification task, and a text matching task.
[0015] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor implements the task performance control method according to any one of the above-mentioned task performance control methods when executing the program.
[0016] The application further provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the task performance control method according to any one of the above-mentioned task performance control methods.
[0017] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned task performance control methods.
[0018] The task performance control method, apparatus, device, storage medium, and program product provided by the present invention can determine the key input words of a target model when executing a contextual learning task based on a causal tracing method; determine the gradient change of the target model on the key input words during the reasoning process; determine neurons at a preset number of locations with the largest gradient changes as key neurons; and control the task performance of the target model by adjusting the activation levels of the key neurons. Through this solution, since neurons at a preset number of locations with the largest gradient changes can be determined as key neurons and the task performance of the target model can be controlled by adjusting the activation levels of the key neurons, it is possible to determine the important neurons of the target model when processing different tasks, thereby achieving control of the target model's task performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 is a flow chart of the task performance control method provided by the present invention;
[0021] Figure 2 It is a structural diagram of the task performance control device provided by the present invention;
[0022] Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0023] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0024] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being more preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0025] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0026] In order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order.
[0027] The embodiments of the present application describe some exemplary embodiments for the purpose of explanation. It should be understood that the present application can be implemented in other ways that are not specifically shown in the drawings.
[0028] like Figure 1 As shown, the embodiment of the present application provides a task performance control method, which can be applied to a task performance control device. The task performance control method may include S101-S104:
[0029] S101. The task performance control device determines the key input words of the target model in executing the context learning task based on the causal tracing method.
[0030] Optionally, the context learning task includes at least one of the following: a question answering task, a sentiment analysis task, a question understanding task, a text classification task, a legal text classification task, a causal relationship classification task, a sentiment classification task, and a text matching task.
[0031] Specifically, the question-answering (QA) task can generate answers to SQuAD 1.1 questions based on documents; the sentiment analysis (SA) task refers to sentiment classification of English tweets on social media, judging the sentiment as positive or negative; the question understanding (QU) task refers to determining whether a query clarification is correct by answering "yes" or "no" in a dialogue; the text classification (TC) task refers to classifying the topic of an English news article into one of four categories; the legal text classification (LTC) task refers to classifying English sentences in the legal field as overturning or non-overturning; the cause-effect classification (CEC) task refers to determining whether the second sentence is logically derived from the first sentence in common sense reasoning; the emotion classification (EC) task refers to classifying the emotion in a post into one of six categories, namely sadness, joy, love, anger, fear or surprise, for sentiment classification in social media; and the text matching (TM) task refers to classifying question pairs in the medical and health fields into two categories.
[0032] It should be noted that for each task, the data can be divided into two parts: one half is used for training to identify key neurons, and the other half is used for testing to evaluate the performance of the target model. The LLama2-7b-chat model and 5-shot ICL can be used for reasoning. The total number of model neurons is calculated by multiplying 32 (the number of model layers) by 11,008 (the size of the hidden state).
[0033] Optionally, the task performance control device determines the key input word of the target model in performing the context learning task based on a cause-effect tracking method, including: adding noise on each input word of the feedforward neural network layer to obtain a fluorescence vector; determining the prediction probability of the fluorescence vector on the corresponding correct label in performing the context learning task to obtain a perturbation matrix; and determining the key input word according to the perturbation matrix.
[0034] Specifically, the task performance control device can add noise on each input word of each feedforward neural network (FFN) layer as a fluorescence vector, and record the prediction probability of the correct label to construct a perturbation matrix A. The rows and columns of the matrix correspond to the index of the FFN layer and the length of the input, respectively. Finally, the input word corresponding to the prediction probability with a larger perturbation degree in the perturbation matrix is selected as the key input word.
[0035] ;
[0036] wherein, represents the prediction after adding noise on the i-th input word of the j-th layer, is the fluorescence vector, and .
[0037] S102: The task performance control device determines the gradient change of the target model on the key input word during the reasoning process.
[0038] Optionally, the task performance control device determines the gradient change of the target model on the key input word during the reasoning process, including: calculating the loss function when the key input word is predicted, and determining the gradient change based on the loss function.
[0039] Specifically, the task performance control device can use the data of a specific task to perform a forward pass, that is, calculate the loss function when a special tag is predicted, as shown below:
[0040] ;
[0041] Among them, s represents the position of the key input word, Represents the focused input set.
[0042] Next, the gradient change is calculated on the training set for the specific task, and the gradient of the gated weights is determined based on the loss value as follows:
[0043] ;
[0044] in, , l represents the number of layers, and d represents the dimension of each layer. Since the size of the FFN gate value of each layer is 4d-dimensional, the gradient change is compressed to d-dimensional to obtain the same Matrices of the same size.
[0045] S103: The task performance control device determines neurons at a preset number of positions with the largest gradient changes as key neurons.
[0046] Optionally, the task performance control device can select n positions with the largest global changes as key neurons, which are considered to be crucial for the task. A single neuron can be represented as .
[0047] S104. The task performance control device controls the task performance of the target model by adjusting the activation level of the key neurons.
[0048] Optionally, the task performance control device controls the task performance of the target model by adjusting the activation level of the key neuron, including: amplifying the activation value of the key neuron by increasing the activation level; or, suppressing the activation value of the key neuron by reducing the activation level.
[0049] Specifically, using these identified key neurons, we can further manipulate them to influence the overall performance of the target model. Specifically, we can amplify the activation level of key neurons by increasing them, or suppress their activation value by decreasing them. The operation is as follows:
[0050] ;
[0051] in, Represents the selected neuron in layer l. The value of this neuron is set to less than 1 during inhibition, greater than 1 during amplification, and equal to 1 during normal inference.
[0052] It should be noted that this application eliminates redundant markers, minimizes the interference of unnecessary neurons, and identifies task-specific neurons by focusing on the most important markers in the task processing process. Compared with traditional neuron localization methods, this application can more effectively identify task-specific neurons. This application also conducted experiments on eight different tasks. The experiment proved that this application can accurately locate task-specific neurons by inhibiting and enhancing the identified neurons.
[0053] In an embodiment of the present application, since neurons at a preset number of positions with the largest gradient changes can be determined as key neurons, and the task performance of the target model can be controlled by adjusting the activation level of the key neurons, the important neurons of the target model when processing different tasks can be determined, thereby achieving control of the task performance of the target model.
[0054] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of method. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0055] The task performance control method provided in the embodiment of the present application can be executed by a task performance control device or a control module for task performance control in the task performance control device. In the embodiment of the present application, the task performance control device provided in the embodiment of the present application is described by taking the task performance control method executed by the task performance control device as an example.
[0056] It should be noted that the embodiment of the present application can divide the task performance control device into functional modules according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of software functional modules. Optionally, the division of modules in the embodiment of the present application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.
[0057] like Figure 2 As shown, an embodiment of the present application provides a task performance control device 200. The task performance control device 200 includes: a determination module 201 and an adjustment module 202; the determination module 201 is used to determine the key input words of the target model in executing the context learning task based on the causal tracing method; determine the gradient change of the target model on the key input words during the reasoning process; determine the neurons at a preset number of positions with the largest gradient change as key neurons; and the adjustment module 202 is used to control the task performance of the target model by adjusting the activation level of the key neurons.
[0058] Optionally, the determination module 201 is used to add noise to each input word of the feedforward neural network layer to obtain a fluorescence vector; determine the predicted probability of the correct label corresponding to the fluorescence vector in executing the context learning task to obtain a perturbation matrix; and determine the key input word based on the perturbation matrix.
[0059] Optionally, the determination module 201 is configured to calculate a loss function when the key input word is predicted, and determine a gradient change according to the loss function.
[0060] Optionally, the adjustment module 202 is configured to amplify the activation value of the key neuron by increasing the activation level; or to suppress the activation value of the key neuron by decreasing the activation level.
[0061] Optionally, the context learning task includes at least one of the following: a question answering task, a sentiment analysis task, a question understanding task, a text classification task, a legal text classification task, a causal relationship classification task, a sentiment classification task, and a text matching task.
[0062] In an embodiment of the present application, since neurons at a preset number of positions with the largest gradient changes can be determined as key neurons, and the task performance of the target model can be controlled by adjusting the activation level of the key neurons, the important neurons of the target model when processing different tasks can be determined, thereby achieving control of the task performance of the target model.
[0063] Figure 3An example of a physical structure diagram of an electronic device is shown below. Figure 3 As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communications bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other via the communications bus 340. The processor 310 may call logic instructions in the memory 330 to execute a task performance control method, which includes: determining the key input words of a target model in executing a context learning task based on a causal tracing method; determining the gradient change of the target model on the key input words during the reasoning process; determining neurons at a preset number of positions with the largest gradient change as key neurons; and controlling the task performance of the target model by adjusting the activation level of the key neurons.
[0064] Furthermore, the logic instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0065] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the task performance control method provided by the above methods, which includes: determining the key input words of the target model in executing the context learning task based on the causal tracing method; determining the gradient change of the target model on the key input words during the reasoning process; determining the neurons at a preset number of positions with the largest gradient change as key neurons; and controlling the task performance of the target model by adjusting the activation level of the key neurons.
[0066] In yet another aspect, the present application also provides a non-transitory computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements a task performance control method as provided by any of the above methods, the method comprising: determining, based on a cause-effect tracking method, a focus input word of a target model in performing a context learning task; determining a gradient change of the target model on the focus input word in performing an inference process; determining neurons at a preset number of positions with the largest gradient change as focus neurons; and controlling a task performance of the target model by adjusting an activation level of the focus neurons.
[0067] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0068] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary universal hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0069] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A task performance control method, characterized in that: include: Determine the key input words of the target model in performing contextual learning tasks based on the causal tracing method; Determining a gradient change of the target model on the key input word during the inference process; The neurons at the preset number of positions with the largest gradient changes are determined as key neurons; The task performance of the target model is controlled by adjusting the activation level of the focus neurons.
2. The task performance control method according to claim 1, characterized in that: The method of determining the key input words of the target model in performing the context learning task based on the causal tracing method includes: Add noise to each input word of the feedforward neural network layer to obtain the fluorescence vector; Determining the predicted probability of the correct label corresponding to the fluorescence vector in performing the context learning task to obtain a perturbation matrix; The key input word is determined according to the disturbance matrix.
3. The task performance control method according to claim 1, characterized in that: Determining the gradient change of the target model on the key input word during the reasoning process includes: A loss function is calculated when the key input word is predicted, and a gradient change is determined according to the loss function.
4. The task performance control method according to claim 1, characterized in that: The step of controlling the task performance of the target model by adjusting the activation level of the key neurons includes: The activation value of the key neuron is amplified by increasing the activation level; or the activation value of the key neuron is suppressed by decreasing the activation level.
5. The task performance control method according to any one of claims 1 to 4, characterized in that: The context learning task includes at least one of the following: a question answering task, a sentiment analysis task, a question understanding task, a text classification task, a legal text classification task, a causal relationship classification task, a sentiment classification task, and a text matching task.
6. A task performance control device, characterized in that: include: Determine modules and adjust modules; The determination module is used to determine the key input words of the target model in performing the context learning task based on the causal tracing method; Determining the gradient change of the target model on the key input word during the inference process; determining neurons at a preset number of positions with the largest gradient changes as key neurons; The adjustment module is used to control the task performance of the target model by adjusting the activation level of the key neurons.
7. The task performance control device according to claim 6, characterized in that: The determination module is used to add noise to each input word of the feedforward neural network layer to obtain a fluorescence vector; determine the predicted probability of the correct label corresponding to the fluorescence vector in performing the context learning task to obtain a perturbation matrix; and determine the key input word based on the perturbation matrix.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the task performance control method according to any one of claims 1 to 5 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the task performance control method according to any one of claims 1 to 5 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the task performance control method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Data processing method and device based on artificial intelligence, equipment and medium
CN112116912A
Selective activation-based pulse neural network continuous learning target recognition system
CN117710789A