Training Method and Device for Self-Attention Model

By introducing linear transformation units and calculating regular term update loss function in the self-attention model, the problem of large changes in model structure in the prior art is solved, and the generalization ability and workload are improved.

CN115186757BActive Publication Date: 2025-07-11BEIJING LONGZHI DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210848602.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-07-11
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

The method of improving the generalization ability of self-attention models in the prior art requires changes to the model structure, resulting in a large workload.

Method used

By introducing a first linear transformation unit, a second linear transformation unit and a third linear transformation unit into the self-attention model, the matrix of the attention head is calculated, and the regular terms are calculated based on these matrices, and the loss function is updated to train the model.

Benefits of technology

Improve the generalization ability of the self-attention model without modifying the model structure, reducing the workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115186757B_ABST
    Figure CN115186757B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of artificial intelligence technologies, and provides a training method and device for a self-attention model. The method includes: obtaining target data, and respectively inputting the target data into a first linear transformation unit, a second linear transformation unit, and a third linear transformation unit in each attention head, so that the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head respectively output a first matrix, a second matrix, and a third matrix; calculating a fourth matrix corresponding to each attention head based on the first matrix, the second matrix, and the third matrix corresponding to each attention head; calculating any one of the following regularization terms based on the fourth matrix corresponding to each attention head: a first regularization term, a second regularization term, and a third regularization term; updating a loss function corresponding to the self-attention model based on any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term; and training the self-attention model by using the updated loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a training method and device for a self-attention model. Background Art

[0002] To improve the generalization ability of the self-attention model (the generalization ability refers to the processing effect of the self-attention model on data that has not been encountered before), the self-attention model is often improved. For example, different structural restrictions are imposed on the self-attention model. For example, local attention + sliding window mechanism is introduced into the self-attention model, so that the information extraction of a single character or word by the self-attention model only involves several adjacent characters or words, thereby reducing the amount of computation and improving the data processing ability of the self-attention model. This method makes great changes to the model structure of the self-attention model and requires a huge amount of work.

[0003] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following technical problems in the related art: the current method for improving the generalization ability of the self-attention model needs to modify the model structure of the self-attention model, resulting in a large amount of work. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure provide a training method and device for a self-attention model, so as to solve the problem in the prior art that the current method for improving the generalization ability of the self-attention model needs to modify the model structure of the self-attention model, resulting in a large amount of work.

[0005] In a first aspect of the embodiments of the present disclosure, a training method for a self-attention model is provided, including: obtaining target data, and inputting the target data into the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head respectively, so that the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head output a first matrix, a second matrix, and a third matrix respectively, where the self-attention model includes multiple self-attention network layers, and each self-attention network layer includes multiple attention heads; calculating a fourth matrix corresponding to each attention head based on the first matrix, the second matrix, and the third matrix corresponding to each attention head; calculating any one of the following regularization terms based on the fourth matrix corresponding to each attention head: a first regularization term, a second regularization term, and a third regularization term; updating the loss function corresponding to the self-attention model based on any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term; and training the self-attention model using the updated loss function.

[0006] In a second aspect of the embodiments of the present disclosure, a training apparatus for a self-attention model is provided, including: an acquisition module configured to acquire target data and input the target data into a first linear transformation unit, a second linear transformation unit, and a third linear transformation unit in each attention head respectively, so that the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head output a first matrix, a second matrix, and a third matrix respectively, where the self-attention model includes multiple self-attention network layers, and each self-attention network layer includes multiple attention heads; a first calculation module configured to calculate a fourth matrix corresponding to each attention head based on the first matrix, the second matrix, and the third matrix corresponding to each attention head; a second calculation module configured to calculate any one of the following regularization terms: a first regularization term, a second regularization term, and a third regularization term based on the fourth matrix corresponding to each attention head; an update module configured to update the loss function corresponding to the self-attention model based on any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term; a training module configured to train the self-attention model by using the updated loss function.

[0007] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the above method are implemented.

[0008] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0009] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: Since the embodiments of the present disclosure obtain target data and input the target data into the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head respectively, so that the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head output the first matrix, the second matrix, and the third matrix respectively. Among them, the self-attention model includes multiple self-attention network layers, and each self-attention network layer includes multiple attention heads; based on the first matrix, the second matrix, and the third matrix corresponding to each attention head, calculate the fourth matrix corresponding to each attention head; based on the fourth matrix corresponding to each attention head, calculate any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term; update the loss function corresponding to the self-attention model based on any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term; train the self-attention model using the updated loss function. Therefore, by adopting the above technical means, it is possible to solve the problem in the prior art that the existing methods for improving the generalization ability of the self-attention model require modifications to the model structure of the self-attention model, resulting in a large workload, and further reduce the workload while improving the generalization ability of the self-attention model. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the prior art descriptions. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0011] Figure 1 is a schematic diagram of the application scenario of the embodiments of the present disclosure;

[0012] Figure 2 is a schematic flowchart of a method for training a self-attention model provided by the embodiments of the present disclosure;

[0013] Figure 3 is a schematic structural diagram of a device for training a self-attention model provided by the embodiments of the present disclosure;

[0014] Figure 4 is a schematic structural diagram of an electronic device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0016] A training method and apparatus for a self-attention model according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0017] Figure 1 It is a schematic diagram of the application scenario of the embodiment of the present disclosure. The application scenario may include terminal devices 101, 102, and 103, a server 104, and a network 105.

[0018] The terminal devices 101, 102, and 103 can be either hardware or software. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with a display screen and supporting communication with the server 104, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.; when the terminal devices 101, 102, and 103 are software, they can be installed in the above-mentioned electronic devices. The terminal devices 101, 102, and 103 can be implemented as multiple software or software modules, or can also be implemented as a single software or software module, and the embodiments of the present disclosure do not limit this. Further, various applications can be installed on the terminal devices 101, 102, and 103, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.

[0019] The server 104 can be a server providing various services. For example, it is a background server that receives requests sent by terminal devices with which it establishes a communication connection. The background server can receive and analyze requests sent by terminal devices and generate processing results. The server 104 can be a single server, or can also be a server cluster composed of several servers, or can also be a cloud computing service center, and the embodiments of the present disclosure do not limit this.

[0020] It should be noted that the server 104 can be either hardware or software. When the server 104 is hardware, it can be various electronic devices providing various services for the terminal devices 101, 102, and 103. When the server 104 is software, it can be multiple software or software modules providing various services for the terminal devices 101, 102, and 103, or can also be a single software or software module providing various services for the terminal devices 101, 102, and 103, and the embodiments of the present disclosure do not limit this.

[0021] The network 105 can be a wired network connected by coaxial cables, twisted pairs, and optical fibers, or a wireless network that can interconnect various communication devices without wiring. For example, Bluetooth, Near Field Communication (NFC), Infrared, etc. The embodiments of the present disclosure do not limit this.

[0022] Users can establish a communication connection with the server 104 via the network 105 through the terminal devices 101, 102, and 103 to receive or send information, etc. It should be noted that the specific types, quantities, and combinations of the terminal devices 101, 102, and 103, the server 104, and the network 105 can be adjusted according to the actual needs of the application scenario. The embodiments of the present disclosure do not limit this.

[0023] Figure 2 It is a schematic flowchart of a method for training a self-attention model provided by an embodiment of the present disclosure. Figure 2 The method for training the self-attention model can be executed by Figure 1 the terminal device or the server of Figure 2 As shown in

[0024] S201, obtain target data, and input the target data into the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head respectively, so that the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head output a first matrix, a second matrix, and a third matrix respectively. Among them, the self-attention model includes multiple self-attention network layers, and each self-attention network layer includes multiple attention heads;

[0025] S202, calculate a fourth matrix corresponding to each attention head based on the first matrix, the second matrix, and the third matrix corresponding to each attention head;

[0026] S203, calculate any one of the following regularization terms based on the fourth matrix corresponding to each attention head: the first regularization term, the second regularization term, and the third regularization term;

[0027] S204, update the loss function corresponding to the self-attention model based on any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term;

[0028] S205, train the self-attention model using the updated loss function.

[0029] The self-attention model can be a neural network architecture such as a Transformer. The self-attention model includes multiple self-attention network layers connected in series. Each self-attention network layer includes multiple attention heads connected in parallel. Each attention head includes a first linear transformation unit, a second linear transformation unit, and a third linear transformation unit. The linear transformation unit is used to provide a linear transformation. The target data can be information such as pictures and texts. In the embodiments of the present disclosure, the original loss function of the self-attention model is corrected or reconstructed by one of the first regularization term, the second regularization term, and the third regularization term, that is, the loss function corresponding to the self-attention model is updated based on the regularization term. Through the above technical solutions, the generalization ability of the self-attention model can be improved without modifying the structure of the self-attention model.

[0030] The embodiments of the present disclosure can be applied to any scenario that requires improving the generalization ability of the self-attention model. For example, in the face recognition scenario, the target data is a face picture. For example, in the text recognition scenario, the target data is text.

[0031] According to the technical solution provided by the embodiments of the present disclosure, the target data is obtained, and the target data is respectively input into the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head, so that the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head respectively output a first matrix, a second matrix, and a third matrix. Among them, the self-attention model includes multiple self-attention network layers, and each self-attention network layer includes multiple attention heads; based on the first matrix, the second matrix, and the third matrix corresponding to each attention head, a fourth matrix corresponding to each attention head is calculated; based on the fourth matrix corresponding to each attention head, any one of the following regularization terms is calculated: the first regularization term, the second regularization term, and the third regularization term; the loss function corresponding to the self-attention model is updated based on any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term; the self-attention model is trained using the updated loss function. Therefore, by adopting the above technical means, it is possible to solve the problem in the prior art that the method for improving the generalization ability of the self-attention model currently requires modifying the model structure of the self-attention model, resulting in a large workload, and thus reducing the workload while improving the generalization ability of the self-attention model.

[0032] In step S202, based on the first matrix, second matrix, and third matrix corresponding to each attention head, calculate the fourth matrix corresponding to each attention head, including: performing a vector inner product operation on the vectors in the first matrix and the vectors in the second matrix corresponding to each attention head according to the positions of the vectors to obtain the target vector corresponding to each attention head; performing a vector multiplication operation on the data in the target vector corresponding to each attention head and the vectors in the third matrix according to the positions of the data in the target vector and the vectors in the third matrix to obtain the fourth matrix corresponding to each attention head.

[0033] The dimensionality data of the first matrix, second matrix, and third matrix are the same, for example, all are of dimension (N, M). For each attention head: The first vector in the first matrix and the first vector in the second matrix perform a vector inner product operation to obtain the first data (the first vector can be a column vector or a row vector, taking the row vector as an example); The second vector in the first matrix and the second vector in the second matrix perform a vector inner product operation to obtain the second data... The Nth vector (i.e., the last vector) in the first matrix and the Nth vector in the second matrix perform a vector inner product operation to obtain the Nth data. Then the first data, second data... Nth data are combined to form a row vector, which is the target vector. The third matrix has N row vectors, and the target vector has N data. The first data in the target vector and the first vector in the third matrix perform a vector multiplication operation to obtain the first vector of the fourth matrix; The second data in the target vector and the second vector in the third matrix perform a vector multiplication operation to obtain the second vector of the fourth matrix... The Nth data in the target vector and the Nth vector in the third matrix perform a vector multiplication operation to obtain the Nth vector of the fourth matrix. Finally, the fourth matrix is obtained.

[0034] Optionally, performing a vector multiplication operation on the data in the target vector corresponding to each attention head and the vectors in the third matrix according to the positions of the data in the target vector and the vectors in the third matrix to obtain the fourth matrix corresponding to each attention head, including: processing the target vector corresponding to each attention head using the softmax function (this step is a normalization process, and finally the sum of the data in the target vector corresponding to each attention head is 1); performing a vector multiplication operation on the data in the target vector corresponding to each attention head after being processed by the softmax function and the vectors in the third matrix according to the positions of the data in the target vector and the vectors in the third matrix to obtain the fourth matrix corresponding to each attention head

[0035] In step S203, based on the fourth matrix corresponding to each attention head, calculate any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term, including: calculating the first regularization term based on the fourth matrix corresponding to each attention head: performing an inner product operation on the fourth matrices corresponding to the attention heads with the same serial number in every two adjacent self-attention network layers in the self-attention model to obtain multiple first calculation results; summing up the multiple first calculation results to obtain the first regularization term.

[0036] For example, if the self-attention model has N self-attention network layers and each self-attention network layer has H attention heads, the first regularization term R1 can be calculated by the following formula:

[0037]

[0038] where i represents the serial number of the attention head in each self-attention network layer, starting from 0, and the maximum value of i is H - 1; j represents the serial number of the self-attention network layer in the self-attention model, starting from 0, and the maximum value of j is N - 1; represents the i-th attention head in the j-th self-attention network layer.

[0039] Illustrative example: If the self-attention model has 3 self-attention network layers and each self-attention network layer has 3 attention heads, then N is 3 and H is 3. The calculation process is as follows: Perform an inner product operation on the fourth matrix of the 0th attention head in the 0th self-attention network layer and the fourth matrix of the 0th attention head in the 1st self-attention network layer, and perform an inner product operation on the fourth matrix of the 0th attention head in the 1st self-attention network layer and the fourth matrix of the 0th attention head in the 2nd self-attention network layer. Add up the two inner product results of the 0th attention head (this time belongs to the content of the first summation formula, and the calculation of the two inner product results of the 1st attention head and the 2nd attention head and the 0th attention head is the same); Add up the two inner product results of the 0th attention head, the two inner product results of the 1st attention head, and the two inner product results of the 2nd attention head (this time belongs to the content of the second summation formula).

[0040] In the previous embodiment, the multiple first calculation results are summed up to obtain the first regularization term, where one first calculation result is equivalent to one inner product result (in the example given, there are 6 inner product results, which can be understood as summing up 6 inner product results at one time, or first performing the first summation on the two inner product results of each attention head, and then performing the second summation on the summation results corresponding to the three attention heads).

[0041] In step S203, based on the fourth matrix corresponding to each attention head, calculate any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term, including: calculating the second regularization term based on the fourth matrix corresponding to each attention head: according to the serial number of the self-attention network layer in the self-attention model, perform matrix inner product operations on the fourth matrices corresponding to the attention heads with the same serial number in every two self-attention network layers in the self-attention model in the reverse multiplication order to obtain multiple second calculation results; sum up the multiple second calculation results to obtain the second regularization term.

[0042] For example, if the self-attention model has N self-attention network layers and each self-attention network layer has H attention heads, the second regularization term R2 can be calculated by the following formula:

[0043]

[0044] Illustrative example: If the self-attention model has 4 self-attention network layers and each self-attention network layer has 3 attention heads, then N is 4 and H is 3. The calculation process is as follows: Perform matrix inner product operation on the 0th attention head of the 0th self-attention network layer and the 0th attention head of the 2nd self-attention network layer, and perform matrix inner product operation on the 0th attention head of the 1st self-attention network layer and the 0th attention head of the 3rd self-attention network layer, and add up the two inner product results of the 0th attention head (this time it belongs to the content of the first summation formula, and the calculation of the two inner product results of the 1st attention head, the 2nd attention head, and the 0th attention head is the same); add up the two inner product results of the 0th attention head, the two inner product results of the 1st attention head, and the two inner product results of the 2nd attention head (this time it belongs to the content of the second summation formula).

[0045] In the previous embodiment, sum up the multiple second calculation results to obtain the second regularization term, where one second calculation result is equivalent to one inner product result (in the example given, there are 6 inner product results, which can be understood as summing up 6 inner product results at one time, or as first performing the first summation on the two inner product results of each attention head, and then performing the second summation on the summation results corresponding to the three attention heads).

[0046] In step S203, based on the fourth matrix corresponding to each attention head, calculate any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term, including: calculating the third regularization term based on the fourth matrix corresponding to each attention head: perform multiple matrix inner product operations on the fourth matrices corresponding to multiple attention heads in each self-attention network layer in the self-attention model to obtain multiple third calculation results, where the number of attention heads in each self-attention network layer is equal to the number of times of matrix inner product operations; sum up the multiple third calculation results to obtain the third regularization term.

[0047] For example, if the self-attention model has N self-attention network layers, and each self-attention network layer has H attention heads, the third regularization term R3 can be calculated by the following formula:

[0048]

[0049] For example: If the self-attention model has 3 self-attention network layers, and each self-attention network layer has 3 attention heads, then N is 3 and H is 3. The calculation process is as follows: For the 0th self-attention network layer: Calculate the matrix inner product operation between the 0th attention head and the 1st attention head, and the matrix inner product operation between the 1st attention head and the 2nd attention head, and add the two inner product results of the 0th self-attention network layer (this belongs to the content of the first summation formula, and the calculation of the two inner product results of the 1st self-attention network layer and the 2nd self-attention network layer and the 0th self-attention network layer is the same); Add the two inner product results of the 0th self-attention network layer, the two inner product results of the 1st self-attention network layer, and the two inner product results of the 2nd self-attention network layer (this belongs to the content of the second summation formula).

[0050] In the previous embodiment, the third regularization term is obtained by summing multiple third calculation results. Among them, one third calculation result is equivalent to one inner product result (in the example given, there are 6 inner product results, which can be understood as summing 6 inner product results at one time, or as first summing the two inner product results of each self-attention network layer and then summing the corresponding summation results of the three self-attention network layers for the second time).

[0051] In step S204, update the loss function corresponding to the self-attention model based on any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term, including: taking the loss function corresponding to the self-attention model obtained before the update as the original loss function; performing a weighted summation operation on the original loss function and the first regularization term according to a preset ratio to obtain the loss function corresponding to the updated self-attention model; or performing a weighted summation operation on the original loss function and the second regularization term according to a preset ratio to obtain the loss function corresponding to the updated self-attention model; or performing a weighted summation operation on the original loss function and the third regularization term according to a preset ratio to obtain the loss function corresponding to the updated self-attention model.

[0052] Update the loss function corresponding to the self-attention model according to the following formula:

[0053] loss = (1 - α)·suploss + α·regulation

[0054] loss is the loss function corresponding to the updated self-attention model; suploss is the original loss function, that is, the loss function corresponding to the self-attention model before update; regulation is one of the first regularization term, the second regularization term, and the third regularization term; α is the partition coefficient, that is, the preset ratio, used to adjust the model training accuracy, and empirically can be taken between 0.1 and 0.2, and this value may need to be adjusted according to different tasks.

[0055] In step S205, training the self-attention model using the updated loss function includes: obtaining training data, where the training data includes target data; calculating the loss value corresponding to the training data using the updated loss function; and updating the model parameters of the self-attention model based on the loss value.

[0056] For example, in the face recognition scenario, the training data is face images; for example, in the text recognition scenario, the training data is text; the target data can be part or all of the training data. The loss function corresponds to the self-attention model, and different types of self-attention models have different loss functions. Updating the model parameters of the self-attention model based on the loss value is based on the principle of minimizing the loss value, and adjusting the model parameters of the self-attention model. That is to say, after updating the model parameters of the self-attention model, calculate the loss value corresponding to the training data using the updated loss function again. At this time, a minimum loss value can be obtained (compared with before updating the model parameters of the self-attention model and when the model parameters of the self-attention model are other values). After updating, the numerical value of the model parameters of the self-attention model is optimal.

[0057] All of the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated here one by one.

[0058] The following is an embodiment of the apparatus of the present disclosure, which can be used to execute the embodiment of the method of the present disclosure. For details not disclosed in the embodiment of the apparatus of the present disclosure, please refer to the embodiment of the method of the present disclosure.

[0059] Figure 3 is a schematic diagram of a training apparatus for a self-attention model provided by an embodiment of the present disclosure. As Figure 3 shown, the training apparatus for the self-attention model includes:

[0060] An obtaining module 301, configured to obtain target data, and input the target data into the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head respectively, so that the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head output a first matrix, a second matrix, and a third matrix respectively, where the self-attention model includes multiple self-attention network layers, and each self-attention network layer includes multiple attention heads;

[0061] The first calculation module 302 is configured to calculate a fourth matrix corresponding to each attention head based on the first matrix, the second matrix, and the third matrix corresponding to each attention head;

[0062] The second calculation module 303 is configured to calculate any one of the following regularization terms based on the fourth matrix corresponding to each attention head: the first regularization term, the second regularization term, and the third regularization term;

[0063] The update module 304 is configured to update the loss function corresponding to the self-attention model based on any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term;

[0064] The training module 305 is configured to train the self-attention model using the updated loss function.

[0065] The self-attention model may be a neural network architecture such as Transformer. The self-attention model includes multiple self-attention network layers connected in series. Each self-attention network layer includes multiple attention heads connected in parallel. Each attention head includes a first linear transformation unit, a second linear transformation unit, and a third linear transformation unit. The linear transformation unit is used to provide a linear transformation. The target data may be information such as pictures and texts. In the embodiments of the present disclosure, the original loss function of the self-attention model is corrected or reconstructed by one of the first regularization term, the second regularization term, and the third regularization term, that is, the loss function corresponding to the self-attention model is updated based on the regularization term. Through the above technical solutions, the generalization ability of the self-attention model can be improved without modifying the structure of the self-attention model.

[0066] The embodiments of the present disclosure can be applied to any scenario that requires improving the generalization ability of the self-attention model. For example, in the face recognition scenario, the target data is a face picture, and in the text recognition scenario, the target data is text.

[0067] According to the technical solution provided by the embodiments of the present disclosure, target data is obtained, and the target data is respectively input into the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head, so that the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head respectively output a first matrix, a second matrix, and a third matrix. The self-attention model includes multiple self-attention network layers, and each self-attention network layer includes multiple attention heads; based on the first matrix, the second matrix, and the third matrix corresponding to each attention head, a fourth matrix corresponding to each attention head is calculated; based on the fourth matrix corresponding to each attention head, any one of the following regularization terms is calculated: a first regularization term, a second regularization term, and a third regularization term; based on any one of the following regularization terms, the loss function corresponding to the self-attention model is updated: the first regularization term, the second regularization term, and the third regularization term; the self-attention model is trained by using the updated loss function. Therefore, by adopting the above technical means, the problem in the prior art that the method for improving the generalization ability of the self-attention model at present requires modification of the model structure of the self-attention model and has a large workload can be solved, and thus, while improving the generalization ability of the self-attention model, the workload can be reduced.

[0068] Optionally, the first calculation module 302 is further configured to perform an inner product operation on the vectors in the first matrix and the vectors in the second matrix corresponding to each attention head according to the positions corresponding to the vectors, so as to obtain a target vector corresponding to each attention head; perform a multiplication operation on the data in the target vector corresponding to each attention head and the vectors in the third matrix according to the positions corresponding to the data in the target vector and the vectors in the third matrix, so as to obtain a fourth matrix corresponding to each attention head.

[0069] The dimension data of the first matrix, the second matrix, and the third matrix are the same, for example, all are of the dimension (N, M). For each attention head: The first vector in the first matrix and the first vector in the second matrix perform a vector inner product operation to obtain the first data (the first vector can be a column vector or a row vector, taking the row vector as an example); the second vector in the first matrix and the second vector in the second matrix perform a vector inner product operation to obtain the second data... The Nth vector (i.e., the last vector) in the first matrix and the Nth vector in the second matrix perform a vector inner product operation to obtain the Nth data. Then the first data, the second data... the Nth data are combined to form a row vector, which is the target vector. The third matrix has N row vectors, and the target vector has N data. The first data in the target vector and the first vector in the third matrix perform a vector multiplication operation to obtain the first vector of the fourth matrix; the second data in the target vector and the second vector in the third matrix perform a vector multiplication operation to obtain the second vector of the fourth matrix... The Nth data in the target vector and the Nth vector in the third matrix perform a vector multiplication operation to obtain the Nth vector of the fourth matrix. Finally, the fourth matrix is obtained.

[0070] Optionally, the first calculation module 302 is further configured to process the target vector corresponding to each attention head by using the softmax function (this step is a normalization process, and finally makes the sum of the data in the target vector corresponding to each attention head equal to one); perform a vector multiplication operation on the data in the target vector corresponding to each attention head after being processed by the softmax function and the vectors in the third matrix according to the corresponding positions of the data in the target vector and the vectors in the third matrix, to obtain the fourth matrix corresponding to each attention head

[0071] Optionally, the second calculation module 303 is further configured with a first regularization term, a second regularization term, and a third regularization term, including: calculating the first regularization term based on the fourth matrix corresponding to each attention head: performing a matrix inner product operation on the fourth matrices corresponding to the attention heads with the same serial number in every two adjacent self-attention network layers in the self-attention model to obtain a plurality of first calculation results; summing the plurality of first calculation results to obtain the first regularization term.

[0072] For example, if the self-attention model has N self-attention network layers, and each self-attention network layer has H attention heads, the first regularization term R1 can be calculated by the following formula:

[0073]

[0074] where, i represents the serial number of the attention head in each self-attention network layer, i starts from 0, and the maximum value of i is H - 1; j represents the serial number of the self-attention network layer in the self-attention model, j starts from 0, and the maximum value of j is N - 1; Represents the $i$-th attention head of the $j$-th self-attention network layer.

[0075] For example: The self-attention model has 3 self-attention network layers, and each self-attention network layer has 3 attention heads. Then $N$ is 3 and $H$ is 3. The calculation process is as follows: The 0-th attention head of the 0-th self-attention network layer performs a matrix inner product operation with the 0-th attention head of the 1-st self-attention network layer, and the 0-th attention head of the 1-st self-attention network layer performs a matrix inner product operation with the 0-th attention head of the 2-nd self-attention network layer. Add up the two inner product results of the 0-th attention head (this belongs to the content of the first summation formula, and the calculation of the two inner product results of the 1-st attention head and the 2-nd attention head and the 0-th attention head is the same); Add up the two inner product results of the 0-th attention head, the two inner product results of the 1-st attention head, and the two inner product results of the 2-nd attention head (this belongs to the content of the second summation formula).

[0076] In the previous embodiment, multiple first calculation results are summed to obtain a first regularization term. Among them, one first calculation result is equivalent to an inner product result (in the example given, there are 6 inner product results, which can be understood as summing 6 inner product results at once, or as first summing the two inner product results of each attention head and then summing the summation results corresponding to the three attention heads).

[0077] Optionally, the second calculation module 303 is further configured to calculate a second regularization term based on the fourth matrix corresponding to each attention head: According to the serial number of the self-attention network layer in the self-attention model, perform a matrix inner product operation on the fourth matrices corresponding to the attention heads with the same serial number in every two self-attention network layers in the self-attention model in the reverse multiplication order to obtain multiple second calculation results; Sum up the multiple second calculation results to obtain the second regularization term.

[0078] For example, if the self-attention model has $N$ self-attention network layers and each self-attention network layer has $H$ attention heads, the second regularization term $R2$ can be calculated by the following formula:

[0079]

[0080] For example: The self-attention model has 4 self-attention network layers, and each self-attention network layer has 3 attention heads. Then N is 4 and H is 3. The calculation process is as follows: The 0th attention head of the 0th self-attention network layer performs a matrix inner product operation with the 0th attention head of the 2nd self-attention network layer. The 0th attention head of the 1st self-attention network layer performs a matrix inner product operation with the 0th attention head of the 3rd self-attention network layer. The two inner product results of the 0th attention head are added together (this belongs to the content of the first summation formula, and the calculation of the two inner product results of the 1st attention head and the 2nd attention head and the 0th attention head is the same); The two inner product results of the 0th attention head, the two inner product results of the 1st attention head, and the two inner product results of the 2nd attention head are added together (this belongs to the content of the second summation formula).

[0081] In the previous embodiment, multiple second calculation results are summed to obtain a second regularization term. Among them, one second calculation result is equivalent to one inner product result (in the example given, there are 6 inner product results, which can be understood as summing 6 inner product results at one time, or as first performing the first summation on the two inner product results of each attention head, and then performing the second summation on the summation results corresponding to the three attention heads).

[0082] Optionally, the second calculation module 303 is further configured to calculate a third regularization term based on the fourth matrix corresponding to each attention head: perform multiple matrix inner product operations on the fourth matrices corresponding to multiple attention heads in each self-attention network layer of the self-attention model to obtain multiple third calculation results, where the number of attention heads in each self-attention network layer is equal to the number of times of matrix inner product operations; sum the multiple third calculation results to obtain the third regularization term.

[0083] For example, if the self-attention model has N self-attention network layers and each self-attention network layer has H attention heads, the third regularization term R3 can be calculated by the following formula:

[0084]

[0085] For example: The self-attention model has 3 self-attention network layers, and each self-attention network layer has 3 attention heads. Then N is 3 and H is 3. The calculation process is as follows: For the 0th self-attention network layer: Calculate the matrix inner product operation between the 0th attention head and the 1st attention head, and the matrix inner product operation between the 1st attention head and the 2nd attention head, and add up the two inner product results of the 0th self-attention network layer (this belongs to the content of the first summation formula, and the calculation of the two inner product results of the 1st self-attention network layer, the 2nd self-attention network layer and the 0th self-attention network layer is the same); Add up the two inner product results of the 0th self-attention network layer, the two inner product results of the 1st self-attention network layer and the two inner product results of the 2nd self-attention network layer (this belongs to the content of the second summation formula).

[0086] In the previous embodiment, multiple third calculation results are summed to obtain a third regularization term. Among them, one third calculation result is equivalent to an inner product result (in the example given, there are 6 inner product results, which can be understood as summing 6 inner product results at one time, or as first summing the two inner product results of each self-attention network layer for the first time, and then summing the corresponding summation results of the three self-attention network layers for the second time).

[0087] Optionally, the update module 304 is further configured to use the loss function corresponding to the self-attention model before update obtained as the original loss function; perform a weighted summation operation on the original loss function and the first regularization term according to a preset ratio to obtain the loss function corresponding to the updated self-attention model; or perform a weighted summation operation on the original loss function and the second regularization term according to a preset ratio to obtain the loss function corresponding to the updated self-attention model; or perform a weighted summation operation on the original loss function and the third regularization term according to a preset ratio to obtain the loss function corresponding to the updated self-attention model.

[0088] Update the loss function corresponding to the self-attention model according to the following formula:

[0089] loss = (1 - α)·suploss + α·regulation

[0090] loss is the loss function corresponding to the updated self-attention model; suploss is the original loss function, that is, the loss function corresponding to the self-attention model before update; regulation is one of the first regularization term, the second regularization term and the third regularization term; α is the partition coefficient, that is, the preset ratio, used to adjust the model training accuracy, and empirically can be taken between 0.1 and 0.2, and this value may need to be adjusted according to different tasks.

[0091] Optionally, the training module 305 is further configured to obtain training data, where the training data includes target data; calculate the loss value corresponding to the training data by using the updated loss function; and update the model parameters of the self-attention model based on the loss value.

[0092] For example, in the face recognition scenario, the training data is face images; for example, in the text recognition scenario, the training data is text; the target data can be a part or all of the training data. The loss function corresponds to the self-attention model, and different types of self-attention models have different loss functions. Updating the model parameters of the self-attention model based on the loss value is to adjust the model parameters of the self-attention model with the principle of minimizing the loss value. That is to say, after updating the model parameters of the self-attention model, the loss value corresponding to the training data is calculated again by using the updated loss function, and at this time, a minimum loss value can be obtained (compared with the situation before updating the model parameters of the self-attention model and the situation where the model parameters of the self-attention model are other values). After updating, the numerical value of the model parameters of the self-attention model is optimal.

[0093] Through the above technical means, when the user enters the current session each time, in the current session, without waiting for the user to issue a specific question statement, one or more question statements that the user is most likely to ask, that is, the question records, can be determined according to the first time when the current question is initiated and the user's historical question records. Furthermore, according to one or more question statements that the user is most likely to ask, the most suitable object can be recommended to the user. Each question record is a statement used by the user to consult questions.

[0094] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.

[0095] Figure 4 It is a schematic diagram of the electronic device 4 provided by the embodiment of the present disclosure. As Figure 4 shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above various method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above device embodiments are implemented.

[0096] The electronic device 4 can be a desktop computer, a notebook, a palm computer, a cloud server, and other electronic devices. The electronic device 4 may include, but is not limited to, the processor 401 and the memory 402. Those skilled in the art can understand, Figure 4This is only an example of the electronic device 4, which does not constitute a limitation on the electronic device 4. It may include more or fewer components than those shown in the figure, or different components.

[0097] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0098] The memory 402 may be an internal storage unit of the electronic device 4. For example, the hard disk or memory of the electronic device 4. The memory 402 may also be an external storage device of the electronic device 4. For example, a plug-in hard disk equipped on the electronic device 4, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. The memory 402 may also include both an internal storage unit and an external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.

[0099] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0100] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present disclosure, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0101] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.

Claims

1. A training method for a self-attention model, characterized in that Including: Obtain target data, and input the target data into the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head respectively, so that the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head output a first matrix, a second matrix, and a third matrix respectively, where the self-attention model includes multiple self-attention network layers, and each self-attention network layer includes multiple attention heads; the target data is a face picture or text. Calculate a fourth matrix corresponding to each attention head based on the first matrix, the second matrix, and the third matrix corresponding to each attention head. Calculate any one of the following regularization terms based on the fourth matrix corresponding to each attention head: the first regularization term, the second regularization term, and the third regularization term. Update the loss function corresponding to the self-attention model based on any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term. Train the self-attention model using the updated loss function.

2. The method according to claim 1, characterized in that, The calculating a fourth matrix corresponding to each attention head based on the first matrix, the second matrix, and the third matrix corresponding to each attention head includes: Perform a vector inner product operation on the vectors in the first matrix and the vectors in the second matrix corresponding to each attention head according to the positions corresponding to the vectors, to obtain a target vector corresponding to each attention head. Perform a vector multiplication operation on the data in the target vector corresponding to each attention head and the vectors in the third matrix according to the positions corresponding to the data in the target vector and the vectors in the third matrix, to obtain a fourth matrix corresponding to each attention head.

3. The method according to claim 1, wherein The calculating any one of the following regularization terms based on the fourth matrix corresponding to each attention head: the first regularization term, the second regularization term, and the third regularization term includes: Calculate the first regularization term based on the fourth matrix corresponding to each attention head: Perform a matrix inner product operation on the fourth matrices corresponding to the attention heads with the same serial number in every two adjacent self-attention network layers in the self-attention model, to obtain a plurality of first calculation results. Sum the plurality of first calculation results to obtain the first regularization term.

4. The method according to claim 1, characterized in that The calculating any one of the following regularization terms based on the fourth matrix corresponding to each attention head: the first regularization term, the second regularization term, and the third regularization term includes: Calculate the second regularization term based on the fourth matrix corresponding to each attention head: Perform a matrix inner product operation on the fourth matrices corresponding to the attention heads with the same serial number in every two self-attention network layers in the self-attention model according to the reverse multiplication order based on the serial number of the self-attention network layer in the self-attention model, to obtain a plurality of second calculation results. Sum the plurality of second calculation results to obtain the second regularization term.

5. The method according to claim 1, wherein The calculating any one of the following regularization terms based on the fourth matrix corresponding to each attention head: the first regularization term, the second regularization term, and the third regularization term includes: Calculate the third regularization term based on the fourth matrix corresponding to each attention head: Perform multiple matrix inner product operations on the fourth matrices corresponding to multiple attention heads in each self-attention network layer of the self-attention model, to obtain multiple third calculation results, where the number of attention heads in each self-attention network layer is equal to the number of times of the multiple matrix inner product operations; Sum the multiple third calculation results to obtain the third regularization term.

6. The method according to claim 1, wherein Update the loss function corresponding to the self-attention model based on any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term, including: Take the loss function corresponding to the self-attention model obtained before the update as the original loss function; Perform a weighted sum operation on the original loss function and the first regularization term according to a preset ratio, to obtain the loss function corresponding to the updated self-attention model; or Perform a weighted sum operation on the original loss function and the second regularization term according to the preset ratio, to obtain the loss function corresponding to the updated self-attention model; or Perform a weighted sum operation on the original loss function and the third regularization term according to the preset ratio, to obtain the loss function corresponding to the updated self-attention model.

7. The method according to claim 1, wherein Training the self-attention model using the updated loss function includes: Obtain training data, where the training data includes the target data; Calculate the loss value corresponding to the training data using the updated loss function; Update the model parameters of the self-attention model based on the loss value.

8. A training device for a self-attention model, characterized in that Includes: An acquisition module, configured to acquire target data, and input the target data into the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head respectively, so that the first linear transformation unit, the second linear transformation unit, and the third linear transformation unit in each attention head output a first matrix, a second matrix, and a third matrix respectively, where the self-attention model includes multiple self-attention network layers, and each self-attention network layer includes multiple attention heads; the target data is a face image or text; A first calculation module, configured to calculate the fourth matrix corresponding to each attention head based on the first matrix, the second matrix, and the third matrix corresponding to each attention head; A second calculation module, configured to calculate any one of the following regularization terms based on the fourth matrix corresponding to each attention head: the first regularization term, the second regularization term, and the third regularization term; An update module, configured to update the loss function corresponding to the self-attention model based on any one of the following regularization terms: the first regularization term, the second regularization term, and the third regularization term; A training module, configured to train the self-attention model using the updated loss function.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multi-head attention memory network for short text sentiment classification

    CN112784532A

  • Liner transformation matrix calculating apparatus, method thereof and program thereof

    US20100195917A1