Attention migration method, data processing method and large language model

By using the first attention layer, structural compression layer and multiple second attention layers in the student model, the parameters of the student model are optimized, and the problem of information loss during the knowledge transfer process in the existing technology is solved, and lossless attention transfer and student model performance improvement are achieved.

CN120146154APending Publication Date: 2025-06-13ALIBABA (CHINA) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510237970.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing technology will lose some information during the migration of knowledge from teacher model to student model, resulting in poor student model performance.

Method used

By acquiring the pre-trained teacher model and the initialized student model, the parameters of the student model are optimized to achieve lossless attention transfer using the first attention layer, the structural compression layer and the multiple second attention layers.

Benefits of technology

This avoids information loss during attention transfer, improves the performance of the student model, and makes it closer to the performance performance of the teacher model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146154A_ABST
    Figure CN120146154A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an attention migration method, a data processing method and a large language model. A pre-trained teacher model and an initialized student model are obtained, the initialized student model comprises a first attention layer, a structure compression layer and a plurality of second attention layers, and parameters of the first attention layer are the same as parameters of the attention layer of the first layer of the teacher model. The parameter of the second attention layer is an initialization parameter, obtaining a training sample and corresponding label information, obtaining prediction information corresponding to the training sample through an initialized student model, and optimizing the initialization parameter according to the label information and the prediction information to obtain a pre-trained student model. Therefore, information loss in the attention migration process can be avoided, lossless migration is realized, and the performance of the student model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an attention transfer method, a data processing method, and a large language model. Background Art

[0002] Edge intelligence is a technology that involves directly performing data processing and AI (Artificial Intelligence) algorithm applications on terminal devices (such as smartphones, in-vehicle infotainment systems, robots, and speakers, etc.). In particular, the deployment of large models for generative AI on the edge, represented by mobile phones, is driving the rapid development of multi-functional and multi-modal applications. This technology not only helps reduce server costs, protect user information security, but also improves real-time response speed and realizes personalized user experiences, showing broad application prospects. With the progress of AI technology, edge large models are gradually becoming an important trend in artificial intelligence applications.

[0003] In the prior art, knowledge distillation technology has been widely applied to the optimization of edge large models. For example, using the KL (Kullback-Leibler Divergence) divergence as a loss function to transfer the attention information of the intermediate layer of the teacher model to the student model, distilling the self-attention module of the last Transformer layer of the teacher model, etc. The core of these solutions lies in using a large teacher model to guide the learning of a small student model, aiming to reduce the model size and computational complexity while maintaining the performance unchanged or close.

[0004] However, in the process of knowledge transfer from the teacher model to the student model in these two methods of the prior art, some information will be lost, resulting in poor performance of the student model. Summary of the Invention

[0005] In view of this, the purpose of the embodiments of the present invention is to provide an attention transfer method, a data processing method, and a large language model, which can avoid information loss during the attention transfer process, achieve lossless transfer, and improve the performance of the student model.

[0006] In a first aspect, an embodiment of the present invention provides an attention transfer method, the method comprising:

[0007] Obtain a pre-trained teacher model and an initialized student model, where the initialized student model includes a first attention layer, a structure compression layer, and a plurality of second attention layers, the parameters of the first attention layer are the same as the parameters of the attention layer of the first layer of the teacher model, and the parameters of the second attention layer are initialized parameters;

[0008] Obtain training samples;

[0009] Obtain the label information corresponding to the training samples;

[0010] Obtain the prediction information corresponding to the training sample through the initialized student model;

[0011] Optimize the initialized parameters according to the label information and the prediction information to obtain a pre-trained student model.

[0012] In some embodiments, the label information corresponding to the training sample is a pre-marked true label.

[0013] In some embodiments, obtaining the label information corresponding to the training sample includes:

[0014] Obtain the label information through the pre-trained teacher model according to the training sample.

[0015] In some embodiments, obtaining the prediction information corresponding to the training sample through the initialized student model includes:

[0016] Generate a first intermediate parameter according to the training sample through the first attention layer of the initialized student model;

[0017] Compress the first intermediate parameter through the structure compression layer of the initialized student model to obtain a second intermediate parameter;

[0018] Obtain the prediction information according to the second intermediate parameter through the multiple second attention layers.

[0019] In some embodiments, obtaining the prediction information according to the second intermediate parameter through the multiple second attention layers:

[0020] Obtain the initialized parameters of each second attention layer;

[0021] Perform a linear transformation and ReLU activation on the attention matrix of the attention layer of the first layer of the teacher model according to the initialized parameters to obtain the attention matrix of each second attention layer;

[0022] Generate an attention coefficient according to the attention matrix of the second attention layer through the normalized exponential function;

[0023] Obtain the prediction information according to the attention coefficient and the second intermediate parameter.

[0024] In some embodiments, optimizing the initialized parameters according to the label information and the prediction information to obtain a pre-trained student model includes:

[0025] Calculate a loss value according to the label information and the prediction information;

[0026] Optimize the initialization parameters according to the loss value to obtain a pre-trained student model.

[0027] In some embodiments, the loss value is calculated by the following formula:

[0028]

[0029] where L(u) is the loss value, u i is the target word, u i-k ,…,u i-1 are the context words before the target word, and θ is a learnable parameter.

[0030] In a second aspect, an embodiment of the present invention provides a data processing method, the method comprising:

[0031] Receiving input information;

[0032] Generating a first intermediate parameter according to the input information through a first attention layer;

[0033] Compressing the first intermediate parameter through a structure compression layer to obtain a second intermediate parameter;

[0034] Obtaining output information according to the second intermediate parameter through a plurality of second attention layers.

[0035] In a third aspect, an embodiment of the present invention provides a large language model applied to the terminal side, the large language model comprising:

[0036] A first attention layer, configured to generate a first intermediate parameter according to input information;

[0037] A structure compression layer, configured to compress the first intermediate parameter to obtain a second intermediate parameter;

[0038] A plurality of second attention layers, configured to obtain output information according to the second intermediate parameter.

[0039] In a fourth aspect, an embodiment of the present invention provides an attention transfer device, the device comprising:

[0040] A model acquisition unit, configured to acquire a pre-trained teacher model and an initialized student model, the initialized student model comprising a first attention layer, a structure compression layer, and a plurality of second attention layers, wherein the parameters of the first attention layer are the same as those of the attention layer of the first layer of the teacher model, and the parameters of the second attention layer are initialization parameters;

[0041] A sample acquisition unit, configured to acquire training samples;

[0042] A label acquisition unit, configured to acquire label information corresponding to the training samples;

[0043] A prediction information acquisition unit, configured to acquire prediction information corresponding to the training sample through the initialized student model;

[0044] A model training unit, configured to optimize the initialized parameters according to the label information and the prediction information to obtain a pre-trained student model.

[0045] In a fifth aspect, an embodiment of the present invention provides a data processing device, where the device includes:

[0046] An information receiving unit, configured to receive input information;

[0047] A first parameter acquisition unit, configured to generate a first intermediate parameter according to the input information through a first attention layer;

[0048] A second parameter acquisition unit, configured to compress the first intermediate parameter through a structure compression layer to obtain a second intermediate parameter;

[0049] An output information acquisition unit, configured to acquire output information according to the second intermediate parameter through a plurality of second attention layers.

[0050] In a sixth aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, where the memory is used to store one or more computer program instructions, and wherein the one or more computer program instructions are executed by the processor to implement the methods described in the first aspect and the second aspect.

[0051] In a seventh aspect, an embodiment of the present invention provides a computer program product, where the computer program product includes a computer program, and when the computer program runs on a computer, the computer executes the methods described in the first aspect and the second aspect.

[0052] In an eighth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions, when executed by a processor, implement the methods described in the first aspect and the second aspect.

[0053] The technical solution of the embodiment of the present invention obtains a pre-trained teacher model and an initialized student model. The initialized student model includes a first attention layer, a structure compression layer, and multiple second attention layers. The parameters of the first attention layer are the same as those of the attention layer of the first layer of the teacher model, and the parameters of the second attention layer are initialized parameters. Training samples and corresponding label information are obtained, prediction information corresponding to the training samples is obtained through the initialized student model, and the initialized parameters are optimized according to the label information and the prediction information to obtain a pre-trained student model. Thus, information loss during the attention transfer process can be avoided, lossless transfer can be achieved, and the performance of the student model can be improved. Description of the Drawings

[0054] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0055] Figure 1 is a schematic diagram of attention transfer in an embodiment of the present invention;

[0056] Figure 2 is a flowchart of the attention transfer method in an embodiment of the present invention;

[0057] Figure 3 is a flowchart of obtaining prediction information through a student model in an embodiment of the present invention;

[0058] Figure 4 is a flowchart of the second attention layer obtaining prediction information in an embodiment of the present invention;

[0059] Figure 5 is a flowchart of the training process in an embodiment of the present invention;

[0060] Figure 6 is a flowchart of the data processing method in an embodiment of the present invention;

[0061] Figure 7 is a schematic diagram of the attention transfer device in an embodiment of the present invention;

[0062] Figure 8 is a schematic diagram of the data processing device in an embodiment of the present invention;

[0063] Figure 9 is a schematic diagram of the electronic device in an embodiment of the present invention. Detailed Embodiments

[0064] The present application is described below based on embodiments, but the present application is not limited to these embodiments. In the detailed description of the present application below, some specific details are described in detail. It is possible for those skilled in the art to fully understand the present application without the description of these details. In order to avoid confusing the essence of the present application, known methods, processes, flows, components and circuits are not described in detail.

[0065] In addition, persons of ordinary skill in the art will appreciate that the drawings provided herein are for illustration purposes and are not necessarily drawn to scale.

[0066] Unless the context clearly requires otherwise, the words "include", "comprising" and similar words throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, the meaning is "including but not limited to".

[0067] In the description of this application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" is two or more.

[0068] The solutions described in this specification and in the examples, if they involve the processing of personal information, will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.

[0069] End-side intelligence refers to the technology of directly processing data and applying AI algorithms on terminal devices (such as smartphones, car computers, robots, speakers, etc.). The end-side big model represented by generative AI deployed on the mobile phone end is setting off a multi-functional and multi-modal wave, which has brought broader application prospects in terms of saving server costs, protecting user information security, improving real-time performance and realizing personalized user experience.

[0070] Knowledge distillation is particularly important in large models on the end, because the computing power and storage space of terminal devices are usually limited. Through knowledge distillation, the size and inference time of the model can be greatly reduced without significantly reducing performance, thereby achieving efficient deployment in resource-constrained environments such as mobile devices and embedded systems.

[0071] Knowledge distillation is a model compression and optimization technique. Its core idea is to use the soft labels of the teacher model to train the student model, enabling it to mimic the behavior of the teacher model while maintaining a small size and low computational complexity. However, the knowledge distillation process itself is a black-box operation, making it difficult to intuitively explain why one student model is better than another. The knowledge transfer from a large teacher model to a small student model is not always direct or simple, and may require multiple iterations and adjustments to achieve better results.

[0072] Essentially, during the process of distilling a large model, the attention of the teacher model and the attention of the student model are used as the loss function through KL divergence to drive the attention distribution of the student model to approach that of the teacher model. How to transfer the attention of the teacher model without loss is a very important issue.

[0073] Figure 1 It is a schematic diagram of attention transfer in an embodiment of the present invention. As Figure 1 shown, during the attention transfer process, the teacher model 1 and the student model 2 are involved.

[0074] Among them, the teacher model (also known as the Teacher model) refers to a pre-trained large language model, which is known for its superior performance. However, its high computational complexity and resource requirements make it difficult to directly deploy on mobile devices or other resource-constrained environments. Such teacher models are usually pre-trained based on various existing open-source models, and they perform well in terms of accuracy and functional richness. However, due to their complex architectures and large number of parameters, directly applying these high-performance models to edge devices faces many challenges, including but not limited to high computational resource consumption, high storage space occupancy, and insufficient real-time processing capabilities. Among them, the teacher model 1 is various pre-trained models that can be implemented through various existing open-source models.

[0075] The student model (Student model) 2 is the target model to be obtained in an embodiment of the present invention. Its design purpose is to be more lightweight and efficient, facilitating deployment in resource-constrained environments. Compared with the teacher model, the student model has fewer parameters and lower computational complexity, making it more suitable for applications in scenarios such as mobile devices or embedded systems. By learning from the teacher model, the goal is to make the student model as close as possible to the performance of the teacher model. This means that while maintaining high prediction accuracy and functionality, the student model also needs to achieve fast response and low power consumption operation to meet the requirements of practical applications. The present invention focuses on optimizing this knowledge transfer process to ensure that the student model not only inherits the advantages of the teacher model but also achieves a significant improvement in resource utilization efficiency.

[0076] Both the teacher model 1 and the student model 2 are large language models (LLMs). Such models are trained using a vast amount of text data to generate natural language text or understand the meaning of language text. By performing unsupervised learning on a large dataset, large language models master the patterns and structures of natural language, thereby simulating the human language cognition and generation process to a certain extent. The core advantage of this deep learning model lies in its ability to obtain powerful general modeling capabilities and excellent generalization capabilities through large-scale pre-training. Therefore, LLMs perform well in a variety of application scenarios and can not only complete basic language processing tasks such as spelling checking and grammar correction, but also handle complex tasks such as text summarization, machine translation, sentiment analysis, dialogue generation, and content recommendation. In the present invention, the teacher model 1, as a pre-trained large language model, is characterized by its high performance, while the student model 2 aims to inherit the advantages of the teacher model and achieve higher resource utilization efficiency by reducing the number of parameters and optimizing the architecture design, so as to be deployed in an environment with limited computing resources. Such a setting not only fully utilizes the potential of large language models but also overcomes the challenges they encounter in practical applications.

[0077] Among them, the teacher model 1 includes n attention layers A 1 -A n . The attention layer is a Transformer Layer (encoder layer). The implementation principle of the Transformer Layer mainly includes key parts such as the self-attention mechanism, multi-head attention mechanism, and feed-forward network.

[0078] Among them, the self-attention mechanism is the core part of the Transformer Layer. It allows the model to compare each element in the input sequence with other elements when processing the sequence, so as to correctly process each element in different contexts. There are three important input matrices in the self-attention mechanism: the query matrix Q (query), the key matrix K (key), and the value matrix V (value). These three matrices are all obtained by different linear transformations of the input sequence. The product of the query matrix Q and the key matrix K passes through a softmax function to obtain a probability distribution with the same length as the input sequence, which represents the importance of each element for the query matrix Q. Finally, multiplying this probability distribution by the value matrix V gives the self-attention vector, which represents the result of weighted averaging the values of each element.

[0079] The multi-head attention mechanism learns different context representations by applying the self-attention mechanism to multiple different query matrices Q, key matrices K, and value matrices V. Specifically, the input sequence is passed through different linear transformations to obtain multiple different query matrices Q, key matrices K, and value matrices V, and then they are input into multiple parallel self-attention mechanisms for processing. In this way, the model can learn rich feature representations from different subspaces.

[0080] The feed-forward network is usually a two-layer fully connected neural network. This feed-forward network processes the output of the multi-head attention mechanism to further extract features. The feed-forward network usually uses the ReLU activation function, and each sub-layer will perform layer normalization to ensure the stability and convergence speed of the model.

[0081] For different types of large language models, the structure of the attention layer may vary. The attention layer can be an encoder architecture, a decoder architecture, or a combination of encoder and decoder architectures.

[0082] For the encoder architecture, the input information to be processed first enters the multi-head self-attention layer, where the information at each position can attend to all positions in the entire input sequence. This mechanism allows the model to process information in different representation subspaces in parallel at the same time point, thus enhancing the model's understanding ability. The initial output obtained from the multi-head self-attention mechanism will be added to the original input through a residual connection and then undergo layer normalization. This process helps to stabilize the training process and accelerate convergence. The normalized result will then be fed into a fully connected feed-forward neural network, which consists of two linear transformations and uses the ReLU activation function between them. This step is used to further transform and enhance the information at each position. The output of the feed-forward neural network will also be added to the data entering the network through a residual connection and then undergo layer normalization again. Such a design not only helps to alleviate the vanishing gradient problem in deep networks but also improves the model's expressiveness. Through the above steps, the encoder architecture can effectively capture long-range dependencies in the input sequence while maintaining computational efficiency. This architecture is the basis of many modern natural language processing models, such as BERT, which utilize the powerful capabilities of this structure to perform various complex language understanding and generation tasks.

[0083] For the decoder architecture, the input information to be processed passes through a masked multi-head self-attention layer. A masking technique is used here to prevent the prediction at position i from seeing information at any "future" position after position i during training. The output of the masked multi-head self-attention mechanism is added to the original input through a residual connection and layer normalization is performed. The normalized data is input to the Encoder-Decoder Multi-Head Attention layer, which allows the decoder to utilize all the outputs of the encoder as queries, keys, and values, enabling the decoder to attend to different parts of the input sequence. The output of the Encoder-Decoder Multi-Head Attention layer is added to the data input to this network through a residual connection, and then layer normalization is performed again. Next, the normalized result is passed to a fully-connected feed-forward neural network for further processing. The output of the feed-forward neural network is also added to the data input to this network through a residual connection, and then layer normalization is performed again.

[0084] For the combination of the encoder and decoder architectures, in the encoding stage, the input information to be processed first passes through a series of encoder layers, each layer including a multi-head self-attention mechanism, a residual connection and layer normalization, and a feed-forward neural network. This series of steps enables the encoder to capture rich information in the input sequence and form a deep representation. In the decoding stage, the decoder first receives a start token as input and processes it through a masked multi-head self-attention mechanism to ensure that only information at the current and previous time steps can be accessed at each time step. Then, through the Encoder-Decoder Multi-Head Attention mechanism, with the final output of the encoder as keys and values and the current output of the decoder as queries, the decoder can attend to information across the entire input sequence. Next, the decoder output passes through a feed-forward neural network and is processed through a residual connection and layer normalization. This process is iterated until an end token is generated, completing the generation of the entire sequence. The advantage of this encoder-decoder framework is that it can effectively handle input and output sequences of different lengths and dynamically determine which input information is important through the attention mechanism, greatly enhancing the model's ability to handle complex tasks.

[0085] Furthermore, the number of heads in each attention layer of the teacher model 1 is Ht, and the hidden layer dimension of each head is Dt, that is, the dimension of the data output by each attention layer is Ht * Dt.

[0086] The student model 2 includes a first attention layer B 1 , a structure compression layer B 2 , and multiple second attention layers B 3 -B m .

[0087] Among them, for the first attention layer B 1 , the first attention layer B 1 has the same parameters as the attention layer A of the first layer of the teacher model 1 . That is to say, the student model uses the first layer of the teacher model as its own first layer, so that it can be ensured that whether in training or inference, the attention and output of the student model and the teacher model on the first layer are the same. In this way, multiple second attention layers B 3 -B m of the student model can be directly obtained based on the original attention of the teacher model

[0088] For the structure compression layer B 2 , since the parameters of the second attention layer of the student model are inconsistent with those of the attention layer of the teacher model, it is necessary to convert the parameters of the teacher model into those of the student model. Specifically, the structure compression layer B 2 is used to compress the Ht heads with a total of Ht*Dt dimensions output by the attention layer A 1 of the first layer of the teacher model into the Ds dimensions of one head of the student model. Among them, the number of heads of the attention layer of the teacher model is Ht times that of the student model (Ht is an integer).

[0089] Among them, the method of compressing Ht*Dt dimensions into Ds dimensions can be implemented based on various existing methods. For example, the space of Ht*Dt dimensions can be mapped to Ds dimensions by means of linear transformation. For another example, the space of Ht*Dt dimensions can also be mapped to Ds dimensions through dimensionality reduction technology to capture and retain the most important features. Among them, the dimensionality reduction technology can adopt principal component analysis (PCA, Principal Component Analysis), autoencoders, etc.

[0090] For the second attention layer, the second attention layer B 3 -B m of the student model reuses the attention matrix of the first attention layer A1 of the teacher model. Assuming that the input length is k, for the teacher model, its attention matrix is k*k dimensional, and for the student model, its attention matrix is also k*k dimensional, so the student model can reuse the attention matrix of the teacher model. For the student model, if the attention coefficients of the second attention layer B 3 -B m are exactly the same, the expression ability of the student model is too weak. Therefore, in the embodiments of the present invention, the first attention layer A 1The attention matrix of the Ht heads (before softmax) is linearly transformed and ReLU-activated for each corresponding element to obtain the attention matrix of one head of the student model, and then converted into attention coefficients through Softmax.

[0091] Specifically, an initialization parameter is set for each second attention layer. The initialization parameter includes a weight value W and a bias value b. The first attention layer A in the teacher model is extracted. 1 The attention matrices of the Ht heads (not processed by softmax) of the teacher model, and the size of these matrices is k*k, where k is the length of the input sequence. Let the set composed of these attention matrices be {F (1) , F (2) , ……, F (Ht)}, where F (g) represents the attention matrix of the g-th head. For each attention matrix F (g) of the teacher model, the attention matrix of the second attention layer is obtained through linear transformation and ReLU activation with the said initialization parameter. Among them, the calculation formula of the linear transformation is as follows:

[0092] Z (g) = W * F (g) + b

[0093] where F (g) represents the attention matrix of the g-th head of the first attention layer A 1 of the teacher model, and Z (g) represents the first intermediate matrix after linear transformation of F (g) , W is the weight value, and b is the bias value.

[0094] Furthermore, since the first attention layer A 1 in the teacher model has Ht heads, therefore, the transformed first intermediate matrix of each head can be obtained through the above formula.

[0095] For each first intermediate matrix, it is processed through the ReLU (Rectified Linear Unit) activation function to obtain the second intermediate matrix. Among them, the calculation formula of the second intermediate matrix is as follows:

[0096] H (g) = ReLU(Z (g) )

[0097] where H (g) is the second intermediate matrix corresponding to the attention matrix of the g-th head of the first attention layer A 1 of the teacher model, and Z (g) represents the first attention layer A of the teacher model.1 The first intermediate matrix corresponding to the attention matrix of the g-th head.

[0098] Finally, merge the Ht second intermediate matrices to obtain the attention matrix of the second attention layer. In the embodiments of the present invention, the attention matrix of the second attention layer is denoted as P, where P i represents the attention matrix corresponding to the second attention layer B i i = 3, 4,..., m. Among them, the merging method can be various existing methods, for example, taking the average or weighted sum of all attention matrices by element, etc.

[0099] Since there are a total of m - 2 second attention layers, therefore, it is necessary to transform m - 2 times to obtain the attention matrix of each second attention layer.

[0100] After obtaining the attention matrix of the second attention layer, generate attention coefficients according to the attention matrix of the second attention layer through the normalized exponential (Softmax) function. The formula for obtaining the attention coefficients is as follows:

[0101] C i = Softmax(P i )

[0102] where P i represents the attention matrix corresponding to the second attention layer B i and C i represents the attention coefficient corresponding to the second attention layer B i .

[0103] In the embodiments of the present invention, by obtaining the initialization parameters of each second attention layer, performing a linear transformation and ReLU activation on the attention matrix of the attention layer of the first layer of the teacher model according to the initialization parameters to obtain the attention matrix of each second attention layer, and generating attention coefficients according to the attention matrix of the second attention layer through the normalized exponential function. Thus, the attention of the teacher model can be directly reused without distillation. At the same time, the attention of different layers of the student model can be different, and the computational complexity of the student model can be reduced.

[0104] Since the attention coefficients of the student model come from the first attention layer of the teacher model, it is not necessary to calculate the values of Q, K, and V. When calculating each second attention layer, only the output of the previous layer needs to be linearly transformed once to obtain V. The entire formula can be expressed as:

[0105] Attention(F, V) = softmax(ReLU(W * F + b))V

[0106] Among them, F represents the attention matrix of the first attention layer A of the teacher model, W represents the weight value, b represents the bias value, Attention(F, V) represents the output information of the second attention layer, and V represents the parameter after the input information of the second attention layer is linearly transformed. 1 The attention matrix of 1 , W represents the weight value, b represents the bias value, Attention(F, V) represents the output information of the second attention layer, and V represents the parameter after the input information of the second attention layer is linearly transformed.

[0107] Figure 2 is the flowchart of the attention transfer method of the embodiment of the present invention. Figure 2 The attention transfer method shown can be executed by various electronic devices. The electronic devices include a memory and a processor. The memory is used to store one or more computer program instructions. Among them, the one or more computer program instructions are executed by the processor to implement the attention transfer method of the embodiment of the present invention. The attention transfer method specifically includes the following steps:

[0108] Step S110, obtain a pre-trained teacher model and an initialized student model.

[0109] In this embodiment, the teacher model is various pre-trained models, which can be implemented through various existing open-source models. The initialized student model includes a first attention layer, a structure compression layer, and multiple second attention layers. The parameters of the first attention layer are the same as those of the attention layer of the first layer of the teacher model. The parameters of the second attention layer are initialized parameters, and the initialized parameters include weight values and bias values.

[0110] Step S120, obtain training samples.

[0111] In this embodiment, the training samples are text information to be processed, which can be set according to the actual application scenario.

[0112] Step S130, obtain the label information corresponding to the training samples.

[0113] In an optional implementation, the label information corresponding to the training samples is a pre-labeled true label. That is to say, the label information is a manually marked true label.

[0114] In another optional implementation, the label information is obtained through the pre-trained teacher model according to the training samples. That is to say, after obtaining the training samples, the teacher model processes the training samples to obtain the corresponding label information.

[0115] Step S140, obtain the prediction information corresponding to the training samples through the initialized student model.

[0116] In this embodiment, the training samples are input into the initialized student model, and the prediction information corresponding to the training samples is obtained through the initialized student model.

[0117] Specifically, Figure 3 is a flowchart of obtaining prediction information through the student model according to an embodiment of the present invention. As Figure 3 shown, obtaining the prediction information corresponding to the training samples through the initialized student model includes the following steps:

[0118] Step S141: Generate first intermediate parameters according to the training samples through the first attention layer of the initialized student model.

[0119] In this embodiment, after receiving the training samples, the initialized student model generates first intermediate parameters according to the training samples through the first attention layer.

[0120] Step S142: Compress the first intermediate parameters through the structure compression layer of the initialized student model to obtain second intermediate parameters.

[0121] In this embodiment, the first intermediate parameters output by the first attention layer are provided to the structure compression layer, and the structure compression layer compresses the first intermediate parameters to obtain second intermediate parameters.

[0122] Specifically, the attention layer of the first layer of the teacher model outputs a total of Ht heads with a dimension of Ht*Dt. The parameters of the first attention layer of the student model are the same as those of the attention layer of the first layer of the teacher model. Therefore, the dimension of the first intermediate parameters output by the first attention layer of the student model is also Ht*Dt. The first intermediate parameters with a dimension of Ht*Dt are compressed into a Ds dimension of one head of the student model through the structure compression layer to obtain second intermediate parameters.

[0123] Step S143: Obtain the prediction information according to the second intermediate parameters through the multiple second attention layers.

[0124] In this embodiment, the second intermediate parameters output by the structure compression layer are provided to the second attention layer, and the multiple second attention layers obtain the prediction information according to the second intermediate parameters.

[0125] Specifically, the input of the first second attention layer is the second intermediate parameters, the input of the subsequent second attention layers is the output of the previous second attention layer, and the output of the last second attention layer is the prediction information.

[0126] Figure 4 is a flowchart of the second attention layer obtaining prediction information according to an embodiment of the present invention. As Figure 4As shown, obtaining the prediction information according to the second intermediate parameter through the multiple second attention layers includes the following steps:

[0127] Step S1431: Obtain the initialization parameters of each second attention layer.

[0128] In this embodiment, an initialization parameter is set for each second attention layer, and the initialization parameter includes a weight value W and a bias value b. Among them, the initialization parameters of different second attention layers can be the same or different.

[0129] Step S1432: Perform a linear transformation and ReLU activation on the attention matrix of the attention layer of the first layer of the teacher model according to the initialization parameter to obtain the attention matrix of each second attention layer.

[0130] In this embodiment, extract the attention matrix of the first attention layer A 1 with Ht heads in the teacher model, and perform a linear transformation on the attention matrix through the initialization parameter to obtain a first intermediate matrix. Among them, the calculation formula of the linear transformation is as follows:

[0131] Z (g) = W * F (g) + b

[0132] Among them, F (g) represents the attention matrix of the g-th head of the first attention layer A 1 in the teacher model, Z (g) represents the first intermediate matrix after linear transformation of F (g) , W is the weight value, and b is the bias value.

[0133] For each first intermediate matrix, it is processed through the ReLU (Rectified Linear Unit) activation function to obtain a second intermediate matrix.

[0134] Among them, the calculation formula of the second intermediate matrix is as follows:

[0135] H (g) = ReLU(Z (g) )

[0136] Among them, H (g) is the second intermediate matrix corresponding to the attention matrix of the g-th head of the first attention layer A 1 in the teacher model, and Z (g) represents the first intermediate matrix corresponding to the attention matrix of the g-th head of the first attention layer A 1 in the teacher model.

[0137] Finally, merge Ht second intermediate matrices to obtain the attention matrix of the second attention layer.

[0138] Step S1433: Generate attention coefficients according to the attention matrix of the second attention layer through the softmax function.

[0139] In this embodiment, after obtaining the attention matrix of the second attention layer, generate attention coefficients according to the attention matrix of the second attention layer through the softmax function. The formula for obtaining attention coefficients is as follows:

[0140] C i = Softmax(P i )

[0141] where P i represents the attention matrix corresponding to the second attention layer B i , and C i represents the attention coefficient corresponding to the second attention layer B i .

[0142] Step S1434: Obtain the prediction information according to the attention coefficient and the second intermediate parameter.

[0143] Since the attention coefficients of the student model come from the first attention layer of the teacher model, it is not necessary to calculate the values of Q, K, and V. When calculating each second attention layer, only need to linearly transform the output of the previous layer to obtain V. The whole formula can be expressed as:

[0144] Attention(F, V) = C * V

[0145] where C represents the attention coefficient of the second attention layer, F represents the attention matrix of the first attention layer A 1 of the teacher model, and V represents the parameter after linearly transforming the input information of the second attention layer.

[0146] Furthermore, the input information of the first second attention layer is the second intermediate parameter, the input information of the subsequent second attention layers is the output information of the previous second attention layer, and the output information of the last second attention layer is the prediction information.

[0147] Step S150: Optimize the initialization parameters according to the label information and the prediction information to obtain a pre-trained student model.

[0148] In this embodiment, after obtaining the prediction information through the student model, train the student model according to the label information and the prediction information.

[0149] Specifically,Figure 5 is a flowchart of the training process of an embodiment of the present invention. As Figure 5 shown, optimizing the initialization parameters according to the label information and the prediction information to obtain a pre-trained student model includes the following steps:

[0150] Step S151, calculate a loss value according to the label information and the prediction information.

[0151] In this embodiment, a loss value is calculated according to the label information and the prediction information through a predetermined loss function, where the loss value is calculated and obtained through the following formula:

[0152]

[0153] where L(u) is the loss value, u i is the target word, u i-k ,…,u i-1 are the context words before the target word, and θ is a learnable parameter.

[0154] where θ is a learnable parameter, that is, the initialization parameter of the embodiment of the present invention, including the weight values and bias values of all the second attention layers.

[0155] Step S152, optimize the initialization parameters according to the loss value to obtain a pre-trained student model.

[0156] In this embodiment, an optimization algorithm is used to update the initialization parameters according to the calculated loss value to obtain a pre-trained student model.

[0157] Specifically, a batch of training samples are input into the initialized student model. For each training sample, it is processed through a first attention layer, a structure compression layer, and multiple second attention layers in turn to obtain prediction information. The cross-entropy loss, that is, the loss value, is calculated based on the prediction information and the label information output by the student model. According to the calculated loss value, all parameters of all the second attention layers perform a backpropagation algorithm to calculate the gradient. An optimization algorithm (such as SGD, Adam, etc.) is used to update the model parameters according to the calculated gradient, and the above steps are repeated until the stop condition is met.

[0158] Under the teacher-student paradigm, transferring the knowledge of the large model on the cloud side to the large model on the edge side usually adopts the distillation method. However, distillation will inevitably bring about the attenuation of "information volume" during the learning process of the student model, which is a lossy process. An embodiment of the present invention proposes a teacher-student model with a progressive (parameter progressive from more to less) Transformer architecture. A structure compression layer is embedded in the Transformer model in this solution. During the training process of the student model, the distillation process and the attention calculation process of the transposed matrix multiplication of query and key are eliminated, and the attention of the teacher model is directly and losslessly transferred to the student model, and the computational cost of the large model on the edge side during the training and inference processes can be greatly reduced.

[0159] Specifically, the student model uses the first layer of the teacher model as its first layer, so as to ensure that the attention and output of the student model and the teacher model on the first layer are exactly the same whether in training or inference. In this way, the attention of the 3rd to nth layers can be directly obtained based on the original attention of the teacher model, realizing the distillation-free attention transfer process. Moreover, the 3rd to nth layers of the student model no longer need to use QKV to calculate the attention coefficient, greatly saving the computational cost of training and inference.

[0160] An embodiment of the present invention obtains a pre-trained teacher model and an initialized student model. The initialized student model includes a first attention layer, a structure compression layer, and multiple second attention layers. The parameters of the first attention layer are the same as the parameters of the attention layer of the first layer of the teacher model, and the parameters of the second attention layer are initialized parameters. Training samples and corresponding label information are obtained, and the prediction information corresponding to the training samples is obtained through the initialized student model. The initialized parameters are optimized according to the label information and the prediction information to obtain a pre-trained student model. Thus, information loss during the attention transfer process can be avoided, lossless transfer can be realized, and the performance of the student model can be improved.

[0161] Figure 6It is a flowchart of the data processing method according to the embodiments of the present invention. After the student model is trained in the embodiments of the present invention, the pre-trained student model can be deployed on the terminal side in the form of an APP (application), a mini-program, a web page, etc. to implement corresponding functions. Among them, the terminal side can be a mobile phone, a tablet computer, a notebook computer, a desktop computer, a vehicle-mounted device, a wearable device, a robot, a speaker, or other devices with data processing functions. Specifically, after the student model is deployed, the terminal side can directly perform data processing and AI algorithm application technologies. As Figure 6 shown, the data processing method includes the following steps:

[0162] Step S210: Receive input information.

[0163] In this embodiment, the terminal side can provide a human-computer interaction interface to the user through an APP, a mini-program, a web page, etc. Through the human-computer interaction interface, the user can input the information to be executed, that is, the terminal side receives the input information of the user. Among them, the input information can be text information, voice information, image information, video information, etc.

[0164] Step S220: Generate a first intermediate parameter according to the input information through the first attention layer.

[0165] In this embodiment, after receiving the input information, the input information is preprocessed to obtain data that can be executed by the TransformerLayer, and a first intermediate parameter is generated according to the data that can be executed through the first attention layer.

[0166] For text information, the preprocessing method is as follows: The original text is segmented into words or subword units (subword units), and the tokenization tool can include WordPiece, BPE (Byte Pair Encoding), etc. Each word or subword is mapped to a unique integer ID to form an ID sequence. This is usually achieved through a vocabulary (Vocabulary). Special tokens (such as [CLS], [SEP]) are added at the beginning and end of the sequence to help the model understand sentence boundaries or other specific task requirements. The sequence is padded (Padding) and truncated (Truncation) to ensure that all input sequences have the same length. Shorter sequences can reach a fixed length by padding (usually 0), while longer sequences need to be truncated. Since the Transformer does not have built-in position awareness capabilities, position encoding needs to be added to each position so that the model can utilize position information. Thus, the text information can be converted into data that can be executed by the Transformer Layer.

[0167] For voice information, the preprocessing method is as follows: convert the original audio signal into a spectrogram or other feature representations, such as Mel Frequency Cepstral Coefficients (MFCCs), Log Mel Spectrogram, etc. If the audio is too long, it can be divided into multiple shorter segments, and each segment is used as an independent input. Normalize the extracted features to improve the stability and convergence speed of the model. Ensure that all input segments have the same length through padding and truncation. Shorter segments can be padded to reach a fixed length, while longer segments need to be truncated. Similar to text processing, add positional encoding for each time step to retain the time order information. Thus, voice information can be converted into data that can be processed by the Transformer Layer.

[0168] For image information, the preprocessing method is as follows: adjust the image to a fixed size to unify the input size. If the image is too large, selectively crop the part of interest. Normalize the pixel values of the image. For example, scale the pixel values to the range of [0, 1] or [-1, 1]. Divide the image into multiple small patches, and each patch is used as an independent input unit. For example, the image can be divided into 16x16 patches. Flatten each patch and perform a linear transformation to convert it into a vector form suitable for Transformer processing. Add positional encoding for each patch to retain the spatial position information. Thus, image information can be converted into data that can be processed by the Transformer Layer.

[0169] For video information, the preprocessing method is as follows: extract key frames from the video, usually by sampling frames at a fixed time interval. Adjust each frame to a fixed size through resizing and cropping, and crop as needed. Normalize the pixel values of each frame. Divide each frame into multiple small patches, similar to the method in image processing. Flatten each patch and perform a linear transformation to convert it into a vector form suitable for Transformer processing. Consider consecutive frames as part of a time series and take into account the temporal relationship between frames. This can be achieved by stacking multiple consecutive frames or encoding the time information into each patch. Add positional encoding for each patch, and also consider the positional encoding in the time dimension to retain spatio-temporal information. Thus, video information can be converted into data that can be processed by the Transformer Layer.

[0170] Further, the data that the obtained Transformer Layer can execute is input into the first attention layer, and the first attention layer generates a first intermediate parameter according to the data that the Transformer Layer can execute.

[0171] Among them, the implementation manner for the first attention layer to obtain the first intermediate parameter can be implemented based on various existing methods. For example, it can be implemented based on an encoder architecture, a decoder architecture, or a combination of an encoder and a decoder architecture. The embodiments of the present invention do not limit this.

[0172] Step S230: Compress the first intermediate parameter through a structure compression layer to obtain a second intermediate parameter.

[0173] In this embodiment, the first intermediate parameter is provided to the structure compression layer, and the structure compression layer compresses the first intermediate parameter to obtain a second intermediate parameter.

[0174] Specifically, the first intermediate parameter is of dimension Ht*Dt, and the compressed second intermediate parameter is of dimension Ds. Among them, the compression method can be implemented based on various existing methods. For example, a linear transformation method can be used to map the space of dimension Ht*Dt to dimension Ds. Another example is that dimensionality reduction technology can also be used to map the space of dimension Ht*Dt to dimension Ds to capture and retain the most important features. Among them, the dimensionality reduction technology can adopt principal component analysis (PCA, Principal Component Analysis), autoencoders, etc.

[0175] Step S240: Obtain output information according to the second intermediate parameter through multiple second attention layers.

[0176] In this embodiment, the second intermediate parameter output by the structure compression layer is provided to multiple second attention layers, and the multiple second attention layers obtain output information according to the second intermediate parameter. Among them, the input of the first second attention layer is the second intermediate parameter, and the input of the subsequent second attention layers is the output of the previous second attention layer. Among them, the implementation manner for the second attention layer to obtain the output parameter can be implemented based on various existing methods. For example, it can be implemented based on an encoder architecture, a decoder architecture, or a combination of an encoder and a decoder architecture. The embodiments of the present invention do not limit this.

[0177] Among them, the output information can be set according to the tasks performed by the large language model. Among them, the tasks performed by the large language model can be text classification, sentiment analysis, machine translation, question answering system, automatic speech recognition (ASR), speech emotion recognition, image classification, object detection, video classification, action recognition, etc.

[0178] In an embodiment of the present invention, a pre-trained teacher model and an initialized student model are obtained. The initialized student model includes a first attention layer, a structure compression layer, and multiple second attention layers. The parameters of the first attention layer are the same as those of the attention layer of the first layer of the teacher model, and the parameters of the second attention layers are initialized parameters. Training samples and corresponding label information are obtained, and prediction information corresponding to the training samples is obtained through the initialized student model. The initialized parameters are optimized according to the label information and the prediction information to obtain a pre-trained student model. Thereby, information loss during the attention transfer process can be avoided, lossless transfer can be achieved, and the performance of the student model can be improved.

[0179] Further, an embodiment of the present invention also provides a large language model applied to the terminal side. The large language model includes a first attention layer, a structure compression layer, and multiple second attention layers. The specific structure can refer to Figure 1 the student model 2 shown. Among them, the first attention layer is used to generate first intermediate parameters according to the input information. The structure compression layer is used to compress the first intermediate parameters to obtain second intermediate parameters. The multiple second attention layers are used to obtain output information according to the second intermediate parameters. Among them, the specific implementation manner of the large language model applied to the terminal side can refer to the above attention transfer method and data processing method, which will not be elaborated in this embodiment of the present invention.

[0180] Figure 7 It is a schematic diagram of the attention transfer device according to an embodiment of the present invention. As Figure 7 shown, the attention transfer device includes: a model acquisition unit 71, a sample acquisition unit 72, a label acquisition unit 73, a prediction information acquisition unit 74, and a model training unit 75. Among them, the model acquisition unit 71 is used to obtain a pre-trained teacher model and an initialized student model. The initialized student model includes a first attention layer, a structure compression layer, and multiple second attention layers. The parameters of the first attention layer are the same as those of the attention layer of the first layer of the teacher model, and the parameters of the second attention layers are initialized parameters. The sample acquisition unit 72 is used to obtain training samples. The label acquisition unit 73 is used to obtain the label information corresponding to the training samples. The prediction information acquisition unit 74 is used to obtain the prediction information corresponding to the training samples through the initialized student model. The model training unit 75 is used to optimize the initialized parameters according to the label information and the prediction information to obtain a pre-trained student model.

[0181] In an embodiment of the present invention, a pre-trained teacher model and an initialized student model are obtained. The initialized student model includes a first attention layer, a structure compression layer, and a plurality of second attention layers. The parameters of the first attention layer are the same as those of the attention layer of the first layer of the teacher model, and the parameters of the second attention layer are initialized parameters. Training samples and corresponding label information are obtained, prediction information corresponding to the training samples is obtained through the initialized student model, and the initialized parameters are optimized according to the label information and the prediction information to obtain a pre-trained student model. Thereby, information loss during the attention transfer process can be avoided, lossless transfer can be achieved, and the performance of the student model can be improved.

[0182] Figure 8 is a schematic diagram of the data processing device according to an embodiment of the present invention. As Figure 8 shown, the data processing device includes an information receiving unit 81, a first parameter obtaining unit 82, a second parameter obtaining unit 83, and an output information obtaining unit 84. Among them, the information receiving unit 81 is used to receive input information. The first parameter obtaining unit 82 is used to generate a first intermediate parameter according to the input information through the first attention layer. The second parameter obtaining unit 83 is used to compress the first intermediate parameter through the structure compression layer to obtain a second intermediate parameter. The output information obtaining unit 84 is used to obtain output information according to the second intermediate parameter through a plurality of second attention layers.

[0183] In an embodiment of the present invention, a pre-trained teacher model and an initialized student model are obtained. The initialized student model includes a first attention layer, a structure compression layer, and a plurality of second attention layers. The parameters of the first attention layer are the same as those of the attention layer of the first layer of the teacher model, and the parameters of the second attention layer are initialized parameters. Training samples and corresponding label information are obtained, prediction information corresponding to the training samples is obtained through the initialized student model, and the initialized parameters are optimized according to the label information and the prediction information to obtain a pre-trained student model. Thereby, information loss during the attention transfer process can be avoided, lossless transfer can be achieved, and the performance of the student model can be improved.

[0184] Figure 9 is a schematic diagram of the electronic device according to an embodiment of the present invention. In this embodiment, the electronic device 9 includes a server, a terminal, etc. As Figure 9 shown, the electronic device 9: includes at least one processor 91; and, a memory 92 communicatively connected to at least one processor 91; and, a communication component 93 communicatively connected to the scanning device, and the communication component 93 receives and transmits data under the control of the processor 91; wherein, the memory 92 stores instructions executable by at least one processor 91, and the instructions are executed by at least one processor 91 to implement the above-mentioned attention transfer method and data processing method.

[0185] Specifically, the electronic device includes: one or more processors 91 and a memory 92. Figure 9 Taking one processor 91 as an example. The processor 91 and the memory 92 can be connected through a bus or other means. Figure 9 Taking the connection through a bus as an example. The memory 92, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The processor 91 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in the memory 92, that is, implementing the above-mentioned attention migration method and data processing method.

[0186] The memory 92 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store an option list, etc. In addition, the memory 92 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 92 may optionally include a memory remotely set relative to the processor 91, and these remote memories can be connected to an external device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0187] One or more modules are stored in the memory 92, and when executed by one or more processors 91, they execute the attention migration method and the data processing method in any of the above method embodiments.

[0188] The above product can execute the method provided in the embodiments of the present application, and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiments of the present application.

[0189] In the embodiment of the present invention, by obtaining a pre-trained teacher model and an initialized student model, the initialized student model includes a first attention layer, a structure compression layer, and a plurality of second attention layers, the parameters of the first attention layer are the same as the parameters of the first attention layer of the teacher model, the parameters of the second attention layer are initialized parameters, obtaining training samples and corresponding label information, obtaining prediction information corresponding to the training samples through the initialized student model, and optimizing the initialized parameters according to the label information and the prediction information to obtain a pre-trained student model. Thereby, information loss during the attention migration process can be avoided, lossless migration can be achieved, and the performance of the student model can be improved.

[0190] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, which is used for a computer to execute the above-mentioned partial or all method embodiments.

[0191] That is, those skilled in the art can understand that all or part of the steps in implementing the above method embodiments can be completed by instructing relevant hardware through a program. This program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0192] The foregoing are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for attention transfer, characterized in that: The method comprises: Obtain a pre-trained teacher model and an initialized student model, wherein the initialized student model includes a first attention layer, a structure compression layer, and a plurality of second attention layers, wherein parameters of the first attention layer are the same as parameters of the attention layer of the first layer of the teacher model, and parameters of the second attention layer are initialization parameters; Get training samples; Obtaining label information corresponding to the training sample; Acquire prediction information corresponding to the training sample through the initialized student model; The initialization parameters are optimized according to the label information and the prediction information to obtain a pre-trained student model.

2. The method according to claim 1, characterized in that The label information corresponding to the training samples is a pre-labeled real label.

3. The method according to claim 1, characterized in that The obtaining of label information corresponding to the training sample includes: The label information is obtained according to the training samples through the pre-trained teacher model.

4. The method according to claim 1, characterized in that: The obtaining prediction information corresponding to the training sample through the initialized student model includes: Generate a first intermediate parameter according to the training sample through a first attention layer of the initialized student model; Compressing the first intermediate parameter through the structure compression layer of the initialized student model to obtain a second intermediate parameter; The prediction information is obtained according to the second intermediate parameters through the multiple second attention layers.

5. The method according to claim 4, characterized in that The obtaining of the prediction information according to the second intermediate parameter through the multiple second attention layers: Get the initialization parameters of each second attention layer; Performing linear transformation and ReLU activation on the attention matrix of the attention layer of the first layer of the teacher model according to the initialization parameters to obtain the attention matrices of each second attention layer; Generate an attention coefficient according to the attention matrix of the second attention layer by a normalized exponential function; The prediction information is obtained according to the attention coefficient and the second intermediate parameter.

6. The method according to claim 1, characterized in that The optimizing the initialization parameters according to the label information and the prediction information to obtain the pre-trained student model comprises: Calculate a loss value according to the label information and the prediction information; The initialization parameters are optimized according to the loss value to obtain a pre-trained student model.

7. The method according to claim 6, characterized in that The loss value is calculated by the following formula: Among them, L(u) is the loss value, u i is the target word, u i-k ,…,u i-1 is the context word before the target word, and θ is a learnable parameter.

8. A data processing method, characterized in that: The method comprises: Receive input information; Generate a first intermediate parameter according to the input information through a first attention layer; compressing the first intermediate parameter through a structure compression layer to obtain a second intermediate parameter; Output information is obtained according to the second intermediate parameters through multiple second attention layers.

9. A large language model applied on a terminal side, characterized in that: The large language model includes: A first attention layer, used to generate a first intermediate parameter according to input information; A structure compression layer, used for compressing the first intermediate parameter to obtain a second intermediate parameter; A plurality of second attention layers are used to obtain output information according to the second intermediate parameters.

10. An attention transfer device, characterized in that: The device comprises: A model acquisition unit, used to acquire a pre-trained teacher model and an initialized student model, wherein the initialized student model includes a first attention layer, a structure compression layer, and a plurality of second attention layers, wherein parameters of the first attention layer are the same as parameters of the attention layer of the first layer of the teacher model, and parameters of the second attention layer are initialization parameters; A sample acquisition unit, used for acquiring training samples; A label acquisition unit, used to acquire label information corresponding to the training sample; A prediction information acquisition unit, used to acquire prediction information corresponding to the training sample through the initialized student model; A model training unit is used to optimize the initialization parameters according to the label information and the prediction information to obtain a pre-trained student model.

11. A data processing device, characterized in that: The device comprises: An information receiving unit, used for receiving input information; A first parameter acquisition unit, configured to generate a first intermediate parameter according to the input information through a first attention layer; A second parameter acquisition unit, configured to compress the first intermediate parameter through a structure compression layer to acquire a second intermediate parameter; An output information acquisition unit is used to acquire output information according to the second intermediate parameters through multiple second attention layers.

12. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 8.

13. A computer program product, comprising a computer program, characterized in that: When the computer program is executed on a computer, the computer executes the method according to any one of claims 1 to 8.

14. A computer-readable storage medium storing computer program instructions, characterized in that: The computer program instructions implement the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Cited By

  • Large language model illusion reduction method for entropy-triggered visual attention backtracking

    CN121982494A

  • A Large Language Model Illusion Reduction Method Based on Entropy-Triggered Visual Attention Backtracking

    CN121982494B