Natural language processing method, electronic device, storage medium and program product

The student network model of zero-vector columns is generated through the knowledge distillation mechanism, which solves the problems of large storage requirements and low computing efficiency of the Transformer architecture model, and achieves savings in storage space and improving computing speed.

CN120234409AActive Publication Date: 2025-07-01INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510714845.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

The existing Transformer architecture model has large storage requirements and low inference operation efficiency, especially in resource-constrained environments.

Method used

The student network model consisting of multiple zero vector column weight matrices is generated through a knowledge distillation mechanism, the storage of zero vector columns is omitted, and non-zero vector columns are used for calculations only in the inference operation.

Benefits of technology

It reduces storage demand, reduces calculation amount, improves inference operation speed, and improves overall computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234409A_ABST
    Figure CN120234409A_ABST
Patent Text Reader

Abstract

The invention discloses a natural language processing method, electronic equipment, a storage medium and a program product, and relates to the technical field of computers, a student network model comprising a plurality of zero vector column weight matrixes is generated through a knowledge distillation mechanism, so that zero vector columns in the weight matrixes can be omitted when the student network model is stored, and the storage efficiency of the student network model is improved. The storage requirement is reduced, a large amount of storage space is further saved, when reasoning operation is carried out on the text data through the stored student network model, the zero vector column does not participate in operation, the calculation amount is reduced, the reasoning operation speed is increased, and the overall operation efficiency is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a natural language processing method, an electronic device, a storage medium, and a program product. Background Art

[0002] In the context of the current booming development of artificial intelligence technologies, the Transformer architecture model (a deep learning model architecture for natural language processing and other sequence-to-sequence tasks) has demonstrated excellent performance and broad application prospects in numerous fields. However, with the continuous expansion of model scale and the increasing complexity of application scenarios, its storage requirements have become a key issue that urgently needs to be addressed. In the related model storage method, during the storage process, the weight matrix of the student model after knowledge distillation needs to be completely stored, which often requires a large amount of storage space. This not only poses high requirements on storage devices but also restricts the application and deployment of the model in resource-constrained environments to a certain extent. Additionally, during the model inference and calculation process, the entire weight matrix of the stored student model participates in the operation, resulting in a large amount of calculation and low operation efficiency. Summary of the Invention

[0003] This application provides a natural language processing method, an electronic device, a storage medium, and a program product to at least solve the problems of large storage resource occupation and low inference operation efficiency in the related technologies.

[0004] This application provides a natural language processing method, including: Obtaining a pre-constructed teacher network model; Generating a target student network model based on the teacher network model and a knowledge distillation mechanism, where the target student network model includes at least one weight matrix, and at least one zero vector column is included in the weight matrix; In response to receiving a natural language processing request, obtaining text data and dividing the text data into multiple tokens; Inputting the tokens into the target student network model, and performing inference calculation on the tokens based on the weight matrix including the zero vector column in the target student network model to determine the output result corresponding to the text data.

[0005] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the following steps of the natural language processing method when executing the computer program: Obtaining a pre-constructed teacher network model; Generating a target student network model based on the teacher network model and a knowledge distillation mechanism, where the target student network model includes at least one weight matrix, and at least one zero vector column is included in the weight matrix; In response to receiving a natural language processing request, obtain text data and divide the text data into multiple tokens; Input the tokens into the target student network model, and based on the weight matrix including zero vector columns in the target student network model, perform inference calculations on the tokens to determine the output result corresponding to the text data.

[0006] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the following steps of the natural language processing method are implemented: Obtain a pre-constructed teacher network model; Based on the teacher network model and the knowledge distillation mechanism, generate a target student network model. The target student network model includes at least one weight matrix, and at least one zero vector column is included in the weight matrix; In response to receiving a natural language processing request, obtain text data and divide the text data into multiple tokens; Input the tokens into the target student network model, and based on the weight matrix including zero vector columns in the target student network model, perform inference calculations on the tokens to determine the output result corresponding to the text data.

[0007] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps of the natural language processing method are implemented: Obtain a pre-constructed teacher network model; Based on the teacher network model and the knowledge distillation mechanism, generate a target student network model. The target student network model includes at least one weight matrix, and at least one zero vector column is included in the weight matrix; In response to receiving a natural language processing request, obtain text data and divide the text data into multiple tokens; Input the tokens into the target student network model, and based on the weight matrix including zero vector columns in the target student network model, perform inference calculations on the tokens to determine the output result corresponding to the text data.

[0008] Through the knowledge distillation mechanism, this application generates a student network model including a weight matrix with multiple zero vector columns, so that when storing the student network model, the zero vector columns in the weight matrix can be omitted, reducing the storage requirement, thereby saving a large amount of storage space. When performing inference operations on text data through the stored student network model, the zero vector columns do not participate in the operations, reducing the computational amount and improving the inference operation speed, thereby improving the overall operation efficiency. Description of the Drawings

[0009] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0010] Figure 1 This is an application environment diagram for a natural language processing method provided by an embodiment of the present application; Figure 2 This is an overall process schematic diagram of a natural language processing method provided by an embodiment of the present application; Figure 3 This is a workflow schematic diagram of a query embedding generation module for a natural language processing method provided by an embodiment of the present application; Figure 4 This is a workflow schematic diagram of a key embedding generation module for a natural language processing method provided by an embodiment of the present application; Figure 5 This is a workflow schematic diagram of a value embedding generation module for a natural language processing method provided by an embodiment of the present application; Figure 6 This is a schematic diagram of the overall structure of a single-head attention module for a natural language processing method provided by an embodiment of the present application; Figure 7 This is a workflow schematic diagram of a feed-forward neural network layer for a natural language processing method provided by an embodiment of the present application; Figure 8 This is a schematic diagram of the distillation loss process between a teacher network model and a student network model for a natural language processing method provided by an embodiment of the present application; Figure 9 This is an internal structure diagram of an electronic device in an embodiment. Detailed implementation manners

[0011] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0012] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0013] It should be noted that the terms "S1", "S2", etc. are only used for the purpose of describing steps, and do not particularly refer to the meaning of order or sequence, nor are they used to limit this application. They are only used to conveniently describe the method of this application and should not be understood as indicating the order of steps. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions is contradictory or cannot be achieved, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0014] According to the background technology, as the model scale continues to expand and the application scenarios become increasingly complex, its storage requirement has become a key problem to be solved urgently. On the one hand, the related model storage methods often require a large amount of storage space, which not only poses high requirements on storage devices, but also restricts the application and deployment of models in resource-constrained environments to a certain extent. To address this challenge, low-bit quantization methods have emerged. By performing low-bit quantization on the numerical values in the model, such as quantizing from 32-bit floating point to 8-bit or even lower bits, the space required to store each parameter can be significantly reduced. This method effectively compresses the storage size of the model without seriously affecting the model performance, enabling larger-scale models or more model instances to be stored under the same storage resources. On the other hand, low-rank matrix representation has also become an important means to reduce storage requirements. In many cases, the weight matrices in the model have certain redundancy and structure. Using low-rank matrix representation technology, the originally high-dimensional and complex weight matrices can be decomposed into a combination form of low-rank matrices. In this way, only the relevant information of these low-rank matrices needs to be saved during storage, reducing the storage overhead compared to directly storing the original weight matrices. However, the above technologies still need to save the complete weight matrices of the student network model, which requires a large amount of storage space, and during the subsequent inference operation process, the entire weight matrix of the student model participates in the operation, resulting in a large amount of computation.

[0015] To solve the above technical problems, the present application provides a natural language processing method, an electronic device, a storage medium, and a program product. Through a knowledge distillation mechanism, a student network model including multiple zero-vector column weight matrices is generated, so that when storing the student network model, the zero-vector columns in the weight matrix can be omitted, reducing the storage requirement and thus saving a large amount of storage space. When performing inference operations on text data through the stored student network model, the zero-vector columns do not participate in the operations, reducing the computational amount and improving the inference operation speed, thereby enhancing the overall operation efficiency.

[0016] To enable those skilled in the art of the present technology to better understand the solution of the present application, the following further elaborates on the present application in conjunction with the accompanying drawings and specific embodiments.

[0017] The natural language processing method provided by the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the data processing platform set on the server 104 through the network. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0018] To enable those skilled in the art of the present technology to better understand the solution of the present application, an explanation of the Transformer architecture model involved in the present application is provided, including: (1) Query embedding generation module; As Figure 3 shown, the query embedding generation module converts its input embedding ( ) into a query embedding ( ). Assuming that the input embedding is a column matrix with rows and 1 column, that is, ; the query embedding is a column matrix with rows and 1 column, that is, ; the weight matrix contained in the query embedding generation module is a matrix with rows and columns, that is, , then the process of the input embedding passing through the query embedding generation module to output the query embedding can be formulated by the formula: .

[0019] (2) Key embedding generation module; As Figure 4 shown, the key embedding generation module converts its input embedding ( ) into a key embedding ( ), assuming the input embedding is a column matrix with rows and 1 column, that is ; the key embedding is a column matrix with rows and 1 column, that is ; the weight matrix contained in the key embedding generation module is a matrix with .

[0020] (3) Value embedding generation module; As Figure 5 shown, the value embedding generation module converts its input embedding ( ) into a value embedding ( ), assuming the input embedding is a column matrix with rows and 1 column, that is ; the value embedding is a column matrix with rows and 1 column, that is ; the weight matrix contained in the value embedding generation module is a matrix with .

[0021] Among them, as Figure 6 shown, inside the model based on the Transformer structure, the query embedding generation module, the key embedding generation module, and the value embedding generation module are always used as a whole, and they share the same input embedding ( ). When using single-head attention, the input embedding ( ) is separately input into the query embedding generation module, the key embedding generation module, and the value embedding generation module; in addition, when using multi-head attention, there are multiple groups of query embedding generation modules, key embedding generation modules, and value embedding generation modules, and the input embedding ( ) is separately input into each group of query embedding generation modules, key embedding generation modules, and value embedding generation modules.

[0022] It should be noted that: the above module input embedding (X) can be the token embedding generated by the token embedding encoder when the token is input into the model, or the embedding inside the model. Taking text input as an example, when text data is input into the model, the text data is segmented into several tokens, and then each token is input into the token embedding encoder to generate a token embedding. When represents the above token embedding, then under normal circumstances ; when represents the internal embedding of the model, , it should be noted that and 's numerical relationship does not affect the application scope of this application.

[0023] (4) Feed-forward neural network layer (also called feed-forward neural network module); As Figure 7 shown, the feed-forward neural network layer converts its input embedding ( ) into an output embedding ( ). Assuming the input embedding is a row, 1-column column matrix, that is ; the output embedding is a row, 1-column column matrix, that is ; the weight matrix contained in the feed-forward neural network layer is a row, column matrix, that is . Then, . In this embodiment, the feed-forward neural network layer refers to the feed-forward neural network layer inside each Transformer module, and its input embedding and output embedding have the same dimension, that is .

[0024] As Figure 2 shown, the embodiment of this application provides a natural language processing method. Taking the example that this method is applied to the Figure 1 terminal, it includes the following steps: S1: Obtain a pre-constructed teacher network model.

[0025] It should be noted that the teacher model refers to a large and complex deep neural network model with an extremely large number of pre-trained parameters during the knowledge distillation process. The teacher model is usually trained on a large dataset, has a high complexity and a large number of parameters, and can exhibit good performance on specific tasks. In this application, the teacher model refers to a model based on the Transformer architecture , also known as the source model.

[0026] S2: Based on the teacher model and the knowledge distillation mechanism, generate a target student network model. The target student network model includes at least one weight matrix, and at least one column of zero vectors is included in the weight matrix.

[0027] It should be noted that the knowledge distillation mechanism refers to the transfer of "knowledge" from a teacher model with an extremely large number of parameters to a lightweight student network model. After the knowledge distillation process, the student network model can exhibit performance comparable to that of the teacher network model. In this application, the target student network model refers to a model based on the Transformer architecture , the model and the model have the same network structure, that is, the dimensions of the weight matrices at corresponding positions are the same. This application uses a neural network training method to distill the target model from the source model , and the neural network training loss function is , , consists of two parts, namely the distillation loss function and the task-related loss function , is a function with the distillation loss function and the task-related loss function as independent variables. , and multiple columns of zero vectors are included in the weight matrix of the target student network model obtained after being processed by the knowledge distillation mechanism. Among them, both the teacher network model and the student network model are large language models, such as the large language model of Question-Answer. When a question is input into the model, the output is the answer corresponding to the question.

[0028] S3: In response to receiving a natural language processing request, obtain text data and divide the text data into multiple tokens.

[0029] It should be noted that natural language processing requests are generally issued by the task node that stores the target student model, and the requests include text data. For example, a token is to split the text data into smaller units, which may be words, characters, or sub-words. Tokens are the most basic units that language models such as GPT (Generative Pre-trained Transformer) operate on during training and inference.

[0030] S4: Input the token into the target student network model, and based on the weight matrix including zero vector columns in the target student network model, perform inference calculations on the token to determine the output result corresponding to the text data.

[0031] It should be noted that when performing inference calculations on the token, the zero vector columns do not participate in the inference calculation process. The text data is the input problem-related data, and the output result is the answer corresponding to the problem.

[0032] In the above embodiment, through the knowledge distillation mechanism, a student network model including multiple weight matrices of zero vector columns is generated, so that when storing the student network model, the zero vector columns in the weight matrix can be omitted, reducing the storage requirement, and thus saving a large amount of storage space. When performing inference operations on the text data through the stored student network model, the zero vector columns do not participate in the operation, reducing the computational amount, improving the inference operation speed, and thus improving the overall operation efficiency.

[0033] In some specific embodiments, generating the target student network model based on the teacher network model and the knowledge distillation mechanism includes: Based on the teacher network model, construct a first student network model, where represents the weight matrix in the first student network model, , represents the first diagonal matrix, represents the randomly initialized matrix. Among them, the teacher network model and the first student network model have the same network structure. The first diagonal matrix refers to a matrix in which the element values in other positions are all zero values except for the element values on the main diagonal. The element values in the randomly initialized matrix will be updated during the training process; Based on the teacher network model and the first student network model, construct a training loss function. The training loss function is a function with the distillation loss function and the task-related loss function as independent variables; Obtain training data, where the training data refers to text data. The training data of this application collects a large-scale text data set, including hundreds of millions of statements, including statements in multiple languages. Exemplarily, the proportion of Chinese statements is 30%, the proportion of English statements is 60%, and the proportion of other languages is 10%. During the model training process, text data is sampled from the large-scale text data set in a random sampling manner. The language type proportion is an exemplary value, and using other numerical ratios does not affect the effect of this application. Input this text data into the large prediction model to generate intermediate state values during model operation (such as query embedding, key embedding, and value embedding, etc.), as well as the final output text. To avoid the order of training samples in the training data set D affecting the performance of the model, at the beginning of model training, all samples in the training data set D are randomly arranged. During each round of training, the model training algorithm reads a batch of samples, such as 1024 samples; Train the first student network model based on the training loss function and training data to update the weight matrix in the first student network model; In response to the current training cycle being the preset training cycle, end the training to obtain the second student network model, where represents the weight matrix in the second student network model, , represents the randomly initialized matrix after training update; Furthermore, sort the diagonal elements in the first diagonal matrix in ascending order, for example, sort the diagonal elements in the first diagonal matrix in descending order; Select multiple element values from the sorting result according to the preset ratio, and set the other unselected element values to zero to generate the second diagonal matrix , where the preset ratio can be set according to actual needs, and its value range is 30% - 70%. The preferred value can be set to 50%, that is, retain the first 50% of the diagonal elements in the sorting result, and set the other element values in the first diagonal matrix to zero to generate the second diagonal matrix ; Based on the product of the second diagonal matrix and , determine that the weight matrix in the third student network model is ; Fix the zero vector columns in the weight matrix of the third student network model, and set the element value states in the remaining columns to be learnable states; Train the third student network model based on the task-related loss function and the text data corresponding to the target task node, so as to update the element values in the remaining columns, where the target task node is the node corresponding to a more specific task scenario; Generate a target student network model in response to the completion of training.

[0034] In the above embodiment, by constructing a teacher network model and a student network model with the same network structure, and training the student network model by setting a diagonal matrix and a training loss function, the obtained student network model presents a unique characteristic in the dimension of the weight matrix, that is, there are a large number of columns with element values of zero in the weight matrix. Based on this, the space required for storing the student network model can be greatly reduced, and the hardware storage cost can be reduced. On the other hand, during the model inference operation process, the existence of zero-value columns can effectively reduce the amount of calculation, significantly improve the inference speed, and thus improve the overall operation efficiency.

[0035] In some specific embodiments, after generating the target student network model, the method further includes: Deploy the target student network model to the target task node, that is, the target student network model is deployed to the corresponding task node through the text data of which task node is used to train the model. When deploying and storing the target student network model, there is no need to save the zero vector columns, and only the position information of these columns needs to be recorded.

[0036] In the above embodiment, by deploying a more targeted student network model to the corresponding task node, the requests of the target task node can be processed efficiently and accurately, and the task processing efficiency can be improved.

[0037] In some specific embodiments, the above method further includes: Obtain the text data corresponding to the historical tasks of multiple task nodes; Calculate the correlation between two task nodes based on the text data, including: Obtain the first text data corresponding to the historical task of the first task node and the second text data corresponding to the historical task of the second task node; Based on the similarity algorithm, calculate the similarity between the first text data and the second text data. Among them, the similarity algorithm can be an algorithm based on string matching, an algorithm based on the bag-of-words model, etc. for calculating text similarity. It is a commonly used algorithm, and the specific calculation process will not be elaborated here; Based on the similarity, evaluate the correlation between the first task node and the second task node, that is, the higher the similarity, the higher the correlation; In response to the similarity being greater than a preset threshold, it is determined that the first task node and the second task node have a high correlation. Based on this, the same student network model is deployed on the first task node and the second task node to process the corresponding tasks, where the preset threshold can be set according to actual needs, such as 90% or the like.

[0038] In the above embodiment, by calculating the correlation between task nodes and deploying the same student network model on two task nodes with high correlation, the number of model training times can be reduced, and the overall working efficiency of the system can be improved.

[0039] In some specific embodiments, based on the training loss function and training data, training the first student network model to update the weight matrix in the first student network model includes: Fix the weight matrix parameters of the teacher network model, randomly initialize the element values and other parameters in the first student network model, where the other parameters refer to the parameters used to train the network model. Based on the Gaussian distribution, perform random initialization processing on the element values in the first student network model. After processing, the mean value of the weight matrix element values is 0, and the standard deviation is 0.01; Set the relevant parameters of the stochastic gradient descent algorithm, and define the current training epoch as and the total number of training epochs as In response to when perform the following operations, where the stochastic gradient descent algorithm is the Adam algorithm, and the relevant parameters of the stochastic gradient descent algorithm and the total number of training epochs can both be set according to actual needs. Exemplarily, the total number of training epochs is 30, and the relevant parameters of the stochastic gradient descent algorithm are that the learning rate is and the momentum coefficient , , , the learning rate strategy is the cosine strategy, and the number of samples in a mini - batch is B = 1024; Increment the current training epoch by 1, that is ; ; Randomly shuffle the order of samples in the training data, and select the target batch of training samples from the training data, that is, randomly shuffle the order of samples in the data set and select a batch B of training samples from the data set ; According to the training samples and the relevant parameters of the stochastic gradient descent algorithm, perform the training operation, calculate the training loss value, and update the weight matrix in the first student network model using the stochastic gradient descent algorithm; Repeat the steps of sample selection and weight matrix update. In response to all training samples in the training data being utilized, end the current training epoch; End the training in response to the current training cycle being a preset training cycle to obtain a second student network model.

[0040] In the above implementation, the model can be trained through the above model training method to improve the performance of the model.

[0041] In some specific implementations, constructing the training loss function includes: Construct the training loss function based on the distillation loss function and the task-related loss function.

[0042] Among them, the training loss function includes: ; Among them, represents the training loss value, represents the task-related loss value, represents a constant parameter, represents the distillation loss value.

[0043] The distillation loss function includes: ; ; Among them, represents the distillation loss value, represents the number of weight matrices in the teacher network model, represents the matrix and the matrix the distillation loss value between them, represents a random matrix, represents a diagonal matrix.

[0044] The matrix and the matrix The calculation method of the distillation loss value between them includes: ; Among them, represents norm, represents the matrix F norm, represents a constant parameter, represents the vector composed of the diagonal elements of the diagonal matrix.

[0045] The task-related loss function includes: ; Among them, represents the training loss value, represents the number of tokens, represents the vector the th element of, Denote the vector 's th element, denote -dimensional one-hot vector, representing the category corresponding to the true token at position , represent the predicted -dimensional probability distribution vector.

[0046] Specifically, the definition of the distillation loss function includes: as Figure 8 shown, without loss of generality, assume that any weight matrix in the source model is , and the weight matrix in the target model corresponding to is . They have the same input. Here, the weight matrix and the weight matrix represent the weight matrix of any embedding generation module, the weight matrix of the key embedding generation module, the weight matrix of the value embedding generation module, or the weight matrix of the feed-forward neural network layer. This application requires that for any weight matrix in the source model and the corresponding weight matrix in the target model , when their input data is the same, the outputs generated are the same or similar. For any weight matrix in the source model and the corresponding weight matrix in the target model , on the premise that the input data is consistent, the output results of the two need to strictly meet the requirements of consistency or approximation, that is, the outputs generated by the two are either exactly the same or infinitely close, so as to ensure the accuracy and effectiveness in the process of model knowledge transfer, making the weight matrix and the weight matrix have outputs that are infinitely close or the same. The method is as follows: Assume that the input data of the weight matrix and the weight matrix is the same, defined as the column vector . The column vector is input into the weight matrix , and the output column vector is obtained. The column vector is input into the weight matrix , and the output column vector is obtained. In the case of the same input , the outputs are the same or approaching, that is, it is required that and are the same or close. In this application, is used to measure the similarity between and When it is zero, is the same as . The smaller is, the closer (more similar) is to . The optimization goal of the model training algorithm is to minimize , where represents the square of the L2 norm of the difference between and represents the norm; ; Assume that the element values of follow a Gaussian distribution. Minimizing the square of the L2 norm of the difference between and can be transformed into minimizing the square of the F norm of the difference between ; Based on this, the optimization goal is transformed into , where represents minimization, is the weight matrix of the source model. The optimization goal is to change so as to minimize . From the perspective of the entire optimization process, the weight matrix of the source model remains fixed, and the weight matrix of the optimized target model is optimized; while minimizing , this application also requires that all the elements of some columns of are zero. Then, during the process of saving the model, these columns do not need to be saved, and only the position information of these columns needs to be recorded. To achieve this optimization goal, this application further represents as the multiplication of two matrices, where ; To achieve the optimization goal that all the elements of some columns of are zero, some of the elements on the main diagonal of the diagonal matrix need to be zero. If some of the elements on the main diagonal of the diagonal matrix and Calculated All elements of some columns in are zero. To achieve the above goal, extract the diagonal matrix of the diagonal elements and integrate them into a column vector . Then, for the vector of norm as a constraint is added to the following formula. By adding the above constraint, the effect of making some diagonal elements in the diagonal matrix zero can be achieved. Based on the above analysis and demonstration, the loss function for training the weight matrix is defined as : ; where is a constant coefficient. Preferably, ; In the above expression, the norm is used as a constraint to find the optimal sparse term. However, the minimization problem of the norm is an NP-hard problem. Therefore, the optimization problem is relaxed to a higher-dimensional norm problem, such as the norm; ; The above analyzed the distillation loss between a set of matrices and . Expanding the above analysis, the expression of the distillation loss function between models can be obtained as follows: ; where , represents the number of weight matrices in the source model , and is also the number of weight matrices in the target model .

[0047] Furthermore, the task-related loss function is defined as follows: For knowledge distillation of a large language model, the task is to predict the next token given a sequence of tokens. Then, the task loss function adopts the cross-entropy loss function. Suppose there is a sequence containing tokens. For each position , the model needs to predict the probability distribution that the next token belongs to the words in the vocabulary. Let be a -dimensional one-hot vector representing the category corresponding to the true token at position . is the dimensional probability distribution vector, then the cross-entropy loss function can be expressed as: ; where represents the training loss value, represents the number of tokens, represents the th element of the vector represents the th element of the vector represents dimensional one-hot vector, representing the category corresponding to the true token at position , represents the predicted dimensional probability distribution vector.

[0048] Based on the above distillation loss function and task-related loss function, the overall loss function for model training can be constructed as: ; where is a function with the distillation loss function and the task-related loss function as independent variables. In this application, adopts a linear function, as follows: ; where is a constant coefficient used to balance the and contributions. Preferably, , better accuracy can be achieved by adjusting .

[0049] In the above embodiment, the overall training loss function is constructed by the constructed task-related loss function and distillation loss function to calculate the loss value, and then the parameters of the large language model to be trained are updated, which can improve the performance of the large language model.

[0050] In some specific embodiments, in response to receiving a natural language processing request, text data is obtained, and the text data is divided into multiple tokens, including: Receiving a natural language processing request sent by a target task node; In response to receiving the natural language processing request, parsing the natural language processing request to obtain the text data corresponding to the natural language processing request; According to the partitioning rules, the text data is partitioned to obtain multiple tokens corresponding to the text data. Among them, the partitioning rules can be set according to actual needs. For example, a series of rules can be formulated according to the grammar rules and lexical characteristics of the language for partitioning. Exemplarily, in Chinese, single Chinese characters and words can be recognized according to a dictionary, and some word-formation rules can also be combined, such as the patterns of "prefix + root" and "root + suffix".

[0051] In the above embodiment, by partitioning the text data into tokens, the learning efficiency of the model can be improved, and the generalization ability of the model can be enhanced.

[0052] In the above natural language processing method, the method includes: obtaining a pre-constructed teacher network model; generating a target student network model based on the teacher network model and the knowledge distillation mechanism. The target student network model includes at least one weight matrix, and at least one column of zero vectors is included in the weight matrix; in response to receiving a natural language processing request, obtaining text data and partitioning the text data into multiple tokens; inputting the tokens into the target student network model, and based on the weight matrix including columns of zero vectors in the target student network model, performing inference calculation on the tokens to determine the output result corresponding to the text data. In this application, through the knowledge distillation mechanism, a student network model including a weight matrix with multiple columns of zero vectors is generated. In terms of storage, by omitting the columns of zero vectors in the weight matrix of the student network model, the storage cost is significantly reduced, and the storage capacity of complex models can be reduced by up to 50% at most, which is more suitable for devices with limited storage resources. In terms of operation, the columns of zero vectors do not participate in the calculation during inference, and the amount of calculation is reduced to 30%-50% of the traditional method, and the inference speed is significantly improved, greatly enhancing the real-time performance of the system. In scenarios with high requirements for response speed such as autonomous driving and intelligent security, data can be quickly processed and accurate decisions can be made, and the application prospect is broad.

[0053] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0054] It should be understood that although Figures 2 - 8 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figures 2 - 8At least a part of the steps therein may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed and completed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turns with at least a part of other steps or sub-steps or stages of other steps.

[0055] In one embodiment, an electronic device is provided. The electronic device may be a terminal, and its internal structure diagram may be as Figure 9 shown. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a natural language processing method is implemented. The display screen of the electronic device may be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, a touchpad, or a mouse, etc.

[0056] Those skilled in the art can understand that Figure 9 the structure shown in

[0057] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. S1: Obtain a pre-built teacher network model; S2: Based on the teacher network model and the knowledge distillation mechanism, generate a target student network model. The target student network model includes at least one weight matrix, and at least one zero vector column is included in the weight matrix; S3: In response to receiving a natural language processing request, obtain text data and divide the text data into multiple tokens; S4: Input the tokens into the target student network model, and based on the weight matrix including the zero vector column in the target student network model, perform inference calculation on the tokens to determine the output result corresponding to the text data.

[0058] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in the embodiment of the natural language processing method during runtime, including: S1: Obtain a pre-constructed teacher network model; S2: Generate a target student network model based on the teacher network model and the knowledge distillation mechanism. The target student network model includes at least one weight matrix, and at least one column of zero vectors is included in the weight matrix; S3: In response to receiving a natural language processing request, obtain text data and divide the text data into multiple tokens; S4: Input the tokens into the target student network model, and perform inference calculation on the tokens based on the weight matrix including a column of zero vectors in the target student network model to determine the output result corresponding to the text data.

[0059] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), external hard drives, magnetic disks, or optical discs that can store computer programs.

[0060] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in the embodiment of the natural language processing method, including: S1: Obtain a pre-constructed teacher network model; S2: Generate a target student network model based on the teacher network model and the knowledge distillation mechanism. The target student network model includes at least one weight matrix, and at least one column of zero vectors is included in the weight matrix; S3: In response to receiving a natural language processing request, obtain text data and divide the text data into multiple tokens; S4: Input the tokens into the target student network model, and perform inference calculation on the tokens based on the weight matrix including a column of zero vectors in the target student network model to determine the output result corresponding to the text data.

[0061] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the steps in the embodiment of the natural language processing method, including: S1: Obtain a pre-constructed teacher network model; S2: Based on the teacher network model and the knowledge distillation mechanism, generate a target student network model, where the target student network model includes at least one weight matrix, and the weight matrix includes at least one column of zero vectors; S3: In response to receiving a natural language processing request, obtain text data and divide the text data into multiple tokens; S4: Input the tokens into the target student network model, and based on the weight matrix including a column of zero vectors in the target student network model, perform inference calculations on the tokens to determine the output result corresponding to the text data.

[0062] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0063] The above has introduced in detail a natural language processing method, device, electronic device, and storage medium provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A natural language processing method, characterized in that, The method includes: Obtaining a pre-built teacher network model; Generating a target student network model based on the teacher network model and a knowledge distillation mechanism, where the target student network model includes at least one weight matrix, and at least one zero vector column is included in the weight matrix; In response to receiving a natural language processing request, obtaining text data and dividing the text data into multiple tokens; Inputting the tokens into the target student network model, and performing inference calculation on the tokens based on the weight matrix including the zero vector column in the target student network model to determine the output result corresponding to the text data.

2. The natural language processing method according to claim 1, wherein Generating a target student network model based on the teacher network model and a knowledge distillation mechanism includes: Based on the teacher network model, a first student network model is constructed, where represents the weight matrix in the first student network model, , represents a first diagonal matrix, represents a randomly initialized matrix; Constructing a training loss function based on the teacher network model and the first student network model, where the training loss function is a function with a distillation loss function and a task-related loss function as independent variables; Obtaining training data; Training the first student network model based on the training loss function and the training data to update the weight matrix in the first student network model; In response to the current training cycle being a preset training cycle, end the training to obtain a second student network model, where represents the weight matrix in the second student network model, , represents the randomly initialized matrix after training update.

3. The natural language processing method according to claim 2, wherein The method further includes: Sorting the diagonal elements in the first diagonal matrix in ascending order; Select multiple element values from the sorting result according to a preset ratio, set the other element values that are not selected to zero, and generate a second diagonal matrix ; Based on the product of the second diagonal matrix and , determine that the weight matrix in the third student network model is ; Fixing the zero vector columns in the weight matrix of the third student network model and setting the element value states in the remaining columns to be learnable states; Training the third student network model based on the task-related loss function and the text data corresponding to the target task node to update the element values in the remaining columns; Upon completion of training, generating the target student network model.

4. The natural language processing method according to claim 3, wherein After generating the target student network model, the method further includes: Deploying the target student network model to the target task node.

5. The natural language processing method according to claim 2, wherein Training the first student network model based on the training loss function and the training data to update the weight matrix in the first student network model includes: Fixing the weight matrix parameters of the teacher network model, randomly initializing the element values and other parameters in the first student network model; Set the relevant parameters of the stochastic gradient descent algorithm, and define the current training epoch as , and the total number of training epochs is , in response to when , perform the following operations: Increment the current training cycle by one; Randomly shuffling the sample order in the training data and selecting a target batch of training samples from the training data; Performing a training operation according to the training samples and the relevant parameters of the stochastic gradient descent algorithm, calculating a training loss value, and updating the weight matrix in the first student network model using the stochastic gradient descent algorithm; Upon utilization of all the training samples in the training data, ending the current training cycle; Upon the current training cycle being a preset training cycle, ending the training to obtain the second student network model.

6. The natural language processing method according to claim 5, wherein Randomly initializing the element values in the first student network model includes: Performing random initialization processing on the element values in the first student network model based on a Gaussian distribution.

7. The natural language processing method according to claim 2, wherein Constructing a training loss function includes: Constructing the training loss function based on the distillation loss function and the task-related loss function.

8. The natural language processing method according to claim 7, wherein The training loss function includes: ; Among them, represents the training loss value, represents the task-related loss value, represents the constant parameter, represents the distillation loss value.

9. The natural language processing method according to claim 7, wherein The distillation loss function includes: ; ; Among them, represents the distillation loss value, represents the number of weight matrices in the teacher network model, represents the matrix and the matrix the distillation loss value between them, represents a random matrix, represents a diagonal matrix.

10. The natural language processing method according to claim 9, characterized in that, The matrix and the matrix The calculation method of the distillation loss value between them includes: ; Among them, denotes norm, denotes the matrix F norm, denotes a constant parameter, denotes the vector composed of the diagonal elements of the diagonal matrix.

11. The natural language processing method according to claim 7, characterized in that, The task-related loss function includes: ; Among them, represents the training loss value, represents the number of tokens, represents the th element of the vector, represents the th element of the vector, represents a one-hot vector of dimension corresponding to the true token at position represents the predicted dimensional probability distribution vector.

12. The natural language processing method according to claim 1, characterized in that In response to receiving a natural language processing request, obtain text data, and divide the text data into a plurality of tokens, including: Receive a natural language processing request sent by a target task node; In response to receiving the natural language processing request, parse the natural language processing request to obtain text data corresponding to the natural language processing request; According to a division rule, divide the text data to obtain a plurality of tokens corresponding to the text data.

13. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the natural language processing method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the natural language processing method according to any one of claims 1 to 12 when executed by a processor.

15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the natural language processing method according to any one of claims 1 to 12 when executed by a processor.

Citation Information

Patent Citations

  • Student network model generation method and device, equipment and storage medium

    CN111598216A

  • Industrial anomaly detection model training method and device based on multi-model fusion

    CN116028891A

  • Knowledge distillation-based text processing method and device, equipment and medium

    CN116050516A

  • Multi-modal learning level mining method and system under small sample condition and medium

    CN116186250A

  • Knowledge distillation method, device, equipment, storage medium and program product

    CN118627590A