A method to improve the accuracy of TinyLlama model
By improving the network structure of the TinyLlama model and enhancing the capabilities of the Transformer encoder and output layer, the problem of low accuracy of the lightweight model under low computing power conditions is solved, and a better understanding of long text and context semantics and an improvement in reply accuracy is achieved.
Patent Information
- Application Number
- CN202411699214.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-26
AI Technical Summary
The lightweight model has low accuracy when processing long text and context semantics, especially under low computing power, which makes it difficult to effectively improve the accuracy of reply.
By improving the network structure of the TinyLlama model, including the construction of an improved Transformer encoder and output layer, using improved Attention module and rotation position encoding, the model's long text and context semantics comprehension under low computing power conditions.
It significantly improves the model's reply accuracy under low computing power conditions and enhances its understanding of long text and context semantics.
Smart Images

Figure CN119204109B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a method for improving the accuracy of a TinyLlama model. Background Art
[0002] Lightweight models are often accompanied by a significant decrease in accuracy. Improving the accuracy of lightweight model responses has many applications, such as question-answering systems, task classification, sentiment analysis, etc. It is a very challenging problem in the field of artificial intelligence.
[0003] Some representative works to solve the problem of model response accuracy are models that use encoders and decoders in the model, which usually consist of two parts: the encoder is responsible for encoding the input sequence into a context vector, and the decoder generates the output sequence from the context vector. Usually more complex than the encoder-only model and has higher computational requirements, especially when processing long sequences. The input and output can be variable-length sequences, and the decoder outputs the generated results of each step based on the encoder's context. Usually more complex than the encoder-only model and has higher computational requirements, especially when processing long sequences. Since a lightweight model is used, the high computational requirements of the encoder-decoder are very disadvantageous.
[0004] In addition, in terms of feature fusion, most of them connect the feature maps of the same scale of the encoder and decoder through the jump connection mechanism, helping the model to fuse shallow information with deep information to achieve effective fusion of coarse-grained features and fine-grained features. However, this process is limited in improving the representation ability of high-resolution feature maps, which is crucial for dense prediction tasks such as human posture estimation. Summary of the invention
[0005] In view of this, the present invention provides a method for improving the accuracy of the TinyLlama model, so as to enhance the model's ability to understand long texts and contextual semantics under low computing power conditions, and improve the accuracy of the model's responses.
[0006] In a first aspect, the present invention provides a method for improving the accuracy of a TinyLlama model, the method comprising:
[0007] Step 1: Get the conversation dataset , represents the i-th data in the conversation dataset S, , perform data preprocessing on the data set S to obtain the preprocessed dialogue data set ;
[0008] Step 2: construct an improved TinyLlama network structure, wherein the improved TinyLlama network structure includes: an input layer, an improved Transformer encoder, an output layer, and the i-th number in the preprocessed conversation dataset I Input into the improved TinyLlama network structure and get the output text.
[0009] Optionally, step 1 includes:
[0010] The conversation dataset All samples included in the preprocessed dialogue dataset are trimmed, and texts with a length greater than 4096 characters are trimmed. Then, text normalization is performed, including: converting to lowercase, removing special characters, processing extra spaces, and removing spelling errors. .
[0011] Optionally, the step 2 includes:
[0012] Step 21: Input the i-th data in the conversation dataset S into the input layer to obtain a real vector logger; the input layer includes four stages, namely: serialization, indexing, embedding and position encoding;
[0013] Step 22, pass the real number vector logger through the encoder to obtain the hidden state sequence X; wherein the encoder includes the first stage to the tenth stage, each stage includes an improved Attention module and a Forword module, the modules and parameters of each stage are completely consistent, and the input of the latter stage is the output of the previous stage; the improved Attention module includes five consecutive parts, namely: root mean square normalization function, fully connected layer, improved attention mechanism, SoftMax activation function and fully connected layer, and the Forword module includes two consecutive parts, namely: root mean square normalization function and fully connected layer;
[0014] Step 23: Input the hidden sequence X into the output layer to generate the probability distribution of the next word; the output layer includes a root mean square normalization function, a fully connected layer, and a SoftMax activation function.
[0015] Optionally, the step 21 includes:
[0016] Serialization: After a piece of text is input, it needs to be tokenized and divided into words or characters to form a token sequence. When segmenting, the llama tokenizer is used;
[0017] Indexing: After obtaining the token sequence, the text is mapped into an input form that the model understands, and the text sequence is converted into an integer index sequence, where the index is the index of the word or character in the corpus;
[0018] Embedding: After the text is indexed, embedding continues to map each token into a real number vector, which is the embedding vector.
[0019] Position encoding: For each position in the token sequence, a position encoding vector is added to provide information about the position of the token in the sequence; position encoding is used to distinguish tokens at different positions and provide the model with contextual information.
[0020] Optionally, the step 22 includes:
[0021] The steps of RMS normalization are as follows:
[0022] 1) Calculate the root mean square, the formula is:
[0023] ;
[0024] 2) Normalize each data point using the root mean square value, the formula is:
[0025] ;
[0026] in is each value in the data set, n is the number of data points; the real vector with position information is normalized by the root mean square to obtain the real vector ;
[0027] The improved attention mechanism adopts a multi-head attention mechanism, and the input is First, the input is linearly transformed to obtain the query Q, key K and value V, and the obtained query Q and key K are rotationally encoded, where the formulas of query Q, key K and value V are:
[0028] ;
[0029] in , , is the weight matrix; set the multi-head attention to 32 heads, and the dimension of each head is 128; for each head, calculate the attention weight, the formula is:
[0030] ;
[0031] is the dimension of the key, which is used for scaling to prevent the gradient from disappearing; finally, the output of each head is concatenated together and transformed through a linear layer, and the formula is: ;
[0032] in is the output weight matrix;
[0033] The output of the Attention module is used as the input of the Forword module and normalized by RMS to obtain a real vector , After passing through the fully connected layer, we get a real number vector after feature extraction ;Will As the input of the second stage, the output ; then As the input of the third stage, this is repeated until the tenth stage. ~ As the input of the second to the tenth stage, ~ As the output of the first stage to the tenth stage respectively, the final output is obtained .
[0034] Optionally, the step 23 includes:
[0035] The final output of the encoder is input into the RMS normalization function to obtain the matrix ;
[0036] The resulting matrix , input to the fully connected layer, and get the matrix ;
[0037] The matrix Input into the SoftMax activation function to get the prediction result, that is, output a text.
[0038] In a second aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute a method for improving the accuracy of the TinyLlama model in the first aspect or any possible implementation of the first aspect.
[0039] In a fourth aspect, an embodiment of the present invention provides an electronic device, comprising: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, enable the device to execute a method for improving the accuracy of the TinyLlama model in the first aspect or any possible implementation of the first aspect.
[0040] In the technical solution provided by the present invention, the method includes obtaining a conversation data set S, performing data preprocessing on the data set S to obtain a preprocessed conversation data set I; constructing an improved TinyLlama network structure, the improved TinyLlama network structure including: an input layer, an improved Transformer encoder, an output layer, and converting the i-th item in the preprocessed conversation data set I into a The input is into the improved TinyLlama network structure to obtain the output text. This method enhances the model's ability to understand long texts and contextual semantics under low computing power conditions, and improves the accuracy of the model's responses. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1 A flow chart of a method for improving the accuracy of the TinyLlama model provided by an embodiment of the present invention;
[0043] Figure 2 A schematic diagram of an improved TinyLlama model provided in an embodiment of the present invention;
[0044] Figure 3 A schematic diagram of an improved Attention module provided in an embodiment of the present invention;
[0045] Figure 4 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] It should be clear that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings.
[0049] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0050] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0051] Figure 1 A flowchart of a method for improving the accuracy of the TinyLlama model provided by an embodiment of the present invention, such as Figure 1 As shown, the method includes:
[0052] Step 1: Get the conversation dataset , represents the i-th data in the conversation dataset S, , perform data preprocessing on the data set S to obtain the preprocessed dialogue data set .
[0053] In the embodiment of the present invention, step 1 includes:
[0054] The conversation dataset All samples included in the preprocessed dialogue dataset are trimmed, and texts with a length greater than 4096 characters are trimmed. Then, text normalization is performed, including: converting to lowercase, removing special characters, processing extra spaces, and removing spelling errors. .
[0055] The preprocessed dialogue data set I is divided for training. The specific operation is as follows: the first 50% of the data set is divided for model training; the remaining 50% of the question data is used as a test set, and the answer data is used as a validation set.
[0056] Step 2: construct an improved TinyLlama network structure, wherein the improved TinyLlama network structure includes: an input layer, an improved Transformer encoder, an output layer, and the i-th number I in the preprocessed conversation dataset I i Input into the improved TinyLlama network structure and get the output text.
[0057] In the embodiment of the present invention, Figure 2 and Figure 3 As shown, step 2 includes:
[0058] Step 21: Input the i-th data in the conversation dataset S into the input layer to obtain a real vector logger; the input layer includes four stages, namely: serialization, indexing, embedding and position encoding;
[0059] Step 22, pass the real number vector logger through the encoder to obtain the hidden state sequence X; wherein the encoder includes the first stage to the tenth stage, each stage includes an improved Attention module and a Forword module, the modules and parameters of each stage are completely consistent, and the input of the latter stage is the output of the previous stage; the improved Attention module includes five consecutive parts, namely: root mean square normalization function, fully connected layer, improved attention mechanism, SoftMax activation function and fully connected layer, and the Forword module includes two consecutive parts, namely: root mean square normalization function and fully connected layer;
[0060] Step 23: Input the hidden sequence X into the output layer to generate the probability distribution of the next word; the output layer includes a root mean square normalization function, a fully connected layer, and a SoftMax activation function.
[0061] In the embodiment of the present invention, step 21 includes:
[0062] Serialization: After a piece of text is input, it needs to be tokenized and divided into words or characters to form a token sequence. When segmenting, the llama tokenizer is used;
[0063] In some embodiments, for example:
[0064] enter:
[0065] {The moonlight shines brightly in front of my bed, I wonder if it is frost on the ground. I look up at the moon and think of my hometown};
[0066] Serialization:
[0067] ['Bos', 'bed', 'in front', 'bright moon', 'light', .... '', 'lowering my head', 'thinking', 'hometown'].
[0068] Indexing: After obtaining the token sequence, the text is mapped into an input form that the model understands, and the text sequence is converted into an integer index sequence, where the index is the index of the word or character in the corpus;
[0069] In some embodiments, for example:
[0070] Serialization:
[0071] ['Bos', 'bed', 'in front', 'bright moon', 'light', ... '', 'lowering', 'thinking', 'hometown'];
[0072] Indexing:
[0073] ['Bos', '10', '3', '5755', '809', .....'', '1354', '564', '155'].
[0074] Embedding: After the text is indexed, embedding continues to map each token into a real number vector, which is the embedding vector.
[0075] In some embodiments, the embedding vector is, for example:
[0076] .
[0077] Position encoding: For each position in the token sequence, a position encoding vector is added to provide information about the position of the token in the sequence; position encoding is used to distinguish tokens at different positions and provide the model with contextual information.
[0078] In some embodiments, the position code is, for example:
[0079] .
[0080] In the embodiment of the present invention, step 22 includes:
[0081] Root Mean Square Normalization (RMS Normalization) is a method of normalizing data to adjust the amplitude of the data so that they have the same range.
[0082] The steps of RMS normalization are as follows:
[0083] 1) Calculate the root mean square, the formula is:
[0084] ;
[0085] 2) Normalize each data point using the root mean square value, the formula is:
[0086] ;
[0087] in is each value in the data set, n is the number of data points; the real vector with position information is normalized by the root mean square to obtain the real vector ;
[0088] The improved attention mechanism adopts a multi-head attention mechanism, and the input is First, the input is linearly transformed to obtain the query Q, key K and value V, and the obtained query Q and key K are rotationally encoded, where the formulas of query Q, key K and value V are:
[0089] ;
[0090] in , , is the weight matrix; set the multi-head attention to 32 heads, and the dimension of each head is 128; for each head, calculate the attention weight, the formula is:
[0091] ;
[0092] is the dimension of the key, which is used for scaling to prevent the gradient from disappearing; finally, the output of each head is concatenated together and transformed through a linear layer, and the formula is: ;
[0093] in is the output weight matrix;
[0094] The output of the Attention module is used as the input of the Forword module and normalized by RMS to obtain a real vector , After passing through the fully connected layer, we get a real number vector after feature extraction ;Will As the input of the second stage, the output ; then As the input of the third stage, this is repeated until the tenth stage. ~ As the input of the second to the tenth stage, ~ As the output of the first stage to the tenth stage respectively, the final output is obtained .
[0095] In the embodiment of the present invention, step 23 includes:
[0096] The final output of the encoder is input into the RMS normalization function to obtain the matrix ;
[0097] The resulting matrix , input to the fully connected layer, and get the matrix ;
[0098] The matrix Input into the SoftMax activation function to get the prediction result, that is, output a text.
[0099] In the embodiment of the present invention, when the model is asked a question, the input is "Who are you?" This input is converted into a real vector logger through the output layer, and logger is then used as an input to enter the encoder. After passing through the encoder, a capture matrix is obtained. , After passing through the output layer, the output is: I am Open Assistant, a language model based on open source protocols and a large number of artificial intelligence models. I can understand multiple languages, generate text from them, and communicate with people.
[0100] This paper adopts an innovative method that combines global and local attention mechanisms. By introducing improved rotational position encoding and optimized attention calculation in the model, it aims to enhance the model's understanding and generation capabilities in complex tasks. In addition, the study will explore the integration of various plug-ins to further enhance the functionality and flexibility of the model. Finally, the effectiveness of the proposed method in improving reasoning efficiency and accuracy is verified through experiments, providing new ideas and practical experience in the field of natural language processing.
[0101] Each step of the embodiment of the present invention may be performed by an electronic device, which includes but is not limited to a mobile phone, a tablet computer, a portable PC, a desktop computer, etc.
[0102] In the technical solution provided by the present invention, the method includes obtaining a conversation data set S, performing data preprocessing on the data set S to obtain a preprocessed conversation data set I; constructing an improved TinyLlama network structure, the improved TinyLlama network structure including: an input layer, an improved Transformer encoder, an output layer, and converting the i-th item in the preprocessed conversation data set I into a The input is into the improved TinyLlama network structure to obtain the output text. This method enhances the model's ability to understand long texts and contextual semantics under low computing power conditions, and improves the accuracy of the model's responses.
[0103] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program. When the program is running, the electronic device where the computer-readable storage medium is located is controlled to execute an embodiment of the method for improving the accuracy of the TinyLlama model.
[0104] Figure 4 A schematic diagram of an electronic device provided by an embodiment of the present invention, such as Figure 4 As shown, the electronic device 21 includes: a processor 211, a memory 212, and a computer program 213 stored in the memory 212 and executable on the processor 211. When the computer program 213 is executed by the processor 211, a method for improving the accuracy of the TinyLlama model in an embodiment is implemented. To avoid repetition, they are not described one by one here.
[0105] The electronic device 21 includes, but is not limited to, a processor 211 and a memory 212. Those skilled in the art will appreciate that Figure 4 It is only an example of the electronic device 21 and does not constitute a limitation of the electronic device 21. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0106] The processor 211 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0107] The memory 212 may be an internal storage unit of the electronic device 21, such as a hard disk or memory of the electronic device 21. The memory 212 may also be an external storage device of the electronic device 21, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card (FlashCard), etc. equipped on the electronic device 21. Further, the memory 212 may also include both an internal storage unit of the electronic device 21 and an external storage device. The memory 212 is used to store computer programs and other programs and data required by network devices. The memory 212 may also be used to temporarily store data that has been output or is to be output.
[0108] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0109] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for improving the accuracy of the TinyLlama model, characterized in that: The method comprises: Step 1: Get the conversation dataset , represents the i-th data in the conversation dataset S, , perform data preprocessing on the data set S to obtain the preprocessed dialogue data set ; Step 2: construct an improved TinyLlama network structure, wherein the improved TinyLlama network structure includes: an input layer, an improved Transformer encoder, an output layer, and the i-th number in the preprocessed conversation dataset I Input into the improved TinyLlama network structure to get the output text; The step 2 comprises: Step 21: Input the i-th data in the conversation dataset S into the input layer to obtain a real vector logger; the input layer includes four stages, namely: serialization, indexing, embedding and position encoding; Step 22, pass the real number vector logger through the encoder to obtain the hidden state sequence X; wherein the encoder includes the first stage to the tenth stage, each stage includes an improved Attention module and a Forword module, the modules and parameters of each stage are completely consistent, and the input of the latter stage is the output of the previous stage; the improved Attention module includes five consecutive parts, namely: root mean square normalization function, fully connected layer, improved attention mechanism, SoftMax activation function and fully connected layer, and the Forword module includes two consecutive parts, namely: root mean square normalization function and fully connected layer; Step 23: Input the hidden sequence X into the output layer to generate the probability distribution of the next word; the output layer includes a root mean square normalization function, a fully connected layer, and a SoftMax activation function.
2. The method according to claim 1, characterized in that: The step 1 comprises: The conversation dataset All samples included in the preprocessed dialogue dataset are trimmed, and texts with a length greater than 4096 characters are trimmed. Then, text normalization is performed, including: converting to lowercase, removing special characters, processing extra spaces, and removing spelling errors. .
3. The method according to claim 1, characterized in that The step 21 comprises: Serialization: After a piece of text is input, it needs to be tokenized and divided into words or characters to form a token sequence. When segmenting, the llama tokenizer is used; Indexing: After obtaining the token sequence, the text is mapped into an input form that the model understands, and the text sequence is converted into an integer index sequence, where the index is the index of the word or character in the corpus; Embedding: After the text is indexed, embedding continues to map each token into a real number vector, which is the embedding vector. Position encoding: For each position in the token sequence, a position encoding vector is added to provide information about the position of the token in the sequence; position encoding is used to distinguish tokens at different positions and provide the model with contextual information.
4. The method according to claim 1, characterized in that: The step 22 comprises: The steps of RMS normalization are as follows: 1) Calculate the root mean square, the formula is: ; 2) Normalize each data point using the root mean square value, the formula is: ; in is each value in the data set, n is the number of data points; the real vector with position information is normalized by the root mean square to obtain the real vector ; The improved attention mechanism adopts a multi-head attention mechanism, and the input is First, the input is linearly transformed to obtain the query Q, key K and value V, and the obtained query Q and key K are rotationally encoded, where the formulas of query Q, key K and value V are: ; in , , is the weight matrix; set the multi-head attention to 32 heads, and the dimension of each head is 128; for each head, calculate the attention weight, the formula is: ; is the dimension of the key, which is used for scaling to prevent the gradient from disappearing; finally, the output of each head is concatenated together and transformed through a linear layer, and the formula is: ; in is the output weight matrix; The output of the Attention module is used as the input of the Forword module and normalized by RMS to obtain a real vector , After passing through the fully connected layer, we get a real number vector after feature extraction ;Will As the input of the second stage, the output ; then As the input of the third stage, this is repeated until the tenth stage. ~ As the input of the second to the tenth stage, ~ As the output of the first stage to the tenth stage respectively, the final output is obtained .
5. The method according to claim 1, characterized in that The step 23 comprises: The final output of the encoder is input into the RMS normalization function to obtain the matrix ; The resulting matrix , input to the fully connected layer, and get the matrix ; The matrix Input into the SoftMax activation function to get the prediction result, that is, output a text.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is run, the computer-readable storage medium is controlled to execute a method for improving the accuracy of the TinyLlama model according to any one of claims 1 to 5.
7. An electronic device, characterized in that: include: one or more processors; Memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, enable the device to perform a method for improving the accuracy of the TinyLlama model as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Chinese intelligent dialogue method based on Transformer
CN112612881A
Conversation generation method fusing basic knowledge and user information
CN116010575A