Natural language processing method based on improved Lama model

By improving the Llama model, building the input layer, language encoder and output layer, and pre-processing the conversation data, the problems of high training cost, high inference complexity and poor model stability in the existing natural language processing technology are solved, and more efficient and stable natural language processing performance is achieved.

CN119961414AActive Publication Date: 2025-05-09QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510117657.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-09
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

While existing natural language processing technologies maintain strong language generation capabilities, they are difficult to reduce training costs and inference complexity, and the model is poorly stable.

Method used

By improving the Llama model, including building input layer, language encoder and output layer, and preprocessing dialogue data, it reduces training costs and inference complexity, and improves model stability.

Benefits of technology

While maintaining strong language generation capabilities, it reduces training costs and inference complexity, improves model stability, and improves the overall performance of natural language processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961414A_ABST
    Figure CN119961414A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, and particularly provides a natural language processing method based on an improved Lama model. The method comprises the following steps: acquiring dialogue data S, and preprocessing the data S to obtain preprocessed dialogue data M; according to the method, an improved Lma model is constructed, the improved Lma model comprises an input layer, a language encoder and an output layer, the preprocessed dialogue data M is input into the improved Lma model, and an output text is obtained.According to the method, through the efficient and light-weight pre-training language model Lma, the training cost and the reasoning complexity are reduced while the powerful language generation capacity is kept, and the training efficiency is improved. And the model stability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a natural language processing method based on an improved Llama model. Background Art

[0002] In recent years, the field of natural language processing (NLP) has made significant progress with the development of deep learning technology. Among them, large-scale pre-trained language models have become an important technical means to promote the performance improvement of NLP tasks.

[0003] Early NLP systems relied on hand-designed rules and finite state machines, such as syntactic analyzers and word segmenters. This approach is effective for specific fields, but has poor versatility and high development costs. With the introduction of statistical language models (such as N-gram) and machine learning algorithms, NLP systems have begun to be able to model language in a data-driven way, but such methods have limited performance in high-dimensional contexts. Summary of the invention

[0004] In view of this, the present invention provides a natural language processing method based on an improved Llama model, so as to reduce training costs and reasoning complexity and improve model stability while maintaining powerful language generation capabilities.

[0005] In a first aspect, the present invention provides a natural language processing method based on an improved Llama model, the method comprising: Step 1: Acquire conversation data S, preprocess the data S, and obtain preprocessed conversation data M; Step 2: construct an improved Llama model, wherein the improved Llama model includes: an input layer, a language encoder, and an output layer. The preprocessed dialogue data M is input into the improved Llama model to obtain output text.

[0006] Optionally, step 1 includes: The conversation data S is preprocessed, including converting to lowercase, removing special characters, processing extra spaces, and removing spelling errors, to obtain preprocessed conversation data M.

[0007] Optionally, the step 2 includes: Step 21: Input the preprocessed conversation data M into the input layer to obtain a real number vector ; The input layer consists of two stages, namely: word segmentation layer and embedding layer; Step 22: Convert the real vector The hidden state sequence is obtained through the language encoder ; The encoder includes the first to twentieth stages, each of which includes four consecutive parts, which are: global and local fusion attention mechanism, normalization, forward propagation, and normalization; the modules and parameters of each stage are exactly the same, and the input of the next stage is the output of the previous stage; Step 23: Hidden state sequence The input is sent to the output layer to generate the probability distribution of the next word; the output layer includes a projection layer and a de-tagged layer.

[0008] Optionally, the step 21 includes: Word segmentation layer: After a piece of text is input, the text is tokenized and divided into words or characters to form a token sequence. After obtaining the token sequence, the text is mapped into an input form that the model understands, and the text sequence is converted into an integer index sequence, where the index is the index of the word or character in the corpus; Embedding: After the text passes through the word segmentation layer, embedding continues to map each token into a real number vector, which is the embedding vector .

[0009] Optionally, the step 22 includes: Using the global and local fusion attention mechanism, the input is First, Enter the global attention mechanism and the local attention mechanism respectively; The calculation formula of the global attention mechanism is: ; Among them, Q, K, and V are query, key, and value matrices respectively; is the dimension of the key, used as a scaling factor; The calculation formula of the local attention mechanism is: ; in, Represents a local window range, where only the keys K and values ​​V at adjacent positions in the sequence are selected for calculation; Computational complexity: From the global Reduced to , where w is the window size and m is the sequence length; The fusion formula is: ; in is the output weight matrix; Take n_1 as the input of the first layer of the encoder, and n_1 is obtained through the global and local attention mechanism. , then Normalize to get V1, input V1 into the forward propagation layer, get output V2, V2 is normalized to get V3, perform weighted summation of V3 and V1, and reassign the obtained matrix to ,Will As the final output of the first stage; The normalization steps are as follows: (1) Calculate the root mean square, the formula is: ; (2) Use the root mean square value to normalize each data point, and the formula is: ; in is each value in the data, and n is the number of data points; As the output of the first stage, As the input of the second stage, stage two goes through the same process as stage one and then outputs ; then As the input of the third stage, this is repeated until the twentieth stage. ~ As the input of the second stage to the twentieth stage respectively, ~ As the output of the first stage to the tenth stage respectively, the final output is obtained .

[0010] Optionally, the step 23 includes: The final output of the encoder Input the projection layer, which can convert the matrix into a text ID sequence, and then input the obtained ID sequence into the de-marker, which converts the ID sequence into text, and finally obtains a text output.

[0011] In a second aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the natural language processing method based on the improved Llama model in the first aspect or any possible implementation of the first aspect.

[0012] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, enable the device to execute the natural language processing method based on the improved Llama model in the first aspect or any possible implementation of the first aspect.

[0013] In the technical solution provided by the present invention, the method includes obtaining dialogue data S, preprocessing the data S to obtain preprocessed dialogue data M; constructing an improved Llama model, the improved Llama model includes: an input layer, a language encoder, and an output layer, and inputting the preprocessed dialogue data set M into the improved Llama model to obtain output text. The method uses an efficient and lightweight pre-trained language model Llama to reduce training costs and reasoning complexity while maintaining powerful language generation capabilities, thereby improving model stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0015] Figure 1 A flowchart of a natural language processing method based on an improved Llama model provided in an embodiment of the present invention; Figure 2 A schematic diagram of an improved Llama model provided in an embodiment of the present invention; Figure 3 A schematic diagram of a language encoder layer provided by an embodiment of the present invention; Figure 4 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0017] It should be clear that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0018] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings.

[0019] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0020] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0021] Figure 1 A flowchart of a natural language processing method based on an improved Llama model provided in an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method includes: Step 1: Acquire conversation data S, preprocess the data S, and obtain preprocessed conversation data M; In the embodiment of the present invention, step 1 includes: The conversation data S is preprocessed, including converting to lowercase, removing special characters, processing extra spaces, and removing spelling errors, to obtain preprocessed conversation data M.

[0022] One million pre-processed conversation data were collected and divided for training. The specific operation was as follows: the first 50% of the data set was divided for model pre-training; the remaining 50% of the question data was used as a test set, and the answer data was used as a validation set.

[0023] Step 2: construct an improved Llama model, wherein the improved Llama model includes: an input layer, a language encoder, and an output layer. The preprocessed dialogue data M is input into the improved Llama model to obtain output text.

[0024] In the embodiment of the present invention, Figure 2 and Figure 3 As shown, step 2 includes: Step 21: Input the preprocessed conversation data M into the input layer to obtain a real number vector ; The input layer consists of two stages, namely: word segmentation layer and embedding layer; In the embodiment of the present invention, step 21 includes: Word segmentation layer: After a piece of text is input, the text is tokenized and divided into words or characters to form a token sequence. After obtaining the token sequence, the text is mapped into an input form that the model understands, and the text sequence is converted into an integer index sequence, where the index is the index of the word or character in the corpus; When segmenting words, the llama word segmenter is used. In some embodiments, for example: enter: {The sun sets over the mountains, and the Yellow River flows into the sea. If you want to see a thousand miles, you must climb to a higher level}; Token sequence: ['Bos', 'white', 'sun', 'depending on', 'mountain', ..... '', 'up', 'one floor', 'floor'].

[0025] Index sequence: ['Bos', '10', '3', '5755', '809', .....'', '1354', '564', '155'].

[0026] Embedding: After the text passes through the word segmentation layer and is converted into an index sequence, embedding continues to map each token into a real number vector, which is the embedding vector n_1; In some embodiments, the embedding vector n_1 is, for example: .

[0027] Step 22: Convert the real vector The hidden state sequence is obtained through the language encoder ; The encoder includes the first to twentieth stages, each of which includes four consecutive parts, which are: global and local fusion attention mechanism, normalization, forward propagation, and normalization; the modules and parameters of each stage are exactly the same, and the input of the next stage is the output of the previous stage; In the embodiment of the present invention, step 22 includes: Using the global and local fusion attention mechanism, the input is First, Enter the global attention mechanism and the local attention mechanism respectively; The calculation formula of the global attention mechanism is: ; Among them, Q, K, and V are query, key, and value matrices respectively; is the dimension of the key, used as a scaling factor; The calculation formula of the local attention mechanism is: ; in, Represents a local window range, where only the keys K and values ​​V at adjacent positions in the sequence are selected for calculation; Computational complexity: From the global Reduced to , where w is the window size and m is the sequence length; The fusion formula is: ; in is the output weight matrix; Take n_1 as the input of the first layer of the encoder, and n_1 is obtained through the global and local attention mechanism. , then Normalize to get V1, input V1 into the forward propagation layer, get output V2, V2 is normalized to get V3, perform weighted summation of V3 and V1, and reassign the obtained matrix to ,Will As the final output of the first stage.

[0028] The normalization steps are as follows: (1) Calculate the root mean square, the formula is: ; (2) Use the root mean square value to normalize each data point, and the formula is: ; in is each value in the data, and n is the number of data points; is the output of the first stage, As the input of the second stage, stage two goes through the same process as stage one and then outputs ; then As the input of the third stage, this is repeated until the twentieth stage. ~ As the input of the second stage to the twentieth stage respectively, ~ As the output of the first stage to the tenth stage respectively, the final output is obtained .

[0029] In the embodiment of the present invention, step 23 includes: The final output of the encoder Input the projection layer, which can convert the matrix into a text ID sequence, and then input the obtained ID sequence into the de-marker, which converts the ID sequence into text, and finally obtains a text output.

[0030] In the embodiment of the present invention, when the model is asked a question, the input is "What is your favorite movie?" This input is converted into a real number vector through the output layer. , Then it enters the encoder as an input, and after passing through the encoder, a capture matrix is ​​obtained ,matrix After passing through the output layer, the output is obtained: "My favorite movie genres are science fiction and horror movies, such as "Interstellar" and "Avengers 4: Endgame". In addition, I also like to watch suspense and comedy movies, such as "Mad Max 4" and "Detective Chinatown 3".'

[0031] This invention aims to take into account both global and local information, avoid the limitations of a single attention mode, and significantly reduce complexity in long sequence tasks while maintaining high performance. In addition, the research will explore the integration of various plug-ins to further enhance the functionality and flexibility of the model. Finally, the effectiveness of the proposed method in improving reasoning efficiency and accuracy will be verified through experiments, providing new ideas and practical experience in the field of natural language processing.

[0032] Each step of the embodiment of the present invention may be performed by an electronic device, which includes but is not limited to a mobile phone, a tablet computer, a portable PC, a desktop computer, etc.

[0033] In the technical solution provided by the present invention, the method includes obtaining dialogue data S, preprocessing the data S to obtain preprocessed dialogue data M; constructing an improved Llama model, the improved Llama model includes: an input layer, a language encoder, and an output layer, and inputting the preprocessed dialogue data M into the improved Llama model to obtain output text. The method uses an efficient and lightweight pre-trained language model Llama to reduce training costs and reasoning complexity while maintaining powerful language generation capabilities, thereby improving model stability.

[0034] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program. When the program is running, the electronic device where the computer-readable storage medium is located is controlled to execute the above-mentioned embodiment of the natural language processing method based on the improved Llama model.

[0035] Figure 4 A schematic diagram of an electronic device provided by an embodiment of the present invention, such as Figure 4 As shown, the electronic device 21 includes: a processor 211, a memory 212, and a computer program 213 stored in the memory 212 and executable on the processor 211. When the computer program 213 is executed by the processor 211, the natural language processing method based on the improved Llama model in the embodiment is implemented. To avoid repetition, they are not described one by one here.

[0036] The electronic device 21 includes, but is not limited to, a processor 211 and a memory 212. Those skilled in the art will appreciate that Figure 4 It is only an example of the electronic device 21 and does not constitute a limitation of the electronic device 21. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0037] The processor 211 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0038] The memory 212 may be an internal storage unit of the electronic device 21, such as a hard disk or memory of the electronic device 21. The memory 212 may also be an external storage device of the electronic device 21, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card (FlashCard), etc. equipped on the electronic device 21. Further, the memory 212 may also include both an internal storage unit of the electronic device 21 and an external storage device. The memory 212 is used to store computer programs and other programs and data required by network devices. The memory 212 may also be used to temporarily store data that has been output or is to be output.

[0039] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0040] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A natural language processing method based on an improved Llama model, characterized in that: The method comprises: Step 1: Acquire conversation data S, preprocess the data S, and obtain preprocessed conversation data M; Step 2: construct an improved Llama model, wherein the improved Llama model includes: an input layer, a language encoder, and an output layer. The preprocessed dialogue data M is input into the improved Llama model to obtain output text.

2. The method according to claim 1, characterized in that The step 1 comprises: The conversation data S is preprocessed, including converting to lowercase, removing special characters, processing extra spaces, and removing spelling errors, to obtain preprocessed conversation data M.

3. The method according to claim 1, characterized in that The step 2 comprises: Step 21: Input the preprocessed conversation data M into the input layer to obtain a real number vector ; The input layer consists of two stages, namely: word segmentation layer and embedding layer; Step 22: Convert the real vector The hidden state sequence is obtained through the language encoder ; The encoder includes the first to twentieth stages, each of which includes four consecutive parts, which are: global and local fusion attention mechanism, normalization, forward propagation, and normalization; the modules and parameters of each stage are exactly the same, and the input of the next stage is the output of the previous stage; Step 23: Hidden state sequence Input to the output layer to generate the probability distribution of the next word; the output layer includes the projection layer and the de-tagged layer.

4. The method according to claim 3, characterized in that The step 21 comprises: Word segmentation layer: After a piece of text is input, the text is tokenized and divided into words or characters to form a token sequence. After obtaining the token sequence, the text is mapped into an input form that the model understands, and the text sequence is converted into an integer index sequence, where the index is the index of the word or character in the corpus; Embedding: After the text passes through the word segmentation layer, embedding continues to map each Token into a real number vector as the embedding vector n_1.

5. The method according to claim 3, characterized in that: The step 22 comprises: Using the global and local fusion attention mechanism, the input is First, Enter the global attention mechanism and the local attention mechanism respectively; The calculation formula of the global attention mechanism is: ; Among them, Q, K, and V are query, key, and value matrices respectively; is the dimension of the key, used as a scaling factor; The calculation formula of the local attention mechanism is: ; in, Represents a local window range, where only the keys K and values ​​V at adjacent positions in the sequence are selected for calculation; Computational complexity: From the global Reduced to , where w is the window size and m is the sequence length; The fusion formula is: ; in is the output weight matrix; Take n_1 as the input of the first layer of the encoder, and n_1 is obtained through the global and local attention mechanism. , then Normalize to get V1, input V1 into the forward propagation layer, get output V2, V2 is normalized to get V3, perform weighted summation of V3 and V1, and reassign the obtained matrix to ,Will As the final output of the first stage; The normalization steps are as follows: (1) Calculate the root mean square, the formula is: ; (2) Use the root mean square value to normalize each data point, and the formula is: ; in is each value in the data, and n is the number of data points; is the output of the first stage, As the input of the second stage, stage two goes through the same process as stage one and then outputs ; then As the input of the third stage, this is repeated until the twentieth stage. ~ As the input of the second stage to the twentieth stage respectively, ~ As the output of the first stage to the tenth stage respectively, the final output is obtained .

6. The method according to claim 3, characterized in that: The step 23 comprises: The final output of the encoder Input the projection layer, which can convert the matrix into a text ID sequence, and then input the obtained ID sequence into the de-marker, which converts the ID sequence into text, and finally obtains a text output.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the natural language processing method based on the improved Llama model according to any one of claims 1 to 6.

8. An electronic device, characterized in that: include: one or more processors; Memory; And one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, which, when executed by the device, enable the device to perform the natural language processing method based on the improved Llama model as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Attention mechanism-based aspect-level text sentiment analysis method and system

    CN115329073A

  • Automatic auxiliary diagnosis method based on multi-modal LLM and model construction method thereof

    CN118098564A

  • Model obtaining method, device and equipment

    CN118114740A

  • Han machine translation system based on dynamic fusion attention model

    CN118520886A

  • Natural language processing

    EP3457332A1