Feature enhancement method based on residual autoregressive cross attention mechanism
Through the residual autoregressive cross-attention mechanism, combined with self-attention and cross-attention mechanisms, the problem of natural language processing models in long text feature extraction is solved, the feature extraction and accuracy of the model are improved, and it is suitable for natural language processing tasks.
Patent Information
- Application Number
- CN202510811766.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-30
AI Technical Summary
Existing natural language processing models have difficulty in capturing long-distance dependencies and fusing different text features when extracting features from long texts, which affects the performance and accuracy of the models.
The residual autoregressive cross-attention mechanism is adopted, which combines the self-attention mechanism and the cross-attention mechanism. The feature enhancement is performed by constructing a residual autoregressive cross-attention module, which includes an autoregressive module, a cross-attention module and a dynamic weight fusion module.
The model's feature extraction capability for natural language text has been significantly improved, enabling it to better capture long-distance dependencies and fuse different text features through a cross-attention mechanism, thereby improving the performance and accuracy of the model while retaining the original information and reducing the difficulty of use.
Smart Images

Figure CN120724151A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and specifically relates to a feature enhancement method based on a residual autoregressive cross-attention mechanism, which is suitable for natural language processing application scenarios. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, natural language processing (NLP) has achieved remarkable results in numerous fields, including outstanding performance in tasks such as text generation, machine translation, and sentiment analysis. The core of these tasks lies in the model's ability to understand and represent text data. However, existing NLP models still face challenges in extracting features from long texts. For example, models may struggle to capture long-range dependencies within text or be unable to effectively integrate features from different texts when processing large batches of text data, thus impacting model performance and accuracy.
[0003] In order to improve the model's ability to extract features from natural language text, researchers have proposed a variety of improvement methods. Among them, the attention mechanism is widely used due to its effectiveness in capturing the internal dependencies of text sequences. However, the traditional attention mechanism still has limitations when processing longer texts. For example, although the self-attention mechanism can capture long-distance dependencies in text, the core of the self-attention mechanism is to calculate the correlation between all word pairs. Therefore, when the input sequence length is n, the time complexity required is n. 2 In addition, a single attention mechanism may not be able to take into account both the local and global features of the text, thus affecting the model's comprehensive understanding of the text. Summary of the Invention
[0004] To solve these problems, the present invention proposes a feature enhancement method based on the residual autoregressive cross-attention mechanism. By combining residual connections, self-attention mechanism and cross-attention mechanism, the model's feature extraction ability for natural language text is effectively improved, and it can better capture long-distance dependencies in the text, thereby improving the performance and accuracy of the model in natural language processing tasks.
[0005] The present invention is implemented as follows: a feature enhancement method based on residual autoregressive cross attention mechanism, comprising the following steps:
[0006] Step S1, construct an unsupervised pre-training dataset (X t ) and supervised fine-tuning dataset (X y ), the unsupervised pre-training dataset is an unlabeled dataset, and the supervised fine-tuning dataset is a labeled dataset;
[0007] Step S2: constructing a large language model (E) based on the decoder, and constructing a residual autoregressive cross-attention module in the large language model (E), comprising an autoregressive module, a cross-attention module, and a dynamic weight fusion module, wherein the dynamic weight fusion module contains a residual connection, and the original input features are retained through the residual connection;
[0008] Step S3, using the unsupervised pre-training dataset (X t ), performing feature enhancement pre-training on the large language model (E) through the residual autoregressive cross-attention module constructed in step S2;
[0009] Step S4, using the supervised fine-tuning dataset (X y ), through supervised fine-tuning, the large language model (E) pre-trained in step S3 is post-trained until the model training is completed;
[0010] In step S5, the large language model (E) trained in step S4 is subjected to inference testing in the NPU environment, and the performance of the trained large language model (E) is evaluated. The evaluation indicators include long-context processing capability and the relevance of the generated text.
[0011] Furthermore, the unlabeled data set in step S1 includes text data.
[0012] Furthermore, the labeled data set in step S1 includes instructions, inputs, and outputs.
[0013] Furthermore, the autoregressive module in step S2 contains a self-attention mechanism for capturing the dependencies of the input long text sequence. The calculation formula of the self-attention mechanism is:
[0014]
[0015] Where Q is the query matrix, K is the key matrix, V is the value matrix, T is the matrix transpose operator, softmax is the Softmax function, that is, the normalized exponential function, and d k is the dimension of matrix K.
[0016] Furthermore, the cross-attention module in step S2 contains a cross-attention mechanism for fusing features of different text data. The calculation formula of the cross-attention mechanism is:
[0017]
[0018] Where Q1 is the query matrix of the first input sequence, K2 is the key matrix of the second input sequence, V2 is the value matrix of the second input sequence, and dk2 is the dimension of matrix K2.
[0019] Furthermore, the dynamic weight fusion calculation formula of the dynamic weight fusion module in step S2 is:
[0020] y f =α·t p +(1-α)·y d
[0021] Where y p is the output of the residual autoregressive cross attention module, y d is the output of the decoder module, y f is the weighted sum, α is the weight of the residual autoregressive cross attention module, α∈[0,1], and 1-α is the weight of the decoder.
[0022] Furthermore, the feature enhancement pre-training in step S3 uses the mean square error loss function to measure the difference between the output of the large language model (E) and the target value. The mean square error MSE calculation formula is:
[0023]
[0024] Where N is the number of samples, t i is the target value, is the output of the large language model.
[0025] Furthermore, the specific process of post-training the large language model (E) pre-trained in step S3 through supervised fine-tuning in step S4 until the model training is completed is as follows: in supervised fine-tuning, the large language model generates a predicted output for a specific task through forward propagation, and uses the cross-entropy loss function to measure the difference between the probability distribution of the predicted output and the probability distribution of the true label. The cross-entropy calculation formula is:
[0026]
[0027] In the formula, N is the number of samples, C is the number of categories, and y i,c is the probability distribution of the true label, Predicting the probability distribution of output for a large language model;
[0028] During post-training, the large language model adjusts parameters according to the cross-entropy loss function so that the probability distribution of the predicted output is closer to the probability distribution of the true label.
[0029] Furthermore, the evaluation method for the long context processing capability in step S5 includes NeedleBench and LV-Eval.
[0030] Furthermore, the evaluation method of the relevance of the generated text in step S5 adopts the N-gram model: for a sentence consisting of n words W = w1w2…w n , w1 is the first word, w n is the last word, and the N consecutive words in a sentence are called N-grams; suppose the i-th word w in the sentence i The probability of occurrence depends only on the previous i-1 words, that is, Then you can i Take a short window of N-1 yuan in front Approximate calculation of w i The probability of occurrence is calculated as follows:
[0031]
[0032] Where, w i The previous N-1 element window w i-(N-1) ,…,w i-1 ; P is probability; Count is the number of statistics; numerator Contains word w i N yuan w i-(N-1) ,…,w i The number of times a word appears in the corpus; the denominator is the number of times N-1 gram appears in the corpus.
[0033] The beneficial effects of the present invention are as follows: (1) by introducing the residual autoregressive cross-attention module, the feature extraction capability of the large language model for natural language text is significantly improved. The residual autoregressive cross-attention module can effectively capture the long-range dependencies of text sequences and fuse different text features through the cross-attention mechanism, thereby enhancing the model's ability to understand text input; (2) the introduction of residual connections ensures that the model does not lose original information while enhancing features, further improving the performance and stability of the model; (3) the present invention reduces the difficulty of using large language models in natural language processing tasks, making it easy for non-professionals to apply it, thereby improving the overall work efficiency of personnel.
[0034] The present invention will be further explained in detail below with reference to the accompanying drawings and specific implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a flow chart of the present invention;
[0036] Figure 2 Schematic diagram of the residual autoregressive cross attention module in step S2 of the present invention;
[0037] Figure 3 This is a flow chart of feature enhancement pre-training in step S3 of the present invention;
[0038] Figure 4 This is a schematic diagram of post-supervised fine-tuning training in step S4 of the present invention;
[0039] Figure 5 This is a curve showing the decrease of the cross entropy loss function during the supervised fine-tuning training in Example 1 of the present invention. DETAILED DESCRIPTION
[0040] Example 1:
[0041] This embodiment provides a feature enhancement method based on residual autoregressive cross attention mechanism, such as Figure 1 As shown, the following steps are included:
[0042] Step S1: construct unsupervised pre-training dataset X t and supervised fine-tuning dataset X y .
[0043] Unsupervised pre-training dataset X t It is an unlabeled dataset that contains a large amount of unlabeled text data for the model to learn common language patterns and features. This data can come from a variety of sources, such as news articles, books, or paper texts. y It is an annotated dataset that contains text data labeled with specific task objectives, usually including instructions, inputs, and outputs, such as text labels or question-answer pairs in text analysis, which are used for targeted fine-tuning based on pre-trained models to adapt to specific natural language processing tasks.
[0044] Step S2: construct a large language model E based on the decoder, and construct a residual autoregressive cross attention module in the large language model E, such as Figure 2 As shown in Figure 3, the module contains an autoregressive module, a cross-attention module, and a dynamic weight fusion module.
[0045] Figure 2 A large language model built based on a decoder is demonstrated, with a residual autoregressive cross-attention mechanism module added after the embedding layer to improve the model's ability to extract long text features.
[0046] (1) The autoregressive module contains a self-attention mechanism, which is used to capture the dependencies of the input long text sequence, thereby reducing the model complexity and improving the computational efficiency. The calculation formula of the self-attention mechanism is:
[0047]
[0048] Where Q is the query matrix (Query), K is the key matrix (Key), V is the value matrix (Value), T is the matrix transpose operator, softmax is the Softmax function (i.e., the normalized exponential function), d k is the dimension of matrix K. The self-attention mechanism captures the dependencies within the sequence by calculating the dot product attention within the input sequence. Specifically: through three learnable weight matrices W Q 、W K and W V , the input sequence X is linearly transformed into the query matrix Q, key matrix K and value matrix V, which is expressed as: Q = XW Q ,K=XW K ,V=XW V The attention score is calculated by taking the dot product of Q and K and dividing it by Scaling is performed to prevent the value from being too large. The attention score is converted into a probability distribution using the Softmax function to represent the attention weight of each position. Finally, the attention weight is multiplied by V to obtain the enhanced feature representation.
[0049] The autoregressive module enhances the model's contextual understanding of text by capturing dependencies within the sequence, such as capturing the correlation between the previous and next text in long sequence data, or capturing the dependencies between words in text data.
[0050] (2) The cross-attention module contains a cross-attention mechanism, which is used to integrate the features of different text data and enhance the model's ability to understand multi-source text input. The calculation formula of the cross-attention mechanism is:
[0051]
[0052] Where Q1 is the query matrix of the first input sequence, K2 is the key matrix of the second input sequence, and V2 is the value matrix of the second input sequence. is the dimension of matrix K2. The cross attention mechanism is used to fuse the features of two different sequences, such as fusing the features of the question text with the features of the context text. Specifically: the query matrix Q1 comes from one input sequence (such as the question text), while the key matrix K2 and the value matrix V2 come from another input sequence (such as the context text). The attention score is calculated by the dot product of Q1 and K2 and divided by Scale. Use the Softmax function to convert the attention score into a probability distribution, which represents the attention weight of the query sequence for each key. Finally, multiply the attention weight by V2 to obtain the fused feature representation.
[0053] The cross-attention module enhances the model's ability to understand multi-source text input by fusing features from different text data, effectively combining the features of question text with those of context text. For example, in a multi-document summarization task, text features from multiple documents can be fused to generate a more accurate summary.
[0054] (3) The dynamic weight fusion module further optimizes the feature representation by fusing the output of the decoder module with the output of the residual autoregressive cross-attention module in a weighted sum manner. The dynamic weight fusion calculation formula is:
[0055] y f =α·y p +(1-α)·y d
[0056] Where y p is the output of the residual autoregressive cross attention module, t d is the output of the decoder module, t f is a weighted sum, α is the weight of the residual autoregressive cross-attention module, α∈[0,1], and 1-α is the weight of the decoder. α is initially set to 0.5 and is continuously and dynamically updated during training to ensure that the model can flexibly adjust the contribution ratio of different modules during training, thereby improving overall performance.
[0057] (4) In particular, the dynamic weight fusion module contains residual connections. During the training process, the original input features are retained through residual connections, ensuring that the original information is not lost while the features are enhanced. At the same time, the autoregressive and cross-attention mechanisms are used to extract richer feature representations. For example, in the text causal generation task, the input features are long texts. The features enhanced by the residual autoregressive cross-attention module can better capture the local and global information in the long text, thereby improving the classification accuracy.
[0058] Step S3, use the unsupervised pre-training dataset X constructed in step S1 t , the feature enhancement pre-training of the large language model E is performed through the residual autoregressive cross-attention module constructed in step S2. The autoregressive module captures the sequence dependency, and the cross-attention module fuses the features of different texts to finally generate the enhanced feature representation.
[0059] like Figure 3 As shown in the figure, the pre-training process begins with data processing. First, the unlabeled data is segmented and masked, and then the processed data is fed into the model for training. During the training process, the model generates a predicted output through forward propagation, and then calculates the loss function and updates the model parameters through backpropagation. Here, the mean squared error (MSE) loss function is used to measure the difference between the model output and the target value. The calculation formula is:
[0060]
[0061] Where N is the number of samples, y i is the target value, is the output of the large language model. The MSE loss function calculates the square of the difference between each sample's prediction and the target value (true value) and averages it across all samples to produce a scalar loss value. This loss function penalizes the model's prediction error; larger errors result in higher loss values, thereby guiding the model to learn more accurate feature representations during pre-training and minimizing the gap between predictions and target values.
[0062] Step S4, use the supervised fine-tuning dataset X constructed in step S1 y Through supervised fine-tuning, the large language model E pre-trained in step S3 is post-trained until model training is complete. Supervised fine-tuning is a key step in post-training. The model further optimizes feature representation through supervised learning, making it better suited to specific downstream tasks. The residual autoregressive cross-attention module also helps the model better understand the underlying features of the text, playing a role in feature enhancement.
[0063] like Figure 4 As shown in Figure 2, the fine-tuning process is based on supervised datasets that contain text data labeled with specific task objectives. In supervised fine-tuning, the model generates task-specific prediction outputs through forward propagation, and then uses the cross-entropy loss function to measure the difference between the probability distribution of these prediction outputs and the probability distribution of the true label. The calculation formula is:
[0064]
[0065] In the formula, N is the number of samples, C is the number of categories, and y i,c is the probability distribution of the true label, The cross-entropy loss function calculates the probability distribution of the output for a large language model. The cross-entropy loss function calculates the log-likelihood of the true label and predicted probability for each sample and averages it across all categories and samples, resulting in a scalar loss value. The cross-entropy loss function is particularly well-suited for instruction fine-tuning tasks. It effectively measures the accuracy of the model's generated output, prompting the model to adjust parameters based on the cross-entropy loss to bring the predicted probability distribution closer to the probability distribution of the true label (the target distribution), thereby improving the model's generation performance and accuracy on specific tasks.
[0066] Figure 5The figure shows the decline curve of the cross-entropy loss function after post-training of the model in this embodiment. The light-colored curve represents the original cross-entropy loss function, and the dark-colored curve represents the smoothed loss function after smoothing the original cross-entropy loss function (the smoothing process can reflect the stability of the loss function). As can be seen from the figure, through supervised fine-tuning, the loss function of the model shows a steady decline during the post-training process. After 4000 steps, when the loss stabilizes at around 0.3, it indicates that the model training is complete.
[0067] In step S5, the large language model E trained in step S4 is subjected to inference testing in the NPU environment, and the performance of the trained model is evaluated.
[0068] This example deploys a trained large language model E in an Ascend NPU environment for inference testing. Ascend NPUs offer high-performance computing and low energy consumption. Through hardware acceleration, the model achieves higher efficiency during inference. For example, in natural language processing tasks, the model can quickly generate high-quality text content.
[0069] During the inference testing phase, the model is rigorously tested on multiple test datasets. These evaluations cover a variety of tasks, including basic language capabilities, multilingual support, instruction following, mathematical reasoning, code generation, and long-context processing. The model's performance is judged by comparing the accuracy of its text generation output with that of other similar models.
[0070] By comparing the results with the true labels of the test dataset, we calculate performance metrics such as precision, recall, and F1 score to comprehensively evaluate the model's performance on specific tasks. Furthermore, we can measure the model's inference speed and resource usage to ensure its feasibility and efficiency in actual deployment.
[0071] Evaluate model performance using metrics such as long-context handling and the relevance of generated text to ensure the model's effectiveness in real-world applications. Long-context handling is a key model characteristic, and key evaluation methods include NeedleBench and LV-Eval. NeedleBench evaluates the model's performance in extremely long contexts, including text up to 120,000 bytes. LV-Eval assesses the model's logical consistency and text generation quality in long text contexts. The relevance of generated text is evaluated using the N-gram model:
[0072] For a sentence consisting of n words W = w1w2…w n , w1 is the first word, w n is the last word, and the N consecutive words in a sentence are called N-grams; suppose the i-th word w in the sentence i The probability of occurrence depends only on the previous i-1 words, that is, Then you can i Take a short window of N-1 yuan in front Approximate calculation of w i The probability of occurrence is calculated as follows:
[0073]
[0074] Where, w i The previous N-1 element window w i-(N-1) ,…,w i-1 ; P is probability; Count is the number of statistics; numerator Contains word w i N yuan w i0(N01) ,…,w i The number of times a word appears in the corpus; the denominator is the number of times N-1 gram appears in the corpus.
[0075] Through the above steps, this embodiment significantly improves the feature extraction capability of the large language model for natural language text. t Combined with the residual autoregressive cross attention module, the model can more effectively capture the long-range dependencies in the text and enhance the ability to understand multi-source input. In the supervised fine-tuning stage, the supervised fine-tuning dataset X is used. y and cross entropy loss function, the model further optimizes the performance of specific tasks, making the predicted probability distribution closer to the target distribution, thereby improving the generation performance and accuracy.
[0076] This embodiment, through the design of residual connections and cross-attention mechanisms, not only retains most of the capabilities of the original large language model, but also flexibly adjusts the contribution ratios of different modules through a dynamic weight fusion module, further improving the model's stability and generalization capabilities. Furthermore, this embodiment reduces the difficulty of using large language models in natural language processing tasks, making them easy for non-experts to apply, thereby improving overall work efficiency.
[0077] Example 2:
[0078] This embodiment is an example of the actual application of embodiment one. This embodiment uses the method of embodiment one to train a large language model with 8B parameters, and evaluates it with the existing models of the same parameter level Qwen2-7B, Llama3-8B, and Llama3.1-8B commonly used in the industry, and compares the scoring results. The data sets used in the evaluation process include MMLU, CMMLU, C-Eval, AED, and APD. Among them, the first three are general data sets used to measure the knowledge and reasoning ability of large language models in different fields and difficulty levels. The MMLU language is English, and CMMLU and C-Eval are Chinese, especially for evaluation in the Chinese context. The latter two, AED and APD, are aerospace professional field data sets, specifically used to evaluate the understanding and generation capabilities in the aerospace field. As shown in Table 1 ("\" means that the data set was not tested):
[0079] Table 1 Comparison results between the 8B parameter model of this embodiment and models of the same level
[0080]
[0081] As can be seen from the table, the MMLU score of the 8B parameter model of this embodiment is higher than that of Llama3.1-8B and Llama3-8B, and is only 0.10 points different from Qwen2-7B. In terms of the Chinese general data set, the scores of CMMLU and C-Eval of this embodiment are almost the same as those of Qwen2-7B. However, in the field of aerospace expertise, the score of this embodiment is much higher than that of Qwen2-7B, both about 30 points higher. The comparison results show that the knowledge and reasoning ability of this embodiment in the general fields and difficulty levels of Chinese and English are almost the same as those of the existing common model Qwen2-7B with the same parameter level, but it performs better in the field of aerospace expertise. Therefore, this embodiment is particularly suitable for use scenarios that are professional knowledge-intensive and supplemented with some general domain knowledge.
[0082] Finally, the above is only used to illustrate the technical solution of the present invention and is not intended to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention (such as the use of various formulas, the sequence of steps, etc.) can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A feature enhancement method based on residual autoregressive cross-attention mechanism, characterized in that: The following steps are involved: Step S1, construct an unsupervised pre-training dataset (X t ) and supervised fine-tuning dataset (X y ), the unsupervised pre-training dataset is an unlabeled dataset, and the supervised fine-tuning dataset is a labeled dataset; Step S2: constructing a large language model (E) based on the decoder, and constructing a residual autoregressive cross-attention module in the large language model (E), comprising an autoregressive module, a cross-attention module, and a dynamic weight fusion module, wherein the dynamic weight fusion module contains a residual connection, and the original input features are retained through the residual connection; Step S3, using the unsupervised pre-training dataset (X t ), performing feature enhancement pre-training on the large language model (E) through the residual autoregressive cross-attention module constructed in step S2; Step S4, using the supervised fine-tuning dataset (X y ), through supervised fine-tuning, the large language model (E) pre-trained in step S3 is post-trained until the model training is completed; In step S5, the large language model (E) trained in step S4 is subjected to inference testing in the NPU environment, and the performance of the trained large language model (E) is evaluated. The evaluation indicators include long-context processing capability and the relevance of the generated text.
2. The feature enhancement method based on residual autoregressive cross-attention mechanism according to claim 1 is characterized in that The unlabeled data set in step S1 includes text data.
3. The feature enhancement method based on residual autoregressive cross-attention mechanism according to claim 1 is characterized in that The labeled data set in step S1 includes instructions, inputs, and outputs.
4. The feature enhancement method based on residual autoregressive cross-attention mechanism according to claim 1, characterized in that The autoregressive module in step S2 contains a self-attention mechanism, which is used to capture the dependencies of the input long text sequence. The calculation formula of the self-attention mechanism is: Where Q is the query matrix, K is the key matrix, V is the value matrix, T is the matrix transpose operator, softmax is the Softmax function, that is, the normalized exponential function, and d k is the dimension of matrix K.
5. The feature enhancement method based on residual autoregressive cross-attention mechanism according to claim 1, characterized in that The cross-attention module in step S2 contains a cross-attention mechanism for fusing features of different text data. The calculation formula of the cross-attention mechanism is: Where Q1 is the query matrix of the first input sequence, K2 is the key matrix of the second input sequence, and V2 is the value matrix of the second input sequence. is the dimension of matrix K2.
6. The feature enhancement method based on residual autoregressive cross-attention mechanism according to claim 1, characterized in that The dynamic weight fusion calculation formula of the dynamic weight fusion module in step S2 is: t f =α·y p +(1-a)·y d Where y p is the output of the residual autoregressive cross attention module, y d is the output of the decoder module, y f is the weighted sum, α is the weight of the residual autoregressive cross attention module, α∈[0,1], and 1-α is the weight of the decoder.
7. The feature enhancement method based on residual autoregressive cross-attention mechanism according to claim 1, characterized in that The feature enhancement pre-training in step S3 uses the mean square error loss function to measure the difference between the output of the large language model (E) and the target value. The mean square error MSE calculation formula is: Where N is the number of samples, y i is the target value, is the output of the large language model.
8. The feature enhancement method based on residual autoregressive cross-attention mechanism according to claim 1, characterized in that The specific process of post-training the large language model (E) pre-trained in step S3 through supervised fine-tuning in step S4 until the model training is completed is as follows: In supervised fine-tuning, the large language model generates a predicted output for a specific task through forward propagation, and uses the cross-entropy loss function to measure the difference between the probability distribution of the predicted output and the probability distribution of the true label. The cross-entropy calculation formula is: In the formula, N is the number of samples, C is the number of categories, and y i,c is the probability distribution of the true label, Predicting the probability distribution of output for a large language model; During post-training, the large language model adjusts parameters according to the cross-entropy loss function so that the probability distribution of the predicted output is closer to the probability distribution of the true label.
9. The feature enhancement method based on residual autoregressive cross-attention mechanism according to claim 1, characterized in that The evaluation method for the long context processing capability in step S5 includes NeedleBench and LV-Eval.
10. The feature enhancement method based on residual autoregressive cross-attention mechanism according to claim 1, characterized in that: The evaluation method of the relevance of the generated text in step S5 adopts the N-gram model: for a sentence consisting of n words W = w1w2…w n , w1 is the first word, w n is the last word, and the N consecutive words in a sentence are called N-grams; suppose the i-th word w in the sentence i The probability of occurrence depends only on the previous i-1 words, that is, Then you can i Take a short window of N-1 yuan in front Approximate calculation of w i The probability of occurrence is calculated as follows: Where, w i The previous N-1 element window w i-(N-1) ,…,w i-1 ; P is probability; Count is the number of statistics; numerator Contains word w i N yuan w i0(N01) ,…,w i The number of times a word appears in the corpus; the denominator is the number of times N-1 gram appears in the corpus.