Text generation method based on mixed sparse attention neural network

By mixing sparse attention neural network and cache identifier technology, the computing complexity and resource allocation of text generation neural networks are optimized, and the calculation complexity and resource consumption of traditional attention mechanisms in high-dimensional complex data processing is solved, and efficient text generation and complex task processing is achieved.

CN119990119APending Publication Date: 2025-05-13SICHUAN UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510063782.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional attention mechanisms have high computational complexity and high resource consumption when processing high-dimensional complex data, which limits its application in large-scale model training, especially in specific areas where data quality is uneven and sample size is limited.

Method used

A hybrid sparse attention neural network is adopted to build a text generation neural network by retaining key identifiers and sparsely processing of unimportant identifiers, combining cache identifier technology to optimize computing complexity and resource allocation.

Benefits of technology

It effectively reduces the computational complexity, optimizes resource allocation, improves the inference speed and efficiency when processing long text and complex tasks, and ensures the consistency and accuracy of text generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990119A_ABST
    Figure CN119990119A_ABST
Patent Text Reader

Abstract

The invention provides a text generation method based on a hybrid sparse attention neural network, and relates to the technical field of natural language processing, and the method comprises the steps: carrying out the word segmentation of a text collection data set, and obtaining a training set, a verification set and a test set; constructing a text generation neural network by using a mixed sparse attention mechanism; inputting the training set into a text generation neural network, and training by using autoregression and autocoding to obtain a trained text generation neural network; inputting the verification set and the test set into the trained text generation neural network, and performing performance evaluation by using three indexes of reasoning time, video memory occupation and confusion to obtain a final text generation neural network; wherein the final text generation neural network is used for analyzing the text collection data to obtain a text generation result, and text generation based on the hybrid sparse attention neural network is completed. According to the method, the problems of high calculation complexity and large resource consumption of long text data are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a text generation method based on a hybrid sparse attention neural network. Background Art

[0002] The development of large-scale language models (LLMs) has greatly promoted the progress of natural language processing (NLP) technology. Relying on massive amounts of text data, these models can capture complex patterns in the data through attention mechanisms and achieve a variety of tasks from language generation to semantic understanding. However, as the scale of the model continues to expand, its demand for computing resources and data quality also increases. In practical applications, the problems of uneven data quality and limited sample size often become challenges, especially in certain specific fields, where data usually have the characteristics of scarce samples, complex formats, and lack of structured labels, which are significantly different from the training data of standard large-scale language models. Therefore, how to effectively apply these large models to process complex data has become a technical problem that needs to be solved urgently. When processing high-dimensional complex data, the traditional attention mechanism has high computational complexity and high resource consumption, which limits its application in large-scale model training. Summary of the invention

[0003] In view of the above-mentioned deficiencies in the prior art, the present invention provides a text generation method based on a hybrid sparse attention neural network, which solves the problems of high computational complexity and large resource consumption of long text data.

[0004] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: a text generation method based on a hybrid sparse attention neural network, comprising:

[0005] S1: Perform word segmentation on the text collection dataset to obtain the training set, validation set and test set respectively;

[0006] S2: Use hybrid sparse attention mechanism to build a text generation neural network;

[0007] S3: inputting the training set into the text generation neural network, and training it by using autoregression and autoencoding to obtain a trained text generation neural network;

[0008] S4: Input the verification set and the test set into the trained text generation neural network, and use the three indicators of inference time, video memory occupancy and perplexity to perform performance evaluation to obtain a final text generation neural network; wherein the final text generation neural network is used to analyze the text collection data to obtain the text generation result, and complete the text generation based on the hybrid sparse attention neural network.

[0009] Furthermore, the S1 includes:

[0010] Clean the collected text data to obtain pure text data;

[0011] Using subword units, performing word segmentation processing on the clean text data to obtain language unit data;

[0012] The language unit data is divided in proportion to obtain a training set, a validation set and a test set.

[0013] Furthermore, the text generation neural network includes:

[0014] The starting module is used to analyze the text data in the training set, extract and generate a set of fixed starting identifiers;

[0015] A cache module, used to generate an identifier based on the starting identifier, calculate the attention weight between the current newly added identifier and the generated identifier, and cache the generated identifier using a cache identifier technology to obtain a cache identifier;

[0016] A generation module is used to perform weighted calculation based on the cache identifier using a hybrid sparse attention mechanism to obtain hybrid sparse attention weights of each identifier; based on the hybrid sparse attention weights, key identifiers are dynamically selected to obtain text generation results.

[0017] Furthermore, the expression of the hybrid sparse attention weight is:

[0018]

[0019] in, represents the mixed sparse attention weight, Q i represents the query matrix of the i-th query, K j represents the key matrix of the jth key, K m represents the key matrix of the mth key, M (i) Indicates that Q i The index set of the related k elements.

[0020] Furthermore, the expression of the autoregressive result is:

[0021] P(x t |x1,x2,...,x t-1 ) = softmax(W·t);

[0022] Among them, P(x t |x1,x2,...x t-1 ) represents the autoregressive result, softmax represents the softmax function, W represents the output weight matrix, and t represents the tth position in the sequence.

[0023] The beneficial effects of the present invention are as follows: a text generation neural network constructed by a hybrid sparse attention mechanism is used to analyze a text collection data set to obtain a text generation result. (1) By retaining key identifiers and sparsely processing unimportant identifiers (i.e., ignoring or reducing their calculation weights), and adopting a cache identifier technology, the generated identifiers will be stored in the cache, thereby avoiding repeated calculation of the contribution of the same identifier to the attention weight. This method effectively reduces computational complexity, optimizes resource allocation, and improves the reasoning speed and efficiency when processing long texts and complex tasks; (2) By locking the starting identifier, the consistency of global semantics in the generation process is ensured to prevent topic drift, especially in the generation of long texts, the theme can be kept consistent; (3) By dynamically screening and highlighting the key information in the text, it is ensured that the model can efficiently process complex semantic structures, accurately focus on the most important content in the generation process, and reduce the interference of irrelevant information; (4) By the comprehensive application of locking the starting identifier, dynamically screening key identifiers, and caching identifier technology, the consistency and accuracy of text generation are improved, and the reasoning efficiency and resource usage are greatly optimized, providing a more efficient and resource-friendly solution for the processing of long texts and complex tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] This specification will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein:

[0025] Figure 1 This is an exemplary flowchart of a text generation method based on a hybrid sparse attention neural network as shown in some embodiments of this specification. DETAILED DESCRIPTION

[0026] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.

[0027] Example

[0028] Figure 1 is an exemplary flow chart of a text generation method based on a hybrid sparse attention neural network according to some embodiments of this specification. Figure 1 As shown, the process includes the following steps. In some embodiments, the process can be executed by a processor.

[0029] S1: Perform word segmentation on the text collection dataset to obtain the training set, validation set and test set respectively.

[0030] Text collection datasets are text data sources covering various knowledge fields and languages. For example, text collection data can include Wikipedia data and Baidu Encyclopedia data.

[0031] The training set is a text dataset used to train the text generation neural network.

[0032] The validation set is a text dataset used to validate the text generation neural network.

[0033] The test set is a text dataset used to test the text generation neural network.

[0034] In some embodiments, the processor can implement S1 based on the following steps: perform data cleaning on the text collection data to obtain pure text data; use sub-word units to segment the pure text data to obtain language unit data; divide the language unit data according to proportion to obtain a training set, a verification set and a test set respectively.

[0035] Clean text data is text data that has been freed of irrelevant information.

[0036] In some embodiments, the processor may clean the collected text data to remove irrelevant HTML tags, special characters, advertisements and other information to obtain pure text data.

[0037] The language unit data is text data which is a basic unit of simulation input.

[0038] In some embodiments, the processor may use sub-word units (such as BPE) to segment the text into smaller language units to obtain language unit data.

[0039] In this way, the sparsity of the data set can be reduced and the processing efficiency can be improved.

[0040] In some embodiments, the processor may divide the language unit data in a ratio of 8:1:1 to obtain a training set, a validation set, and a test set; wherein the training set, the validation set, and the test set all cover texts in various fields and categories.

[0041] S2: Use the hybrid sparse attention mechanism to build a text generation neural network.

[0042] The text generation neural network is used to generate long texts and obtain text generation results. There are many types of text generation neural networks. For example, the type of text generation neural network can include a multi-layer Transformer structure.

[0043] In some embodiments, the input of the text generation neural network may be language unit data, and the output of the text generation neural network may be a text generation result.

[0044] In some embodiments, the structure of the text generation neural network is as follows:

[0045] The text generation neural network includes a starting module, a cache module and a generation module. The output of the starting module is used as the input of the cache module, the output of the cache module is used as the input of the generation module, and the output of the generation module is used as the final output of the text generation neural network.

[0046] The start module is used to analyze the text data in the training set, extract and generate a set of fixed start identifiers. The input of the start module may include language unit data, and the output may include start identifiers.

[0047] The starting identifier is the starting content that is strongly associated with the text topic and represents the core topic of the text, ensuring that the generated content has a clear direction and consistency at the beginning. Each identifier generated thereafter adjusts the direction of the generated content through its correlation with the starting identifier to ensure that topic drift does not occur.

[0048] In some embodiments, the processor may utilize a starting module to analyze text information in a training set, obtain a starting identifier, and fix it.

[0049] By locking the start tokens, the model can always keep its focus on the start content of the text during the generation process. In this way, the topic coherence of the generated text can be ensured, and the topic drift phenomenon that occurs as the generation progresses can be avoided; by locking the start tokens, the model establishes a global semantic framework, ensuring that the generated text revolves around the core topic and avoiding the generation of irrelevant or off-topic content.

[0050] The cache module is used to generate an identifier based on the starting identifier, calculate the attention weight between the current newly added identifier and the generated identifier, and cache the generated identifier using the cache identifier technology to obtain a cache identifier. The input of the cache module may include the starting identifier, and the output may include the cache identifier.

[0051] In some embodiments, the cache module may include a multi-layer Transformer structure with a self-attention mechanism, which increases attention to the starting identifier by assigning a higher weight to the starting identifier, maintains the topic focus of the generated content, and generates identifiers; wherein the input of each layer of the Transformer structure is a text sequence representation after word segmentation processing (i.e., language unit data), and through self-attention calculation, the self-attention result of each identifier is obtained for generating identifiers.

[0052] In some embodiments, the expression of the self-attention result can be:

[0053]

[0054] Among them, Attension represents the self-attention result, Q represents the query vector, K represents the key vector, V represents the value vector, and d k Indicates the dimension of the key vector.

[0055] Cache identifiers are identifiers that have been generated in the text generation neural network.

[0056] In some embodiments, the processor may utilize a cache module to cache the generated identifiers, and through incremental calculation, process only the newly added identifiers to obtain cached identifiers.

[0057] By using cached tokens technology, the generated tokens are cached to avoid repeated calculations, especially in long text generation tasks. As the text is gradually generated, only the newly added tokens are processed through incremental calculations, reducing a lot of repetitive work. It can not only improve the reasoning efficiency, but also significantly reduce the memory usage, making the model more efficient in long text generation and large-scale tasks.

[0058] The generation module is used to perform weighted calculation based on the cache identifier using a hybrid sparse attention mechanism to obtain a hybrid sparse attention weight of each identifier; based on the hybrid sparse attention weight, dynamically select a key identifier to obtain a text generation result. The input of the generation module may include a cache identifier, and the output may include a text generation result.

[0059] The hybrid sparse attention weight is the weight that combines the local and global attention results.

[0060] In some embodiments, the processor can utilize a hybrid sparse attention mechanism, combining local and global attention calculations. For local areas, only the attention of adjacent identifiers is calculated; for global key identifiers, global attention calculations are performed to obtain hybrid sparse attention weights for each identifier; based on the hybrid sparse attention weights, identifiers are screened to obtain key identifiers. As text generation progresses, the model will re-evaluate which identifiers play a key role in the current context, that is, dynamically screen key identifiers and adjust the focus of attention to obtain text generation results.

[0061] In this way, each input element is restricted to interact with only some sequence elements, and the number of attention weights required for calculation is significantly reduced. This can greatly improve computational efficiency while retaining key contextual information.

[0062] In some embodiments, the expression of the hybrid sparse attention weight may be:

[0063]

[0064] in, represents the mixed sparse attention weight, Q i represents the query matrix of the i-th query, K j represents the key matrix of the jth key, K m represents the key matrix of the mth key, M (i) Indicates that Q i The index set of the related k elements.

[0065] The text generation result is the generated text data result.

[0066] In some embodiments, the processor may filter the identifiers based on the hybrid sparse attention weights, retaining only the top k most relevant attention weights, to obtain a text generation result.

[0067] By precisely focusing on key tokens, the model is able to dynamically select these tokens. As text generation progresses, the model will continuously evaluate which tokens play a key role in the current context and adjust its focus accordingly. This process enables the model to more accurately process semantic structures in complex tasks by dynamically filtering and highlighting the core information in the text. By retaining key tokens and removing unnecessary ones, the hybrid sparse attention mechanism optimizes the model's allocation of computing resources and concentrates computing power on the most important parts, which not only improves the inference speed but also reduces resource consumption; it significantly enhances the model's understanding and expression of deep semantics, ensuring that the generated content is more accurate and logical.

[0068] S3: Input the training set into the text generation neural network, and perform training using autoregression and autoencoding to obtain a trained text generation neural network.

[0069] In some embodiments, the text generation neural network can be obtained by training with a training set. For example, the training set can be input into the initial text generation neural network, and trained using autoregression and autoencoding. In the autoregression task, the structure of the language is modeled by predicting the next identifier after the current identifier; in the autoencoding task, the model optimizes the representation ability of the language by masking some identifiers and predicting their values. Using the Adam optimizer and multi-GPU parallel training strategy, the parameters of the initial text generation neural network are iteratively updated, and learning rate scheduling, gradient accumulation and regularization methods are used to ensure the stability of the model. When the preset conditions are met, the model training is completed, and a trained text generation neural network is obtained. Among them, the preset conditions can be that the number of iterations reaches a threshold, etc.

[0070] In some embodiments, the expression of the autoregressive result may be:

[0071] P(x t |x1,x2,...,x t-1 ) = softmax(W·t);

[0072] Among them, P(x t |x1,x2,...x t-1 ) represents the autoregressive result, softmax represents the softmax function, W represents the output weight matrix, and t represents the tth position in the sequence.

[0073] S4: Input the verification set and the test set into the trained text generation neural network, and use the three indicators of inference time, video memory occupancy and perplexity to perform performance evaluation to obtain a final text generation neural network; wherein the final text generation neural network is used to analyze the text collection data to obtain the text generation result, and complete the text generation based on the hybrid sparse attention neural network.

[0074] Inference time is a metric that reflects how efficiently a model processes input data and generates output.

[0075] Memory usage is a measure of the memory requirements of the model during inference.

[0076] Perplexity is a metric used to evaluate the generative ability of a model.

[0077] In some embodiments of the present specification, a text generation neural network constructed by a hybrid sparse attention mechanism is used to analyze a text collection data set to obtain a text generation result. (1) By retaining key identifiers and sparsely processing unimportant identifiers (i.e., ignoring or reducing their calculation weights), and using cache identifier technology, the generated identifiers will be stored in the cache, thereby avoiding repeated calculation of the contribution of the same identifier to the attention weight. This method effectively reduces computational complexity, optimizes resource allocation, and improves the reasoning speed and efficiency when processing long texts and complex tasks; (2) By locking the starting identifier, the consistency of global semantics in the generation process is ensured to prevent topic drift, especially in the generation of long texts to maintain the consistency of the theme; (3) By dynamically screening and highlighting the key information in the text, it is ensured that the model can efficiently process complex semantic structures, accurately focus on the most important content in the generation process, and reduce the interference of irrelevant information; (4) By locking the starting identifier, dynamically screening key identifiers and caching identifier technology, the coherence and accuracy of text generation are improved, and the reasoning efficiency and resource usage are greatly optimized, providing a more efficient and resource-friendly solution for the processing of long texts and complex tasks.

Claims

1. A text generation method based on a hybrid sparse attention neural network, characterized in that: include: S1: Perform word segmentation on the text collection dataset to obtain the training set, validation set and test set respectively; S2: Use hybrid sparse attention mechanism to build a text generation neural network; S3: inputting the training set into the text generation neural network, and training it by using autoregression and autoencoding to obtain a trained text generation neural network; S4: Input the verification set and the test set into the trained text generation neural network, and use the three indicators of inference time, video memory occupancy and perplexity to perform performance evaluation to obtain a final text generation neural network; wherein the final text generation neural network is used to analyze the text collection data to obtain the text generation result, and complete the text generation based on the hybrid sparse attention neural network.

2. The text generation method based on hybrid sparse attention neural network according to claim 1, characterized in that: The S1 includes: Clean the collected text data to obtain pure text data; Using subword units, performing word segmentation processing on the clean text data to obtain language unit data; The language unit data is divided in proportion to obtain a training set, a validation set and a test set.

3. The text generation method based on hybrid sparse attention neural network according to claim 1, characterized in that: The text generation neural network comprises: The starting module is used to analyze the text data in the training set, extract and generate a set of fixed starting identifiers; A cache module, used to generate an identifier based on the starting identifier, calculate the attention weight between the current newly added identifier and the generated identifier, and cache the generated identifier using a cache identifier technology to obtain a cache identifier; A generation module is used to perform weighted calculation based on the cache identifier using a hybrid sparse attention mechanism to obtain hybrid sparse attention weights of each identifier; based on the hybrid sparse attention weights, key identifiers are dynamically selected to obtain text generation results.

4. The text generation method based on hybrid sparse attention neural network according to claim 3 is characterized in that: The expression of the hybrid sparse attention weight is: in, represents the mixed sparse attention weight, Q i represents the query matrix of the i-th query, K j represents the key matrix of the jth key, K m represents the key matrix of the mth key, M (i) Indicates that Q i The index set of the related k elements.

5. The text generation method based on hybrid sparse attention neural network according to claim 1, characterized in that: The expression of the autoregressive result is: P(x t |x1,x2,...,x t-1 )=softmax(W·t); Among them, P(x t |x1,x2,...x t-1 ) represents the autoregressive result, softmax represents the softmax function, W represents the output weight matrix, and t represents the tth position in the sequence.

Citation Information

Cited By

  • Government-affair-oriented Internet of Things data sharing and data quality inspection method, system and device based on large model, and storage medium

    CN120670415A