Text sequence generation method and device of intelligent agent, equipment, medium and product
Through the combination of partial cache and complete cache models, the problems of slow generation of ultra-long text sequences and high repetition in the existing technology are solved, and the lossless acceleration and efficient generation of 100K text sequences are achieved, which improves the natural language processing capabilities of the language model and the language interaction capabilities of the agent.
Patent Information
- Application Number
- CN202510205897.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
AI Technical Summary
The existing language models have the problem of unsatisfactory acceleration effects in the generation of ultra-long text sequences, especially when generating ultra-long text sequences, they are prone to repeated and meaningless messy symbols, and cannot effectively improve language interaction capabilities.
The partial cache model is used to generate candidate draft text, the target draft text is generated through the complete cache model, and the target text sequence is output based on the matching degree of the candidate draft text and the target draft text. The dynamic key-value pruning of the partial cache model and the text processing of the complete cache model are used to achieve lossless acceleration of the ultra-long text sequence.
The lossless acceleration of ultra-long text sequence generation is achieved, and the world's first 100K text sequence acceleration is achieved, maintaining the lossless accuracy of the target language model, and achieving more than 3 times acceleration in different prefix lengths, model architecture and model scales, significantly improving the natural language processing capabilities of the language model and the language interaction capabilities of the agent.
Smart Images

Figure CN120146041A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, specifically to the field of data processing technology, and more specifically to a method, device, equipment, medium and product for generating text sequences of an intelligent agent. Background Art
[0002] Artificial Intelligence (AI for short) is an important driving force for the new round of scientific and technological revolution and industrial transformation, and is a new key technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. As an important part of intelligent science, artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine (i.e., an intelligent agent) that can respond in a way similar to human intelligence.
[0003] Natural language processing is a very core research direction in the field of artificial intelligence technology, mainly reflected in language models based on deep learning. Usually, such models refer to deep learning models trained with a large amount of text data, enabling the model to generate natural language text or understand the meaning of language text. These models can provide in-depth knowledge and language production on various topics through training on a huge dataset, and its core purpose is to simulate the human language cognition and generation process to a certain extent.
[0004] The generation of text sequences by a language model is an important aspect reflecting the language processing ability of the language model, mainly reflected in the accuracy and generation speed of text sequence generation, and can be applied to various scenarios such as interactive chat and document analysis. Usually, with the increase in the scale of processed text data (such as the generation of long text sequences), the existing language models are not suitable for processing ultra-long text sequences. To accelerate the generation speed of ultra-long text sequences, it is necessary to consider lossless acceleration of the text processing process of the language model. However, the lossless acceleration effect in the existing solutions is not ideal, and lossless acceleration for the generation of ultra-long text sequences cannot be performed. Summary of the Invention
[0005] In view of at least one of the above problems existing in the process of generating text sequences by a language model in the prior art, embodiments of the present invention aim to provide a method, device, equipment, medium and product for generating text sequences of an intelligent agent that can achieve lossless acceleration of generating ultra-long text sequences by a language model, so as to significantly improve the language interaction ability of the intelligent agent.
[0006] One aspect of an embodiment of the present invention provides a method for generating a text sequence of an agent, which includes: generating a candidate draft text corresponding to the received prefix input text based on a partial cache model; generating a target draft text corresponding to the candidate draft text through a complete cache model; and outputting a target text sequence according to the matching degree between the candidate draft text and the target draft text.
[0007] According to an embodiment of the present invention, in generating a candidate draft text corresponding to the received prefix input text based on a partial cache model, it includes: outputting logical distribution information corresponding to the prefix input text according to the linear layer data of the partial cache model.
[0008] According to an embodiment of the present invention, in generating a candidate draft text corresponding to the received prefix input text based on a partial cache model, it further includes: obtaining candidate word data corresponding to the prefix input text according to the logical distribution information; generating a candidate text path corresponding to the candidate word data through a preset tree attention rule; and generating a candidate draft text according to the candidate text path.
[0009] According to an embodiment of the present invention, in obtaining candidate word data corresponding to the prefix input text according to the logical distribution information, it includes: obtaining preselected word data corresponding to the prefix input text according to the logical distribution information; and obtaining candidate word data corresponding to the preselected word data through a preset context penalty rule.
[0010] According to an embodiment of the present invention, in generating a candidate draft text according to the candidate text path, it includes: querying draft text data in a preset text item sequence table that matches the candidate text path; and generating a candidate draft text according to the candidate text path and the draft text data.
[0011] According to an embodiment of the present invention, in outputting a target text sequence according to the matching degree between the candidate draft text and the target draft text, it includes: obtaining a valid text sequence according to the matching degree between the candidate draft text and the target draft text; and outputting a target text sequence when the valid text capacity of the valid text sequence meets a preset target text capacity.
[0012] According to an embodiment of the present invention, after obtaining a valid text sequence according to the matching degree between the candidate draft text and the target draft text, it further includes: updating the preset text item sequence table, the partial cache model, and the complete cache model according to the valid text sequence.
[0013] Another aspect of an embodiment of the present invention provides a text sequence generation device for an intelligent agent, which includes a candidate text generation module, a target text generation module, and a target sequence output module. The candidate text generation module is used to generate a candidate draft text corresponding to the received prefix input text based on a partial cache model; the target text generation module is used to generate a target draft text corresponding to the candidate draft text through a complete cache model; and the target sequence output module is used to output a target text sequence according to the matching degree between the candidate draft text and the target draft text.
[0014] Another aspect of an embodiment of the present invention provides an electronic device, including one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned text sequence generation method for the intelligent agent.
[0015] Another aspect of an embodiment of the present invention provides a computer-readable storage medium, on which executable instructions are stored. When the instructions are executed by a processor, the processor is caused to execute the above-mentioned text sequence generation method for the intelligent agent.
[0016] Another aspect of an embodiment of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the above-mentioned text sequence generation method for the intelligent agent is implemented.
[0017] The text sequence generation method for the intelligent agent provided by the embodiment of the present invention can at least partially solve the problem of relatively low intelligence level existing in the process of intelligent agent text language processing in the related art, and thus can at least achieve one of the following technical effects:
[0018] For the first time globally, the generation of ultra-long text sequences is accelerated to 100K, while maintaining the lossless accuracy of the target language model. Moreover, it can always achieve more than 3 times acceleration for different prefix lengths, model architectures, and model scales, greatly reducing the required time of the traditional autoregressive process (nearly 5 hours) to only 90 minutes. In addition, as the length of the text sequence generation increases, it also has a higher acceleration ability than the traditional autoregressive technology, while enhancing the diversity of ultra-long sequence generation, directly breaking through the lossless acceleration bottleneck of ultra-long text sequences in the prior art, significantly improving the natural language processing ability of the language model in terms of low latency and high throughput, significantly enhancing the language interaction ability and intelligence level of the intelligent agent, and achieving extremely successful scientific research value and commercial application value.
[0019] It should be understood that the above general description and the following specific embodiments are only exemplary and explanatory, and they do not limit the scope of what the present invention intends to claim. Description of the Drawings
[0020] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the accompanying drawings:
[0021] Figure 1 Schematically shows an application scenario diagram of a method, apparatus, device, medium, and program product for generating a text sequence of an agent according to an embodiment of the present invention;
[0022] Figure 2 Schematically shows a flowchart of a method for generating a text sequence of an agent according to an embodiment of the present invention;
[0023] Figure 3 Schematically shows a flowchart of an application scenario of a method for generating a text sequence of an agent according to an embodiment of the present invention;
[0024] Figure 4A Schematically shows a comparison table of acceleration performance data of a method for generating a text sequence of an agent according to an embodiment of the present invention on different architecture models;
[0025] Figure 4B Schematically shows a line graph of the changes in acceleration and acceptance rate corresponding to the increase in text capacity of a method for generating a text sequence of an agent according to an embodiment of the present invention;
[0026] Figure 4C Schematically shows a comparison table of data on the acceleration of a method for generating a text sequence of an agent according to an embodiment of the present invention on different scale models;
[0027] Figure 5 Schematically shows a structural block diagram of a device for generating a text sequence of an agent according to an embodiment of the present invention; and
[0028] Figure 6 Schematically shows a block diagram of an electronic device suitable for implementing a method for generating a text sequence of an agent according to an embodiment of the present invention.
[0029] The above-mentioned accompanying drawings are a part of the description of the embodiments of the present invention, which illustrate the exemplary embodiments of the present invention. The accompanying drawings and the description of the specification are used together to explain the principles of the embodiments of the present invention. It should be understood that the above general description of the accompanying drawings and the following specific embodiments are only exemplary and explanatory, and they cannot limit the scope claimed by the present invention. Detailed Description
[0030] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer and more understandable, the spirit of the content disclosed by the present invention will be clearly explained below with reference to the drawings and detailed descriptions. After any person skilled in the relevant technical field understands the embodiments of the content of the present invention, they can make changes and modifications based on the technologies taught by the content of the present invention, without departing from the spirit and scope of the content of the present invention.
[0031] The schematic embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention. Additionally, elements / components with the same or similar reference numerals used in the drawings and embodiments represent the same or similar parts.
[0032] Regarding the use of "first", "second",... etc. in the present invention, they do not particularly refer to the order or sequence, nor are they used to limit the present invention. They are only used to distinguish elements or operations described with the same technical terms.
[0033] Regarding the directional terms used in the present invention, such as: up, down, left, right, front or back, etc., they only refer to the directions in the drawings. Therefore, the directional terms used are for explanation and not for limiting the present creation.
[0034] Regarding the use of "comprising", "including", "having", "containing", etc. in the present invention, they are all open-ended terms, meaning including but not limited to.
[0035] Regarding the use of "and / or" in the present invention, it includes any one or all combinations of the described things.
[0036] Regarding "a plurality of" in the present invention, it includes "two" and "more than two"; regarding "a plurality of groups" in the present invention, it includes "two groups" and "more than two groups".
[0037] Regarding the terms "substantially", "about", etc. used in the present invention, they are used to modify any quantity or error that can vary slightly, but these slight variations or errors will not change their essence. Generally, the range of such slight variations or errors modified by such terms can be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments or other values in some embodiments. Those skilled in the art should understand that the aforementioned values can be adjusted according to actual needs and are not limited thereto.
[0038] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used here should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0039] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning that a person skilled in the art usually understands this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", generally, it should be interpreted according to the meaning that a person skilled in the art usually understands this expression (for example, "a system having at least one of A, B, or C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). A person skilled in the art should also understand that substantially any disjunctive conjunctions and / or phrases representing two or more alternative items, whether in the specification, claims, or drawings, should be understood as giving the possibility of including one of these items, either side of these items, or both items. For example, the phrase "A or B" should be understood as including the possibility of "A" or "B", or "A and B".
[0040] To improve the natural language text processing ability of language models, various solutions have been proposed in the prior art to address the problems of poor accuracy and slow generation speed during the generation of long text sequences. For example, a solution to extend long sequence generation by introducing Triforce to build a hierarchical reasoning decoding system and a solution to accelerate the high-throughput reasoning process for medium and long sequences by MagicDec. However, the above solutions usually adopt KV cache (Key-Value Cache) compression in the drafting stage, and their overall applicability is only limited to application scenarios with long prefixes and short outputs, and is not suitable for ultra-long sequence generation tasks. Therefore, the demand for efficient reasoning covering extended input contexts and long outputs is a challenge that the prior art currently cannot solve.
[0041] In addition, although language models can currently generate text exceeding 100K, due to the biases inherent in the model architectures used (such as the Transformer architecture), generating overly long content usually leads to repetition problems and may result in meaningless jumbled symbols. Therefore, how to achieve lossless acceleration in ultra-long sequence generation tasks is a very important application direction.
[0042] In view of at least one of the above problems existing in the text sequence generation process of language models in the prior art, embodiments of the present invention aim to provide a text sequence generation method, apparatus, device, medium and product for an agent that can achieve lossless acceleration of text sequence generation of a language model. For the first time, the generation of ultra-long text sequences is accelerated to 100K, while maintaining the lossless output of the target language model. Moreover, it can always achieve more than 3 times acceleration at different prefix lengths, model architectures and model scales, greatly reducing the time required for the traditional autoregressive process to only 90 minutes. In addition, as the length of the text sequence generation increases, it also has a higher acceleration ability than traditional autoregressive techniques, while enhancing the diversity of ultra-long sequence generation and significantly improving the language interaction ability of the agent.
[0043] One aspect of the embodiments of the present invention provides a text sequence generation method for an agent, which includes: generating a candidate draft text corresponding to the received prefix input text based on a partial cache model; generating a target draft text corresponding to the candidate draft text through a complete cache model; and outputting a target text sequence according to the matching degree between the candidate draft text and the target draft text.
[0044] Figure 1 Schematically shows an application scenario diagram of the text sequence generation method, apparatus, device, medium and program product for an agent according to an embodiment of the present invention.
[0045] As Figure 1 shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or fiber optic cables, etc.
[0046] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0047] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers and desktop computers, etc.
[0048] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0049] It should be noted that the text sequence generation method of the agent provided by the embodiments of the present invention can generally be executed by server 105. Correspondingly, the text sequence generation device of the agent provided by the embodiments of the present invention can generally be set in server 105. The text sequence generation method of the agent provided by the embodiments of the present invention can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the text sequence generation device of the agent provided by the embodiments of the present invention can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.
[0050] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0051] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the Figures 2 to 4C scenario described below, the text sequence generation method of the agent of the disclosed embodiments will be described in detail through
[0052] As Figure 2 shown, one aspect of the embodiments of the present invention provides a text sequence generation method of an agent, which includes operations S201 to S203.
[0053] In operation S201, a candidate draft text corresponding to the received prefix input text is generated based on a partial cache model;
[0054] In operation S202, a target draft text corresponding to the candidate draft text is generated through a complete cache model; and
[0055] In operation S203, a target text sequence is output according to the matching degree between the candidate draft text and the target draft text.
[0056] The agent can be the execution subject of the above text sequence generation method of the embodiments of the present invention, or can also be the execution party controlled by the text sequence generation method. Specifically, it can be a humanoid intelligent robot or other AI devices, which usually have actuators (such as microphones and speakers) and can complete specific action tasks. Other AI devices can be electronic devices with high-performance processors that can achieve intelligent interaction (such as language interaction) with users, and they can have specific intelligent AI operating systems or operation platforms.
[0057] The text sequence can be a sequence data composed of an ordered string of texts, which can be a word, a sentence, an article or even a novel. In the text sequence, each text part has a corresponding position, and the subsequent text part may be affected by the previous text part, and is concatenated by natural language logic. The text sequence can be used to represent natural language texts, such as sentences or paragraphs, and can be used in natural language processing tasks.
[0058] The partial cache model can be a language model (such as Large Language Model, abbreviated as LLM) optimized for partial key-value caching (Partial Key-Value Cache) for long text sequence generation tasks. The key-value cache can store the calculation results of previous tokens for reuse in subsequent generations, thus effectively avoiding redundant calculations. Among them, the design of the partial cache can reduce the loading time of the partial cache model and the number of times the model is repeatedly loaded. The partial cache model is trained with a large amount of data for text generation, has the logical processing ability of natural language texts, and can execute the processing process of natural language logic according to the input text data, and generate text sequence data that conforms to the logic of human written expression and / or oral expression.
[0059] The prefix input text can be the input text data input into the partial cache model according to the generation requirements of the text sequence. For example, it can be a word, a sentence or a short paragraph of text, specifically such as "Please write an absurd tragedy novel full of comedy elements with the theme of 'Yellow Land', with a word count of more than 100,000 words". As Figure 3 shown, the partial cache model 310 (LLM with Partial KV Cache) can process the prefix input text and generate the candidate draft text corresponding to the prefix input text.
[0060] As Figure 3As shown, the candidate draft text 302 (Candidates) may include text sequence data generated by the partial cache model for the prefix input text according to the model text generation logic. For example, for the prefix input text corresponding to "Please write an absurd tragedy novel full of comedy elements with the theme of 'Yellow Earth' and the number of words exceeding 100,000", the partial cache model can combine the cached data to generate a preliminary draft text (i.e., a combination of tokens) of this novel as a component of the candidate draft text. Among them, the partial cache model can perform "semantic understanding" on the prefix input text, execute text search according to the understood content, and sort the text according to the language logic, and finally can generate the main text composition content of the candidate draft text.
[0061] The full cache model can be a language model (such as Large Language Model, abbreviated as LLM) optimized for partial key-value caching (Full Key-Value Cache) for long text sequence generation tasks. The full cache model can also be trained with a large amount of data for text generation, has the logical processing ability of natural language text, can execute the processing process of natural language logic according to the input text data, and generate text sequence data that conforms to human thinking logic.
[0062] The full cache model has a more complete key-value cache compared to the partial cache model, and can be specifically implemented using the same language model as the partial cache model. In some embodiments, the full cache model and the partial cache model can be used as different text processing modules in the same language model, such as a partial cache module and a full cache module, and the specific implementation is not limited.
[0063] After the candidate draft text corresponding to the output of the partial cache model (LLM with Partial KV Cache) is output, the candidate draft text can be used as the input data of the full cache model 320 (LLM with Full KV Cache) as shown, and then generate the target draft text corresponding to the candidate draft text again for the semantic understanding content of the prefix input text and the understanding content of the candidate draft text 302 (Candidates) as shown. Figure 3 As shown, the candidate draft text 302 (Candidates) of the understanding content, generate the target draft text corresponding to the candidate draft text. Figure 3 As shown, the candidate draft text 302 (Candidates) of the understanding content, generate the target draft text corresponding to the candidate draft text.
[0064] Compare the above candidate draft text and the target draft text, and use the proportion of the text capacity of the same text content between the two in the candidate draft text or the target draft text as the matching degree between the two. At the same time, the same text content after the comparison of the two can be output as the valid output until the target text sequence that can meet the ultra-long text sequence can be output. The so-called ultra-long text sequence (Ultra Long Sequence) refers to the output text capacity of the target text sequence reaching more than 100K.
[0065] Generate the candidate draft text through the partial cache model. Since the partial cache model can only retain the most critical key-value cache content through dynamic key-value pruning, compared with the traditional autoregressive technology solution, it can avoid the process of traversing all the cache content and only match the critical key-value cache content, making the generation speed of the candidate draft text faster. Then, match the candidate draft text with the target draft text generated by the complete cache model. When the text capacity of the target text sequence exceeds 100K, it can still ensure a faster generation speed and also ensure its generation accuracy, so as to achieve a more powerful lossless acceleration ability for ultra-long text sequences.
[0066] With the text sequence generation method of the intelligent agent in the above embodiments of the present invention, through the text processing of the partial cache model and the complete cache model, and using the matching degree of the above candidate draft text and the target draft text to output the target text sequence, for the first time globally, the ultra-long text sequence generation is accelerated to 100K while maintaining the lossless accuracy of the target language model; moreover, compared with the traditional autoregressive technology, it can also adapt to different prefix input lengths, model architectures, and model scales, and specifically can achieve a lossless acceleration ability of more than 3 times; as the generation length increases, while achieving higher acceleration, the generation diversity of ultra-long text sequences is enhanced, which can significantly improve the language interaction ability and intelligent level of the language model.
[0067] To enable those skilled in the art to have a clearer understanding of the above text sequence generation method of the intelligent agent in the embodiments of the present invention, the following Figures 3 to 4C description is further provided.
[0068] Figure 3 Schematically shows a flowchart of an application scenario of the text sequence generation method of the intelligent agent according to an embodiment of the present invention.
[0069] As Figures 2 to 3 shown, according to an embodiment of the present invention, in operation S201 for generating a candidate draft text corresponding to the received prefix input text based on the partial cache model, it includes:
[0070] Output the logical distribution information corresponding to the prefix input text based on the linear layer data of the partial caching model.
[0071] For the partial caching model, the linear layer data can be the quantity data of the linear layer (Linear Layer). The quantity of different linear layers can determine the quantity of the logical distribution information generated during the data processing process (such as the forward propagation process) of the prefix input text.
[0072] The logical distribution information can be the logical distribution probability (Logit) of a certain word, which can be embodied in matrix form and is used to represent the possibility of selecting this word from the preset vocabulary. For example, if the quantity of the linear layer determined by the linear layer data is 3, then the quantity of the logical values (Logit) of the corresponding generated logical distribution information can be 4.
[0073] The prefix input text can be arbitrary words, sentences or paragraphs, serving as the "introduction" for generating the target text sequence as an input prefix, so as to be able to combine the semantic understanding of this "introduction" to conduct reasoning on words or sentences, and finally output the target text sequence.
[0074] Through the logical distribution information, more accurate selection of words in the preset vocabulary can be achieved, so that it can as much as possible conform to the semantic understanding of the user's prefix input text and ensure the accuracy of lossless acceleration.
[0075] Such as Figures 2 to 3 As shown, according to an embodiment of the present invention, in operation S201 for generating a candidate draft text corresponding to the received prefix input text based on the partial caching model, it further includes:
[0076] Obtain candidate word data corresponding to the prefix input text according to the logical distribution information;
[0077] Generate a candidate text path corresponding to the candidate word data through the preset tree attention rule;
[0078] Generate a candidate draft text according to the candidate text path.
[0079] Corresponding to the same preset vocabulary, different logical distribution information corresponds to different probability distributions of words, which can determine the possibility of whether this word conforms to the semantic understanding of the prefix input text and is selected. Specifically, according to the parsing of the prefix input text, semantic-related words can be selected in the preset vocabulary, and the selection possibility of each word can be expressed by its corresponding logical distribution information. The larger the logical value corresponding to the logical distribution information, the greater the selection possibility of the corresponding word, and some words with greater possibility can be selected as candidate word data.
[0080] The preset tree attention rule can be a setting rule for modeling based on a tree structure, which is applicable to processing data with hierarchical features. By constructing and operating on a tree structure, it captures the hierarchical dependency logical relationships in text data. Specifically, the top k candidate words (such as candidate 4-grams) can be retrieved according to the logical distribution information (Logits) to form candidate word data, where k is a positive integer approximately equal to 1, such as Figure 3 As shown, if the candidate word data includes "is", "the", "a", "an", "uncle", "cousin", "good", "of", "to", etc., then based on Figure 3 As shown in the preset tree attention rule 301 (Tree-base Attention), a tree-shaped logical structure based on candidate words can be constructed through logical combination.
[0081] According to the above preset tree attention rule, different logical combinations of words in the tree-shaped logical structure can form different candidate text paths. That is to say, a candidate text path can be understood as the word combination logic formed by a certain word combination of candidate words according to the tree-shaped logic. For example, the word combination logic of "is the uncle of" can form a candidate text path.
[0082] Based on these candidate text paths, the main text content of the candidate draft text 302 (i.e., Candidates, which can be understood as candidate tokens) as shown in Figure 3 can be generated. As shown in Figure 3 , the candidate draft text 302 can be formed in the form of k rows × 4 columns as follows:
[0083] is the uncle of
[0084] is the uncle\n
[0085] ……
[0086] is an good to
[0087] is the father of
[0088] is the mother of
[0089] ……
[0090] Among them, each row can correspond to a candidate text path. Each candidate text path can define the combination logic of the selected candidate words and generate the corresponding text content, and finally jointly form the main text content of the candidate draft text.
[0091] Therefore, the candidate word data selected using the logical distribution information can be logically combined through a preset tree attention rule to form a candidate draft text that conforms to semantic understanding, so as to be able to construct a multi-draft parallel self-drafting architecture (i.e., Multi-token Parallel Self-Drafting) of a candidate draft text as shown in Figure 3 , which speeds up the generation speed of the candidate draft text, avoids the situation of a large amount of repeated text content or even meaningless messy symbols, and ensures the generation accuracy.
[0092] As shown in Figures 2 to 3 , according to an embodiment of the present invention, in obtaining candidate word data corresponding to the prefix input text according to the logical distribution information, it includes:
[0093] Obtaining preselected word data corresponding to the prefix input text according to the logical distribution information;
[0094] Obtaining candidate word data corresponding to the preselected word data through a preset context penalty rule.
[0095] According to the semantic understanding of the prefix input text, words can be selected in a preset word list, specifically selected according to the logical value (Logit) of the logical distribution information corresponding to each word. Specifically, multiple words with the largest sorted numerical values of the logical values can be selected as the preselected word data.
[0096] As shown in Figure 3 , the preset context penalty rule 304 (Contextual Penalty) can be a setting rule for screening words in the preselected word data. Through the preset context penalty rule, the repeated words in the above preselected word data can be screened. Specifically, the penalty value can be set according to the similarity of two words. When the similarity of two words is greater, the penalty value is greater, and then one of the words can be selected as the final candidate word.
[0097] Therefore, the candidate word data can be part of the text data in the preselected word data after being screened by the preset context penalty rule.
[0098] Thereby, the number of repeated texts in the candidate word data can be greatly reduced, the occurrence frequency of repeated texts in the candidate draft text can be greatly reduced, and the redundant calculation can also be further reduced, avoiding the situation of repeated text content or even meaningless messy symbols, thus significantly improving the lossless acceleration ability.
[0099] As shown in Figures 2 to 3 , according to an embodiment of the present invention, in generating a candidate draft text according to the candidate text path, it includes:
[0100] Querying the draft text data in the preset text item sequence table that matches the candidate text path;
[0101] Generate candidate draft text based on the candidate text path and draft text data.
[0102] The preset text item sequence table can be a sequence storage table for text items (n-grams), that is, an n-gram table. The n-gram table can store the target draft text finally output by the complete cache model and the text item content after comparing it with the candidate draft text. As Figure 3 shown, the preset text item sequence table can store and express text items (n-grams) and their corresponding frequencies in the form of row numbers (row id). As Figure 3 shown, for the text item with row number 97475, it is "is the father of", and the corresponding frequency is 5.
[0103] According to the logical content of the candidate text path defined by the above preset tree attention rules (such as the number k of selected words), during the process of generating the above candidate draft text, the first k (such as Figure 3 shown as Top-k) text item contents in the preset text item sequence table can be selected and combined with the text combination content defined by the candidate text path to form the above candidate draft text.
[0104] It can be seen that the text content of the candidate draft text can include both text items generated by the candidate text path defined by the preset tree attention rules and the content that re-uses the existing text items in the preset text item list.
[0105] Therefore, the re-use of the existing draft text item content can be realized (such as Figure 3 shown as text item re-use Token Reutilization). By this means, the rapid generation of the candidate draft text can also be realized, while ensuring the generation accuracy of the final target text sequence, avoiding the situation of text content duplication or even meaningless random symbols caused by the deviation of the model architecture itself, and further improving the lossless acceleration ability of text sequence generation.
[0106] As Figures 2 to 3 shown, according to an embodiment of the present invention, in operation S203, when outputting the target text sequence according to the matching degree between the candidate draft text and the target draft text, it includes:
[0107] Obtain the valid text sequence according to the matching degree between the candidate draft text and the target draft text;
[0108] When the valid text capacity of the valid text sequence meets the preset target text capacity, output the target text sequence.
[0109] As shown Figure 3 in FIG., the candidate draft text 302 is compared with the target draft text 303 to obtain the matching degree of the text contents of the two, that is, the proportion of the text capacity of the same text content (which can be relative to the candidate draft text and / or the target draft text). Based on the size of this matching degree, the same text content of the two can be output as the valid text sequence.
[0110] Therefore, the valid text sequence can be the same text sequence in the candidate draft text and the target draft text. In addition, the valid text sequence finally output by the complete cache model 320 can be the one with the largest text capacity among the output text sequences (as shown Figure 3 in FIG. as the longest and valid, that is, Longest Valid).
[0111] The preset target text capacity can be the capacity condition for outputting the valid text sequence as the target text sequence. Generally, it is the starting value of the text capacity of the ultra-long text sequence. For example, a text capacity of 100K is used as the preset target text capacity. For example, only when the generated valid text sequence satisfies its valid text capacity of 100K can it be output as the final target text sequence. Otherwise, it is necessary to repeat the generation process of the candidate draft text of the partial cache model and the target draft text of the complete cache model, and repeat the generation of the valid text sequence until the generated valid text sequence satisfies the preset target text capacity. At this time, the longest and valid valid text sequence is output as the target text sequence.
[0112] Therefore, through the matching verification of the candidate draft text and the target draft text, a valid and longest text capacity target text sequence can be generated, thus constructing a self-output scheme of the longest and valid (Longest&Valid) target draft text that can achieve parallel verification (Parallel Verification). Compared with the traditional autoregressive technical scheme, it can significantly improve the generation speed of the target text sequence, effectively avoid the appearance of redundant data, and realize the further improvement of the lossless acceleration ability.
[0113] As shown Figures 2 to 3 in FIG., according to an embodiment of the present invention, after obtaining the valid text sequence according to the matching degree of the candidate draft text and the target draft text, it further includes:
[0114] Updating the preset text item sequence list, the partial cache model, and the complete cache model according to the valid text sequence.
[0115] During the process of generating the target text sequence, each time the valid text sequence that completely caches the model output needs to be used as updated data to update the text items in the preset text item sequence list. The text items in the valid text sequence that are different from the text items already stored in the above-mentioned preset text item sequence list are updated and stored in the preset text item sequence list to ensure that the reuse of the draft text (i.e., Figure 3 the text item reuse shown in Token Reutilization) can be achieved during the generation process of the next candidate draft text (Candidates), thereby effectively avoiding the appearance of redundant text and improving the lossless acceleration ability of text generation.
[0116] At the same time, the relevant text information of the above-mentioned valid text sequence will also be used to update the data of the partial cache model and the complete cache model. Specifically, the key-value data can be updated, so that during the generation process of the next candidate draft text, the appearance of redundant text or incorrect words can be further avoided, thereby further improving the lossless acceleration ability of text generation.
[0117] In summary, the text sequence generation method (Token Swift) of the above-mentioned agent in the embodiment of the present invention can generate a draft text sequence with a self-drafting function, and then transfer it to a language model with a complete cache for verification using a tree attention mechanism. This process can ensure that the finally generated target text sequence can be consistent with the prediction of the model, thereby effectively achieving lossless acceleration.
[0118] Figure 4A Schematically shows a comparison table of the acceleration performance data of the text sequence generation method of the agent according to the embodiment of the present invention on different architecture models; Figure 4B Schematically shows a line graph of the changes in the corresponding acceleration and acceptance rate with the increase in text capacity of the text sequence generation method of the agent according to the embodiment of the present invention; Figure 4C Schematically shows a comparison table of the acceleration data of the text sequence generation method of the agent according to the embodiment of the present invention on different scale models.
[0119] Therefore, through the text sequence generation method of the above-mentioned agent in the embodiments of the present invention, for the first time globally, the generation of ultra-long text sequences is accelerated to 100K, while maintaining the lossless accuracy of the target language model. Moreover, it can always achieve more than 3 times acceleration for different prefix lengths, model architectures, and model scales, greatly reducing the required time of the traditional autoregressive process (nearly 5 hours) to only 90 minutes. In addition, as the length of the text sequence generation increases, it also has a higher acceleration ability than the traditional autoregressive technology, while enhancing the diversity of ultra-long sequence generation, directly breaking through the bottleneck of lossless acceleration of ultra-long text sequences in the prior art, significantly improving the natural language processing ability of the language model in terms of low latency and high throughput, significantly enhancing the language interaction ability of the agent, and achieving extremely successful scientific research value and commercial application value.
[0120] Specifically, taking the text sequence generation method of the agent in the embodiments of the present invention (TokenSwift solution) as an example, the verification content of the relevant technical effects is provided as follows:
[0121] As Figure 4A shown in the table, at all generation lengths, the TokenSwift solution provided in the embodiments of the present invention shows better acceleration performance on models with different architectures (MHA, GQA) than all baselines (Medusa, TriForce). At the same time, it also shows excellent robustness and has little impact when tested with different prefix lengths.
[0122] As Figure 4B shown in the change line chart, as the generation time increases, the text sequence generation speed (speed up) of the TokenSwift solution improves more and more significantly. Among them, there are two key factors driving this trend: First, as the number of tokens increases, the loading time of the autoregressive KV cache will be longer, while the TokenSwift solution alleviates this problem through dynamic KV pruning. Second, the acceptance rate will increase as the number of tokens increases, which is mainly due to the higher n-gram acceptance rate (Acceptance). As Figure 4B shown, as the n-gram pool composed of the generated tokens becomes larger and larger, the candidate n-grams become more diverse and accurate.
[0123] The impact of frequent model reloading varies with the model scale. The larger the model scale, the more parameters, and the more time is required. As Figure 4CAs shown in the table, the TokenSwift solution performs excellently on model architectures of different scales, and the larger the model scale, the more obvious the acceleration advantage. Specifically, when generating a text sequence token of 100K, the TokenSwift solution can save up to 5.54 hours for the 14B model.
[0124] Therefore, through the text sequence generation method of the above-mentioned agent in the embodiments of the present invention, for the first time, lossless acceleration of ultra-long text sequences is achieved at a length of 100K. Moreover, based on this method, acceleration of more than 3 times can always be achieved for different prefix lengths, model architectures, and model scales. At the same time, the diversity of the generated content can be significantly enhanced during the lossless acceleration process.
[0125] Based on the above-mentioned text sequence generation method of the agent, the present invention also provides a text sequence generation device for the agent. The following will be combined with Figure 5 to describe this device in detail.
[0126] Figure 5 The structural block diagram of the text sequence generation device for the agent according to the embodiments of the present invention is schematically shown.
[0127] As Figure 5 shown, the text sequence generation device 500 for the agent in this embodiment includes a candidate text generation module 510, a target text generation module 520, and a target sequence output module 530.
[0128] The candidate text generation module 510 is used to generate a candidate draft text corresponding to the received prefix input text based on a partial cache model. In one embodiment, the candidate text generation module 510 can be used to perform the operation S201 described above, which will not be elaborated here.
[0129] The target text generation module 520 is used to generate a target draft text corresponding to the candidate draft text through a complete cache model. In one embodiment, the target text generation module 520 can be used to perform the operation S202 described above, which will not be elaborated here.
[0130] The target sequence output module 530 is used to output a target text sequence according to the matching degree between the candidate draft text and the target draft text. In one embodiment, the target sequence output module 530 can be used to perform the operation S203 described above, which will not be elaborated here.
[0131] According to an embodiment of the present invention, any multiple of the candidate text generation module 510, the target text generation module 520, and the target sequence output module 530 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the candidate text generation module 510, the target text generation module 520, and the target sequence output module 530 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system in a package, an application specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the candidate text generation module 510, the target text generation module 520, and the target sequence output module 530 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.
[0132] Figure 6 FIG. schematically shows a block diagram of an electronic device suitable for implementing a text sequence generation method of an agent according to an embodiment of the present invention.
[0133] The above-mentioned electronic device provided by an embodiment of the present invention includes one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned text sequence generation method of the agent.
[0134] As Figure 6 shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 601 may also include on-board memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0135] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 may also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.
[0136] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the I / O interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read therefrom is installed into the storage portion 608 as needed.
[0137] The present invention also provides a computer-readable storage medium having executable instructions stored thereon, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned text sequence generation method of the agent.
[0138] Among them, the computer-readable storage medium may be included in the device / device / system described in the above embodiments; or it may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0139] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 602 and / or RAM 603 described above and / or one or more memories other than the ROM 602 and RAM 603.
[0140] An embodiment of the present invention further includes a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned text sequence generation method of the agent.
[0141] Wherein, the computer program contains program codes for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program codes are used to enable the computer system to implement the method provided by the embodiment of the present invention.
[0142] When the computer program is executed by the processor 601, it executes the above-mentioned functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-mentioned systems, apparatuses, modules, units, etc. described above can be implemented by computer program modules.
[0143] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 609, and / or be installed from the removable medium 611. The program codes included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0144] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or be installed from the removable medium 611. When the computer program is executed by the processor 601, it executes the above-mentioned functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-mentioned systems, devices, apparatuses, modules, units, etc. described above can be implemented by computer program modules.
[0145] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0147] In addition, all actions of obtaining information, signals, or data in the present invention are carried out on the premise of complying with the corresponding data protection laws, regulations, and policies of the country where it is located and obtaining authorization from the owner of the corresponding device.
[0148] Those skilled in the art can understand that the features described in the various embodiments and / or claims of the present invention can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments and / or claims of the present invention can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0149] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A text sequence generation method for an intelligent agent, characterized in that: include: Generate candidate draft text corresponding to the received prefix input text based on the partial cache model; Generate a target draft text corresponding to the candidate draft text through a complete cache model; as well as Outputting a target text sequence according to the matching degree between the candidate draft text and the target draft text.
2. The method according to claim 1, characterized in that In the generating of the candidate draft text corresponding to the received prefix input text based on the partial cache model, the method includes: According to the linear layer data of the partial cache model, the logical distribution information corresponding to the prefix input text is output.
3. The method according to claim 2, characterized in that In the generating of the candidate draft text corresponding to the received prefix input text based on the partial cache model, the method further includes: Acquire candidate word data corresponding to the prefix input text according to the logical distribution information; Generate a candidate text path corresponding to the candidate word data by using a preset tree attention rule; The candidate draft text is generated according to the candidate text path.
4. The method according to claim 3, characterized in that In the step of obtaining candidate word data corresponding to the prefix input text according to the logical distribution information, the following steps are included: Acquire pre-selected word data corresponding to the prefix input text according to the logical distribution information; The candidate word data corresponding to the pre-selected word data is obtained through a preset context penalty rule.
5. The method according to claim 3, characterized in that: In generating the candidate draft text according to the candidate text path, the method includes: Querying the preset text item sequence table for draft text data that matches the candidate text path; The candidate draft text is generated according to the candidate text path and the draft text data.
6. The method according to claim 1, characterized in that In the step of outputting the target text sequence according to the matching degree between the candidate draft text and the target draft text, the step includes: Acquire a valid text sequence according to the matching degree between the candidate draft text and the target draft text; When the effective text capacity of the effective text sequence meets the preset target text capacity, the target text sequence is output.
7. The method according to claim 6, characterized in that After obtaining a valid text sequence according to the matching degree between the candidate draft text and the target draft text, the method further includes: The preset text item sequence table, the partial cache model and the complete cache model are updated according to the valid text sequence.
8. A text sequence generation device for an intelligent agent, characterized in that: include: A candidate text generation module, used for generating candidate draft text corresponding to the received prefix input text based on the partial cache model; A target text generation module, used for generating a target draft text corresponding to the candidate draft text through a complete cache model; as well as The target sequence output module is used to output a target text sequence according to the matching degree between the candidate draft text and the target draft text.
9. An electronic device, comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Content reasoning and generating method, electronic equipment, storage medium and program product
CN122433922A
Content reasoning and generation methods, electronic devices, storage media, and application products.
CN122433922B
Detecting artificial intelligence generated text
US12596869B2
Detecting artificial intelligence generated text
US20240378380A1