Text reasoning method, question answering method, translation method and electronic equipment

By analyzing the importance indicators and pruning of the Token vectors input by long text, the problem of slow inference speed in long text inference is solved, and a faster text inference process and a better user experience is achieved.

CN120069058APending Publication Date: 2025-05-30ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411997178.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, large language models in long text input scenarios are slower in reasoning, especially when generating the first token, it takes a long time to affect the user experience.

Method used

By inputting and vectorizing the text to be processed, pruning and decoding are performed based on the importance indicator of the token vector, the hidden layer vector of some token vectors is dynamically selected for decoding, reducing the computational amount of the feedforward network module.

Benefits of technology

It effectively reduces the time required for a large language model to generate the first token, improves text inference speed, improves user experience, and improves inference efficiency while ensuring model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069058A_ABST
    Figure CN120069058A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text reasoning method, a question answering method, a translation method, electronic equipment, a storage medium and a computer program product. The method comprises the following steps: performing input processing on a to-be-processed text to obtain a first sequence formed by Token codes corresponding to the to-be-processed text; performing vectorization processing on the Token codes in the first sequence to obtain a second sequence formed by Token vectors corresponding to the to-be-processed text; and pruning and decoding the Token vectors on the basis of importance indexes calculated in the decoding process of each Token vector in the second sequence, and obtaining reasoning output corresponding to the to-be-processed text. According to the text reasoning method, Token-level pruning is carried out on the feed-forward network which consumes the most time, the time required for generating the first Token by the large language model is effectively shortened on the basis of ensuring less model precision loss, and better large language model use experience is brought to a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to a text inference method, a question answering method, a translation method, an electronic device, a storage medium, and a computer program product. Background Art

[0002] With the enhancement of the long text understanding ability of large language models, their applications in long text input are increasing, such as question answering for single or multiple document contents, summarization of long articles, etc. However, the inference of large language models in the long text input scenario often requires a large amount of computing time. For a Qwen2 model (Thousand Questions model) with 7 billion parameters, if a user inputs only one document with a length of 18,000 Tokens for a single inference request and asks a question, the large language model needs 3,000 milliseconds to generate a complete reply. In the prior art, to improve the processing speed of long texts, the text processing process is usually optimized in three aspects: reducing video memory occupancy, accelerating the decoding process, and enhancing the generalization ability of the model to long texts.

[0003] However, the inference speed of the text inference methods in the prior art still needs to be improved. Summary of the Invention

[0004] An embodiment of the present application provides a text inference method, which can improve the inference speed for the input long text.

[0005] Correspondingly, an embodiment of the present application also provides a question answering method, a translation method, an electronic device, a storage medium, and a computer program product to ensure the implementation and application of the above text inference method.

[0006] To solve the above problems, an embodiment of the present application discloses a text inference method, and the method includes:

[0007] Perform input processing on the text to be processed to obtain a first sequence composed of Token encodings corresponding to the text to be processed;

[0008] Perform vectorization processing on the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be processed;

[0009] Based on the importance indicators calculated during the decoding process for each Token vector in the second sequence, perform pruning and decoding processing on the Token vectors to obtain the inference output corresponding to the text to be processed.

[0010] An embodiment of the present application discloses a question answering method, and the method includes:

[0011] Perform input processing on the problem text to obtain a first sequence composed of Token encodings corresponding to the problem text;

[0012] Perform vectorization processing on the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the problem text;

[0013] Based on the importance indicators calculated during the decoding process for each Token vector in the second sequence, perform pruning and decoding processing on the Token vectors to obtain the answer text corresponding to the problem text.

[0014] An embodiment of the present application discloses a translation method, and the method includes:

[0015] Perform input processing on the text to be translated to obtain a first sequence composed of Token encodings corresponding to the text to be translated;

[0016] Perform vectorization processing on the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be translated;

[0017] Based on the importance indicators calculated during the decoding process for each Token vector in the second sequence, perform pruning and decoding processing on the Token vectors to obtain the translation result corresponding to the text to be translated.

[0018] An embodiment of the present application further discloses an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0019] The memory stores computer-executable instructions;

[0020] The processor executes the computer-executable instructions stored in the memory to implement the method as described in the embodiment of the present application.

[0021] An embodiment of the present application further discloses a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method as described in the embodiment of the present application.

[0022] An embodiment of the present application further discloses a computer program product, including a computer program / computer-executable instructions, characterized in that when the computer program / computer-executable instructions are executed by a processor in an electronic device, they implement the method as described in the embodiment of the present application.

[0023] Compared with the prior art, the embodiments of the present application include the following advantages:

[0024] By performing input processing on the text to be processed, a first sequence composed of Token encodings corresponding to the text to be processed is obtained; the Token encodings in the first sequence are vectorized to obtain a second sequence composed of Token vectors corresponding to the text to be processed; based on the attention scores corresponding to each Token vector in the second sequence, some of the hidden layer vectors corresponding to the Token vectors are dynamically selected for decoding processing based on the attention mechanism, and the Tokens input to the feed-forward network module are pruned, effectively reducing the time required to generate the first Token when performing text inference by calling a large language model while ensuring less loss of model accuracy, and bringing a better experience of using the large language model to users. Brief Description of the Drawings

[0025] Figure 1 is a flowchart of the steps of the text inference method disclosed in an embodiment of the present application;

[0026] Figure 2 is a schematic diagram of the time-consuming proportion of each module in the decoding layer of a large language model in the prior art;

[0027] Figure 3 is a schematic diagram of the decoding process in the inference method disclosed in an embodiment of the present application;

[0028] Figure 4 is a flowchart of the steps of the question answering method disclosed in an embodiment of the present application;

[0029] Figure 5 is a flowchart of the steps of the translation method disclosed in an embodiment of the present application;

[0030] Figure 6 is a schematic diagram of the structure of an exemplary device provided by an embodiment of the present application. Detailed Description of the Embodiments

[0031] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] The text inference method disclosed in an embodiment of the present application is applied to the process of performing text inference by calling a large language model. The large language model performs encoding processing based on the attention mechanism. The large language model includes multiple stacked decoding layers, and each decoding layer further includes: a residual connection and a normalization layer, a self-attention module, and a feed-forward network module.

[0033] Among them, the encoding layer is used to perform text processing and vectorization on the text to be processed; the self-attention module is used to calculate the context relevance of each position in the second sequence input to the decoding layer. The self-attention mechanism enables the large language model to establish global dependencies in the second sequence, thereby better capturing long-range dependencies in the second sequence.

[0034] The feed-forward network module is used to perform non-linear transformation and mapping on the representation of each position.

[0035] The residual connections and layer normalization are added between the decoder layers to stabilize the model training process and accelerate model convergence.

[0036] For the specific implementation manners of the residual connections, layer normalization, and the feed-forward network module, reference can be made to the prior art, and details are not described herein in the embodiments of the present application.

[0037] Refer to Figure 1 , the text inference method disclosed in the embodiments of the present application is applied to the process of invoking a large language model for text inference, and includes: steps 102 to 106.

[0038] Step 102, perform input processing on the text to be processed to obtain a first sequence composed of Token encodings corresponding to the text to be processed.

[0039] The text inference method disclosed in the embodiments of the present application can be applied to the inference process executed by invoking a large language model.

[0040] In the embodiments of the present application, "Token" refers to the smallest unit of text, which may be a word, character, or symbol.

[0041] In the prior art, the inference process of a large language model includes four steps: input processing, vector embedding, decoder operation, and output generation. In the process of applying a large language model to perform inference on input text, for a piece of text input by a user to the large language model, the large language model first preprocesses the input text to perform word segmentation and encoding on the input text. For example, a tokenizer is used to decompose the input text of the large language model into a series of words or sub-words (i.e., Tokens), and these words or sub-words are converted into a sequence of digital forms, and these numbers are usually the index numbers of the words in the model dictionary. The encoding process converts the word-segmented text sequence into a digital sequence that can be processed by the model, and the digital sequence is the first sequence composed of Token encodings corresponding to the text to be processed.

[0042] For the specific implementation of input processing the text to be processed to obtain the first sequence composed of Token encodings corresponding to the text to be processed, refer to the prior art and will not be elaborated in the embodiments of the present application.

[0043] Step 104: Vectorize the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be processed.

[0044] After tokenization and encoding, the large language model converts the first sequence into a sequence of high-dimensional vectors through an embedding layer, which is denoted as the "second sequence" in the embodiments of the present application. This process converts words or subwords into points in a mathematical space, such that words with similar meanings or contexts are close to each other in the space. This vector representation enables the model to understand the semantic and contextual relationships of words.

[0045] For the specific implementation of vectorizing the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be processed, refer to the prior art and will not be elaborated in the embodiments of the present application.

[0046] Step 106: Based on the importance indicators calculated during the decoding process for each Token vector in the second sequence, perform pruning and decoding processing on the Token vectors to obtain the inference output corresponding to the text to be processed.

[0047] Among them, the importance indicator can be the attention score corresponding to the Token vector during the decoding process.

[0048] Next, in the decoder operation step, the large language model performs decoding operations based on the attention mechanism on the vectors in the second sequence through the decoder to obtain a sequence of numbers.

[0049] In the prior art, in each decoding layer of the large language model, first, the self-attention module encodes each Token vector in the second sequence, and then, the encoding results of all Token vectors in the second sequence are further input into the feed-forward network module for decoding mapping. During the inference process of the large language model, there are usually two stages: the first inference stage (Prefilling stage), in which the model parallelly performs inference calculations on all input Tokens in this stage, stores the necessary intermediate results in the cache, and finally generates the first Token; the second inference stage (Decoding stage), in which the model cyclically uses the latest generated Token as the input, combines the cache content for inference calculations, and generates the next Token until a certain stop condition is reached.

[0050] By analyzing the inference process of large language models commonly used in the prior art, it is found that the feed-forward network accounts for the largest proportion of the time consumption in the first inference stage. For example, Figure 2 taking the analysis results of the time consumption proportion of each module in the decoding stage of two commonly used large language models as an example, the time consumption proportion of the feed-forward network in the first inference stage exceeds 60%.

[0051] For a large language model with 7 billion parameters (such as the Qwen2 model), even in the ideal case where the model only needs to process one request at a time, when the user inputs a document with a length of 18,000 Tokens and asks a question, the large language model needs to calculate for 3,000 milliseconds to generate a complete response, and 80% of the time is used to calculate the first Token, which means that the user needs to wait at least 2,400 milliseconds to get the response of the first Token. For large language models with higher precision and more parameters, the user needs to wait longer to get the response of the first Token. This waiting time causes great damage to the user experience.

[0052] To solve the above problems, the inference method disclosed in the embodiments of the present application counts the time consumption of each main module in the first inference stage (i.e., the Prefilling stage) of the large language model, and performs pruning at the Token level on the feed-forward network with the most time consumption, effectively reducing the time required for the large language model to generate the first Token on the basis of ensuring less loss of model accuracy, and bringing a better user experience of using the large language model to the user.

[0053] In some optional embodiments, based on the importance indicators calculated during the decoding process of each Token vector in the second sequence, pruning and decoding processing are performed on the Token vector to obtain the inference output corresponding to the text to be processed, including: based on the importance indicators calculated during the decoding process of each second Token vector and each first Token vector in the second sequence, selecting the hidden layer vector corresponding to the first Token vector and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing to obtain the inference output corresponding to the text to be processed, where the first Token vector is: the Token vector within a specified observation window, and the second Token vector is: the Token vector outside the specified observation window.

[0054] Wherein, the specified observation window is: the hidden layer vectors corresponding to the last N Tokens in the second sequence, where N is a positive integer greater than or equal to 32.

[0055] The length of the specified observation window requires a lot of experiments. If the specified observation window is too long, it will increase the additional attention calculation and reduce the inference performance of the model. In addition, if the specified observation window is too long, it will also introduce more irrelevant noises. While if the specified observation window is too short, it may not capture enough important information. Through experiments, the value of the specified observation window can include: the hidden layer vectors between 32 and 64 in the ending part of the second sequence.

[0056] On the other hand, through testing, the hidden layer vectors corresponding to the first few Tokens (generally 4) in the second sequence are retained and not discarded in any decoding layer, and more accurate inference results can be obtained.

[0057] In some optional embodiments, the importance index includes: attention scores. Based on the importance index calculated during the decoding process between each second Token vector and each first Token vector in the second sequence, select the hidden layer vector corresponding to the first Token vector and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing to obtain the inference output corresponding to the text to be processed, including: based on the attention scores calculated respectively between each second Token vector and each first Token vector in the second sequence in a preset multi-layer decoding layer, and the attention score threshold corresponding to the corresponding decoding layer, select the hidden layer vector corresponding to the first Token vector and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing to obtain the inference output corresponding to the text to be processed.

[0058] During the inference process of calling the large language model, the attention score is an intermediate product of the self-attention module of the large language model, which records the importance of any Token to another Token and can be used as an importance index of the Token. In the prior art, the attention score is not the output of the self-attention module. In the embodiments of the present application, by improving the process of calling the large language model for inference, while learning the global dependency relationship between the hidden layer vectors corresponding to each Token vector in the second sequence through the self-attention module, a simplified attention network is added to obtain the attention matrix, so as to obtain the attention scores corresponding to each Token vector in the second sequence. The simplified attention network obtains the query vector matrix Q and the vector matrix K to be queried of the self-attention module, and then calculates the attention score matrix based on the query vector matrix Q and the vector matrix K to be queried.

[0059] In some optional embodiments, the calculation formula of the attention score matrix can be expressed as follows:

[0060]

[0061] Among them, M represents the matrix composed of the attention scores of all heads, h represents the head identifier, H represents the number of heads in the self-attention module, L represents the length of the query vector, represents the attention score matrix of all tokens in a certain head with other tokens, Q represents the query vector matrix, K represents the vector matrix to be queried, and d represents the matrix size of the attention score matrix.

[0062] For the detailed introduction and principle of the calculation formula 1 of the attention score matrix, please refer to the prior art and will not be elaborated in the embodiments of the present application.

[0063] In the embodiments of the present application, after using the observation window, the calculation formula of the attention score matrix can be expressed as follows:

[0064]

[0065] In the above formula 2, h represents the head identifier, H represents the number of heads in the self-attention module, L represents the length of the query vector, and M ′ h represents the attention score matrix of head h when using the observation window, N represents the size of the specified observation window, represents the attention score matrix of all tokens in the specified observation window in a certain head with other tokens. In the attention score matrix, the attention score can be calculated by the formula calculate.

[0066] Furthermore, the large language model scores the "importance" of all tokens according to the self-attention scores between the hidden layer vectors corresponding to the tokens. For example, after calculating the self-attention scores of the hidden layer vectors of all tokens relative to the tokens within the specified observation window and summing them in the N dimension, the attention scores of the hidden layer vectors of each token are obtained. Then, the tokens can be sorted in descending order of the attention scores, and the hidden layer vectors of the important tokens (i.e., the tokens with high attention scores) are selected and input into the feed-forward network module for calculation, while the hidden layer vectors corresponding to the unimportant tokens skip the calculation of the feed-forward network module part, thereby greatly reducing the calculation amount of the feed-forward network module.

[0067] In some alternative embodiments, the decoding layer includes a self-attention module and a feed-forward network module. Based on the attention scores respectively calculated for each second Token vector and each first Token vector in the second sequence in a preset multi-layer decoding layer, and the attention score threshold corresponding to the corresponding decoding layer, the hidden layer vectors corresponding to the first Token vector and some of the hidden layer vectors corresponding to the second Token vectors are selected for decoding processing to obtain the inference output corresponding to the text to be processed, including: successively taking each decoding layer as the current decoding layer, and performing the following decoding operations through the current decoding layer to obtain the decoding output corresponding to the current decoding layer; obtaining the inference output corresponding to the text to be processed based at least on the decoding output corresponding to the last decoding layer; wherein the decoding operations include: obtaining the hidden layer vector output by the self-attention module in the current decoding layer; dynamically selecting some of the hidden layer vectors as the retained hidden layer vectors based on the attention scores calculated for each second Token vector and each first Token vector in the current decoding layer and the attention score threshold corresponding to the current decoding layer; inputting the retained hidden layer vectors and the hidden layer vectors corresponding to the first Token vector into the feed-forward network module in the current decoding layer for processing to obtain a processing result; and integrating the processing result and the hidden layer vectors other than the retained hidden layer vectors to obtain the decoding output corresponding to the current decoding layer.

[0068] In the embodiments of the present application, decoding processing is performed by successively setting the expulsion Tokens layer by layer according to the importance of the decoding layer. That is, the attention score threshold corresponding to the decoding layer is positively correlated with the importance of the decoding layer.

[0069] Optionally, the attention score threshold corresponding to the decoding layer is preset according to the importance of the decoding layer.

[0070] For example, the shallow decoding layers of the large language model need to retain more Tokens. In some alternative embodiments, the attention score thresholds of the first ten decoding layers of the large language model can be set to 100%, that is, during the decoding process, the complete Token vector sequence is decoded. For other decoding layers of the large language model, the attention score thresholds can be set according to the calculated importance.

[0071] Correspondingly, the successive taking of each decoding layer as the current decoding layer includes: successively taking each decoding layer after the L-th layer as the current decoding layer, where L is an integer greater than or equal to 10.

[0072] Among them, the importance of the decoding layer can be set according to the test results or determined according to the cosine similarity between the input and output of the feed-forward network of the current layer. The greater the cosine similarity between the input and output of the feed-forward network of the current layer, the higher the importance. Optionally, the attention score threshold of other decoding layers can be set to a value between 80% and 95%.

[0073] Since calculating the similarity is rather troublesome, in some other alternative embodiments, the attention score thresholds of other layers can all be set to a fixed value between 80% and 95%, for example, 90%. In this way, the vast majority of the model accuracy can be ensured.

[0074] In some alternative embodiments, in order to balance the inference accuracy and inference speed of the model, based on the inference results of the test samples, the importance of each inference layer can be calculated in advance according to the cosine similarity between the input and output of the feed-forward network of each layer, so as to obtain the attention scores and thresholds of each inference layer in advance.

[0075] In some alternative embodiments, based on the attention scores calculated from each second Token vector and each first Token vector in the second sequence in the current decoding layer, and the attention score threshold corresponding to the current decoding layer, dynamically select some of the hidden layer vectors as the reserved hidden layer vectors, including: obtaining the attention scores corresponding to each second Token vector as the attention scores corresponding to the corresponding hidden layer vectors; selecting K hidden layer vectors in descending order of the corresponding attention scores as the reserved hidden layer vectors, where the sum of the attention scores corresponding to the K hidden layer vectors is greater than or equal to the attention score threshold corresponding to the current decoding layer, and K is an integer greater than 1. Among them, the attention score threshold is less than 1.

[0076] Among them, the input hidden layer vector of the current layer decoding layer is the hidden layer vector updated by the previous decoding layer according to the decoding output.

[0077] The following Figure 3 illustrates the decoding process of the current layer decoding layer with reference to the decoding process schematic diagram shown.

[0078] First, perform decoding processing based on the attention mechanism on the hidden layer vectors input to the current decoding layer through the self-attention module of the current decoding layer to obtain the hidden layer vectors output by the self-attention module.

[0079] Then, using the method for calculating attention scores described above, calculate the attention score matrix between each Token in the second sequence based on the Q and K matrices in the intermediate parameters generated during the decoding process by the self-attention module in the current decoding layer, and intercept the matrix row data corresponding to the specified observation window (such as the last N rows) for summation, so as to obtain the attention scores of each Token with respect to each Token in the specified observation window. The attention score of each Token is the attention score of the hidden layer vector corresponding to the Token output by the self-attention module in the current, for example, represented by the matrix S in Figure 3 as shown.

[0080] Then, sort the attention scores of each hidden layer vector and retain the top K hidden layer vectors with high attention scores, such that the proportion of the sum of the attention scores of the retained hidden layer vectors to the sum of the scores of all the originally retained Tokens is greater than the attention score threshold set for the current decoding layer. Taking the Token subscripts in the second sequence as 0 to 7 as an example, assuming K = 3, retain the hidden layer vectors corresponding to the 3 Tokens with the highest attention scores, such as the hidden layer vectors corresponding to the Tokens with subscripts 0, 1, and 5, as the retained hidden layer vectors.

[0081] After that, input the retained hidden layer vectors and the hidden layer vectors corresponding to the specified observation window into the feed-forward network module of the current decoding layer for calculation to obtain the decoding output corresponding to the current decoding layer. Among them, the decoding output is the updated feature information of the current decoding layer. Assuming that the Token vector subscript in the specified observation window is 7, the hidden layer vectors corresponding to the Tokens with subscripts 0, 1, and 5 (i.e., the retained hidden layer vectors) output by the self-attention module in the current decoding layer can be extracted and concatenated with the Token vector with subscript 7 in the specified observation window to obtain a feature sequence, which is input into the feed-forward network module for calculation to obtain the processing results corresponding to the Tokens with subscripts 0, 1, 5, and 7.

[0082] Next, update the updated feature information to the feature information (i.e., the hidden layer vectors) output by the self-attention module in the current decoding layer according to the corresponding Token subscripts through residual connection to retain the information of the pruned Tokens. For example, use the processing results corresponding to the Tokens with subscripts 0, 1, 5, and 7 to update the hidden layer vectors corresponding to the Tokens with subscripts 0, 1, 5, and 7 output by the self-attention module of the current decoding layer, so as to obtain the decoding output of the current decoding layer. It can be seen from the foregoing decoding process that the information of the pruned Tokens with subscripts 2, 3, 4, and 6 is retained in the hidden layer vectors output by the current decoding layer.

[0083] During the process of the feed-forward network module calculating the reserved hidden layer vector, the feed-forward network module uses the following structure to calculate the decoding output: FFN(T) = (T′W U ⊙A(T′W G ))W D ). Among them, W U , W G and W D are the weight matrices of the feed-forward network module, and the dimensions are fixed values; T is the input of the feed-forward network module; T′ is the vector obtained after normalizing T, T′ = Norm(T), and the dimension of T′ is (query vector length, hidden layer vector size), and the hidden layer vector size is a fixed value. In the long text inference scenario, the dimensions of the weight matrix and the hidden layer vector size are very large, and the computational intensity of the feed-forward network module will be very high. It can be seen that by pruning Tokens and reducing the size of the input of the feed-forward network module, thereby reducing the size of T′, the computational overhead of the feed-forward network module can be effectively reduced.

[0084] The specific method for the feed-forward network module to calculate the reserved hidden layer vector to obtain the decoding output can be referred to the prior art and will not be elaborated in the embodiments of the present application.

[0085] The inventor found by observing the attention scores of each layer that even if the importance of a certain Token is very low in the previous layer (i.e., the attention score is very low), it may still become important again in subsequent layers. Therefore, the practice of directly and roughly discarding a certain Token is unreasonable. Based on this discovery, in the embodiments of the present application, a method of discarding Tokens according to the attention score ratio threshold is proposed. A certain Token may have a very low score in the previous layer and is discarded in this layer, but in subsequent layers, once its importance increases, it may still participate in the calculation. By using this method, not only can the inference speed of the model be improved, but also the inference accuracy of the model can be guaranteed.

[0086] Specifically, taking the input of a prompt with a length of 1000 Tokens to the large language model disclosed in the embodiments of the present application as an example, first set the specified observation window length N = 32. The first ten layers are calculated normally. For the subsequent layers, first calculate the attention of the last 32 Tokens with all other Tokens to obtain the attention scores of all Tokens relative to these 32 Tokens, and sum them in the 32 - dimension, that is, obtain the attention scores corresponding to 968 Tokens. Sort these attention scores and retain the top K Tokens such that the sum of the attention scores of these K Tokens accounts for more than the attention score threshold (such as 90%) of the attention scores of all 968 Tokens, obtaining K and the subscripts of these K Tokens. At this time, K is generally much less than 968. Adding the last retained 32 Tokens, the total retained length is K + 32. At this time ′ the length of the query vector of T is reduced from 1000 to K + 32, and the computational amount is reduced by nearly 1000 / (K + 32) times.

[0087] In the output generation step, the large language model will restore the digital sequence (i.e., the decoding output) output by the decoding layer into a natural language text that can be understood by humans through a tokenizer as the inference output corresponding to the text to be processed.

[0088] In some optional embodiments, after vectorizing the Token encoding in the first sequence to obtain the second sequence composed of Token vectors corresponding to the text to be processed, the method includes: in response to obtaining the inference acceleration flag set by the user and the length of the second sequence satisfying the first length threshold, or, in response to the second sequence satisfying the first length threshold, performing pruning and decoding processing on the Token vectors based on the importance indicators calculated during the decoding process of each Token vector in the second sequence to obtain the inference output corresponding to the text to be processed; in response to not obtaining the inference acceleration flag set by the user, or, the length of the second sequence not satisfying the first length threshold, performing decoding processing on each Token vector in the second sequence based on the attention mechanism to obtain the inference output corresponding to the text to be processed.

[0089] Among them, the first length threshold can be set according to specific application scenarios. For example, the first length threshold can be set to 1000. When the Token length in the second sequence is greater than or equal to 1000, it is considered that the length of the second sequence satisfies the first length threshold; otherwise, it is considered that the length of the second sequence does not satisfy the first length threshold.

[0090] As described above, in the case of a relatively long text to be processed, the inference method described in the foregoing embodiments can effectively improve the inference speed of the first Token. It can also be seen from the inference method described in the foregoing embodiments that in the case of a relatively short text to be processed, the inference speed will not be significantly improved, and moreover, the inference accuracy may be affected. Therefore, in the stage of applying the above inference method, the user can set an inference acceleration flag according to the specific inference scenario to initiate the operation of pruning Tokens in the foregoing embodiments. Alternatively, the inference application can automatically determine whether to initiate the operation of pruning Tokens in the foregoing embodiments according to the length of the text to be processed. For example, the Token pruning is only initiated when the second sequence length of the Tokens corresponding to the text to be processed is greater than or equal to the first length threshold. Alternatively, the Token pruning is only initiated when the second sequence length of the Tokens corresponding to the text to be processed is greater than or equal to the first length threshold and the inference acceleration flag set by the user. When the second sequence length of the Tokens corresponding to the text to be processed is less than the first length threshold, or when the inference acceleration flag set by the user is not set, the Token pruning is not initiated, but all the hidden layer vectors output by the self-attention module are input to the feed-forward network module for calculation.

[0091] In summary, the text inference method disclosed in the embodiments of the present application obtains a first sequence composed of Token encodings corresponding to the text to be processed by performing input processing on the text to be processed; performs vectorization processing on the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be processed; based on the attention scores corresponding to the Token vectors in the second sequence, dynamically selects some of the hidden layer vectors corresponding to the Token vectors for decoding processing based on the attention mechanism, performs pruning on the Tokens input to the feed-forward network module and then executes decoding output, effectively reducing the time required for the large language model to generate the first Token while ensuring less precision loss when performing text inference using the large language model, improving the text inference speed, and bringing a better large language model usage experience to users.

[0092] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.

[0093] Based on the above embodiments, this embodiment further provides a question answering method, which is applied to the process of reasoning on the question text based on a large language model to obtain an answer text. As Figure 4 shown, the method includes: step 402 to step 406.

[0094] Step 402, perform input processing on the question text to obtain a first sequence composed of Token encodings corresponding to the question text.

[0095] Step 404, perform vectorization processing on the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the question text.

[0096] Step 406, based on the importance indicators calculated during the decoding process for each Token vector in the second sequence, perform pruning and decoding processing on the Token vectors to obtain the answer text corresponding to the question text.

[0097] For the specific implementation of performing input processing on the question text to obtain a first sequence composed of Token encodings corresponding to the question text, refer to the specific implementation of performing input processing on the text to be processed in the previous embodiments to obtain a first sequence composed of Token encodings corresponding to the text to be processed, which will not be elaborated here.

[0098] For the specific implementation of performing vectorization processing on the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the question text, refer to the specific implementation of performing vectorization processing on the Token encodings in the first sequence in the previous embodiments to obtain a second sequence composed of Token vectors corresponding to the question text, which will not be elaborated here.

[0099] Optionally, the performing pruning and decoding processing on the Token vectors based on the importance indicators calculated during the decoding process for each Token vector in the second sequence to obtain the answer text corresponding to the question text includes:

[0100] Based on the importance indicators calculated during the decoding process for each second Token vector and each first Token vector, select the hidden layer vectors corresponding to the first Token vectors and partial hidden layer vectors corresponding to the second Token vectors for decoding processing to obtain the answer text corresponding to the question text, where the first Token vectors are: Token vectors within a specified observation window, and the second Token vectors are: Token vectors outside the specified observation window.

[0101] Optionally, the importance metric includes: an attention score. Based on the importance metric calculated during the decoding process between each second Token vector and each first Token vector in the second sequence, selecting the hidden layer vector corresponding to the first Token vector and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing to obtain the answer text corresponding to the question text, including:

[0102] Based on the attention scores respectively calculated between each second Token vector and each first Token vector in the second sequence in a preset multi-layer decoding layer, and the attention score threshold corresponding to the corresponding decoding layer, selecting the hidden layer vector corresponding to the first Token vector and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing to obtain the answer text corresponding to the question text.

[0103] Optionally, the specified observation window is: the hidden layer vectors corresponding to the last N Tokens in the second sequence, where N is a positive integer greater than or equal to 32.

[0104] Optionally, the decoding layer includes: a self-attention module and a feed-forward network module. Based on the attention scores respectively calculated between each second Token vector and each first Token vector in the second sequence in a preset multi-layer decoding layer, and the attention score threshold corresponding to the corresponding decoding layer, selecting the hidden layer vector corresponding to the first Token vector and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing to obtain the answer text corresponding to the question text, including:

[0105] Successively taking each decoding layer as the current decoding layer, and performing the following decoding operations through the current decoding layer to obtain the decoding output corresponding to the current decoding layer;

[0106] Obtaining the answer text corresponding to the question text based at least on the decoding output corresponding to the last decoding layer;

[0107] Among them, the decoding operations include:

[0108] Obtaining the hidden layer vector output by the self-attention module in the current decoding layer;

[0109] Based on the attention scores calculated between each second Token vector and each first Token vector in the second sequence in the current decoding layer, and the attention score threshold corresponding to the current decoding layer, dynamically selecting some of the hidden layer vectors as the retained hidden layer vectors;

[0110] Input the retained hidden layer vector and the hidden layer vector corresponding to the first Token vector into the feed-forward network module in the current decoding layer for processing to obtain a processing result;

[0111] Integrate the processing result and the hidden layer vectors other than the retained hidden layer vector to obtain the decoding output corresponding to the current decoding layer.

[0112] Optionally, the step of sequentially using each decoding layer as the current decoding layer includes:

[0113] Sequentially use each decoding layer after the L-th layer as the current decoding layer, where L is an integer greater than or equal to 10.

[0114] Optionally, the attention score threshold corresponding to the decoding layer is preset according to the importance of the decoding layer.

[0115] Optionally, after vectorizing the Token encoding in the first sequence to obtain the second sequence composed of Token vectors corresponding to the problem text, the method includes:

[0116] In response to obtaining the inference acceleration flag set by the user and the length of the second sequence satisfying the first length threshold, or in response to the second sequence satisfying the first length threshold, perform the step of pruning and decoding the Token vectors based on the importance metrics calculated during the decoding process of each Token vector in the second sequence to obtain the answer text corresponding to the problem text;

[0117] In response to not obtaining the inference acceleration flag set by the user, or the length of the second sequence not satisfying the first length threshold, perform decoding processing on each Token vector in the second sequence based on the attention mechanism to obtain the answer text corresponding to the problem text.

[0118] For the specific implementation of pruning and decoding the Token vectors based on the importance metrics calculated during the decoding process of each Token vector in the second sequence to obtain the answer text corresponding to the problem text, refer to the specific implementation of pruning and decoding the Token vectors based on the importance metrics calculated during the decoding process of each Token vector in the second sequence to obtain the inference output corresponding to the text to be processed in the foregoing embodiments.

[0119] In summary, the question answering method disclosed in the embodiments of the present application processes the input of the question text to obtain a first sequence composed of Token encodings corresponding to the question text; vectorizes the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the question text; based on the attention scores corresponding to each Token vector in the second sequence, dynamically selects some of the hidden layer vectors corresponding to the Token vectors for decoding processing based on the attention mechanism, prunes the Tokens input to the feed-forward network module, and performs inference through the feed-forward network module to obtain the answer text corresponding to the question text. This method effectively reduces the time required for the large language model to generate the first Token while ensuring less accuracy loss when invoking the large language model for text inference, bringing a better large language model usage experience to users.

[0120] Based on the above embodiments, this embodiment further provides a translation method, which is applied to the scenario of text translation based on a large language model. As Figure 5 shown, the method includes: steps 502 to 506.

[0121] Step 502, perform input processing on the text to be translated to obtain a first sequence composed of Token encodings corresponding to the text to be translated.

[0122] Step 504, vectorize the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be translated.

[0123] Step 506, based on the importance indicators calculated during the decoding process for each Token vector in the second sequence, prune and decode the Token vectors to obtain the translation result corresponding to the text to be translated.

[0124] For the specific implementation manner of performing input processing on the text to be translated to obtain a first sequence composed of Token encodings corresponding to the text to be translated, refer to the specific implementation manner of performing input processing on the text to be processed to obtain a first sequence composed of Token encodings corresponding to the text to be processed in the previous embodiments, which will not be elaborated here.

[0125] For the specific implementation manner of vectorizing the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be translated, refer to the specific implementation manner of vectorizing the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be processed in the previous embodiments, which will not be elaborated here.

[0126] Optionally, pruning and decoding the Token vectors based on the importance metrics calculated during the decoding process for each Token vector in the second sequence, and obtaining the translation result corresponding to the text to be translated, includes:

[0127] Based on the importance metrics calculated during the decoding process for each second Token vector and each first Token vector in the second sequence, select the hidden layer vectors corresponding to the first Token vectors and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing, and obtain the translation result corresponding to the text to be translated, where the first Token vectors are: Token vectors within a specified observation window, and the second Token vectors are: Token vectors outside the specified observation window.

[0128] Optionally, the importance metrics include: attention scores. Based on the importance metrics calculated during the decoding process for each second Token vector and each first Token vector in the second sequence, selecting the hidden layer vectors corresponding to the first Token vectors and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing, and obtaining the translation result corresponding to the text to be translated, includes:

[0129] Based on the attention scores calculated respectively for each second Token vector and each first Token vector in the second sequence in a preset multi-layer decoding layer, and the attention score threshold corresponding to the corresponding decoding layer, select the hidden layer vectors corresponding to the first Token vectors and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing, and obtain the translation result corresponding to the text to be translated.

[0130] Optionally, the specified observation window is: the hidden layer vectors corresponding to the last N Token in the second sequence, where N is a positive integer greater than or equal to 32.

[0131] Optionally, the decoding layer includes: a self-attention module and a feed-forward network module. Based on the attention scores calculated respectively for each second Token vector and each first Token vector in the second sequence in a preset multi-layer decoding layer, and the attention score threshold corresponding to the corresponding decoding layer, selecting the hidden layer vectors corresponding to the first Token vectors and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing, and obtaining the translation result corresponding to the text to be translated, includes:

[0132] Successively use each decoding layer as the current decoding layer, and perform the following decoding operations through the current decoding layer to obtain the decoding output corresponding to the current decoding layer;

[0133] Obtain the translation result corresponding to the text to be translated based at least on the decoding output corresponding to the last decoding layer;

[0134] Wherein, the decoding operation includes:

[0135] Obtain the hidden layer vector output by the self-attention module in the current decoding layer;

[0136] Based on the attention scores calculated between each second Token vector and each first Token vector in the second sequence in the current decoding layer, and the attention score threshold corresponding to the current decoding layer, dynamically select some of the hidden layer vectors as the retained hidden layer vectors;

[0137] Input the retained hidden layer vectors and the hidden layer vectors corresponding to the first Token vectors into the feed-forward network module in the current decoding layer for processing to obtain a processing result;

[0138] Integrate the processing result and the hidden layer vectors other than the retained hidden layer vectors to obtain the decoding output corresponding to the current decoding layer.

[0139] Optionally, the step of sequentially taking each decoding layer as the current decoding layer includes:

[0140] Sequentially take each decoding layer after the L-th layer as the current decoding layer, where L is an integer greater than or equal to 10.

[0141] Optionally, the attention score threshold corresponding to the decoding layer is preset according to the importance of the decoding layer.

[0142] Optionally, after vectorizing the Token encoding in the first sequence to obtain the second sequence composed of the Token vectors corresponding to the text to be translated, the method includes:

[0143] In response to obtaining the inference acceleration flag set by the user and the length of the second sequence satisfying the first length threshold, or in response to the second sequence satisfying the first length threshold, execute the step of pruning and decoding the Token vectors based on the importance indicators calculated during the decoding process of each Token vector in the second sequence to obtain the translation result corresponding to the text to be translated;

[0144] In response to not obtaining the inference acceleration flag set by the user, or the length of the second sequence not satisfying the first length threshold, perform decoding processing on each Token vector in the second sequence based on the attention mechanism to obtain the translation result corresponding to the text to be translated.

[0145] For the specific implementation of pruning and decoding the Token vectors based on the importance metrics calculated during the decoding process of each Token vector in the second sequence, and obtaining the translation result corresponding to the text to be translated, refer to the specific implementation of pruning and decoding the Token vectors based on the importance metrics calculated during the decoding process of each Token vector in the second sequence in the previous embodiments, and obtaining the inference output corresponding to the text to be processed.

[0146] In summary, the translation method disclosed in the embodiments of the present application obtains a first sequence composed of Token encodings corresponding to the text to be translated by performing input processing on the text to be translated; vectorizes the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be translated; dynamically selects hidden layer vectors corresponding to some of the Token vectors for decoding processing based on the attention mechanism based on the attention scores corresponding to each Token vector in the second sequence, prunes the Tokens input to the feed-forward network module, and performs inference through the feed-forward network module to obtain the translation result corresponding to the text to be translated. This method effectively reduces the time required to generate the first Token during the translation inference process of the large language model while ensuring less precision loss when using the large language model for translation, which helps to improve the translation speed and brings a better experience of using the large language model to users.

[0147] Based on the above embodiments, this embodiment further provides a text inference device, which includes:

[0148] A first sequence acquisition module, configured to perform input processing on the text to be processed to obtain a first sequence composed of Token encodings corresponding to the text to be processed;

[0149] A second sequence acquisition module, configured to vectorize the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be processed;

[0150] An inference module, configured to prune and decode the Token vectors based on the importance metrics calculated during the decoding process of each Token vector in the second sequence, and obtain the inference output corresponding to the text to be processed.

[0151] Optionally, the inference module is further configured to:

[0152] Based on the importance metrics calculated during the decoding process between each second Token vector in the second sequence and each first Token vector, select the hidden layer vectors corresponding to the first Token vectors and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing to obtain the inference output corresponding to the text to be processed, where the first Token vectors are: Token vectors within a specified observation window, and the second Token vectors are: Token vectors outside the specified observation window.

[0153] Optionally, the specified observation window is: the hidden layer vectors corresponding to the last N Tokens in the second sequence, where N is a positive integer greater than or equal to 32.

[0154] Optionally, the importance metrics include: attention scores. The process of selecting the hidden layer vectors corresponding to the first Token vectors and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing based on the importance metrics calculated during the decoding process between each second Token vector in the second sequence and each first Token vector to obtain the inference output corresponding to the text to be processed includes:

[0155] Based on the attention scores calculated respectively between each second Token vector in the second sequence and each first Token vector in a preset multi-layer decoding layer, and the attention score thresholds corresponding to the respective decoding layers, select the hidden layer vectors corresponding to the first Token vectors and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing to obtain the inference output corresponding to the text to be processed.

[0156] Optionally, the decoding layer includes: a self-attention module and a feed-forward network module. The process of selecting the hidden layer vectors corresponding to the first Token vectors and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing based on the attention scores calculated respectively between each second Token vector in the second sequence and each first Token vector in a preset multi-layer decoding layer, and the attention score thresholds corresponding to the respective decoding layers to obtain the inference output corresponding to the text to be processed includes:

[0157] Successively take each decoding layer as the current decoding layer, and perform the following decoding operations through the current decoding layer to obtain the decoding output corresponding to the current decoding layer;

[0158] Obtain the inference output corresponding to the text to be processed based at least on the decoding output corresponding to the last decoding layer;

[0159] Among them, the decoding operations include:

[0160] Obtain the hidden layer vectors output by the self-attention module in the current decoding layer;

[0161] Based on the attention scores calculated in the current decoding layer between each second Token vector and each first Token vector in the second sequence, and the attention score threshold corresponding to the current decoding layer, dynamically select some of the hidden layer vectors as the retained hidden layer vectors;

[0162] Input the retained hidden layer vectors and the hidden layer vectors corresponding to the first Token vectors into the feed-forward network module in the current decoding layer for processing to obtain a processing result;

[0163] Integrate the processing result and the hidden layer vectors other than the retained hidden layer vectors to obtain the decoding output corresponding to the current decoding layer.

[0164] Optionally, the step of dynamically selecting some of the hidden layer vectors as the retained hidden layer vectors based on the attention scores calculated in the current decoding layer between each second Token vector and each first Token vector in the second sequence, and the attention score threshold corresponding to the current decoding layer, includes:

[0165] Obtain the attention scores corresponding to each second Token vector as the attention scores corresponding to the corresponding hidden layer vectors;

[0166] Select K hidden layer vectors as the retained hidden layer vectors in the order of the corresponding attention scores from large to small, where the sum of the attention scores corresponding to the K hidden layer vectors is greater than or equal to the attention score threshold corresponding to the current decoding layer, and K is an integer greater than 1. The attention score threshold is less than 1.

[0167] Optionally, the step of sequentially using each decoding layer as the current decoding layer includes:

[0168] Sequentially use each decoding layer after the L-th layer as the current decoding layer, where L is an integer greater than or equal to 10.

[0169] Optionally, the attention score threshold corresponding to the decoding layer is preset according to the importance of the decoding layer.

[0170] Optionally, after vectorizing the Token encoding in the first sequence to obtain the second sequence composed of Token vectors corresponding to the text to be processed, the method includes:

[0171] In response to obtaining the inference acceleration flag set by the user and the length of the second sequence satisfying the first length threshold, or, in response to the second sequence satisfying the first length threshold, perform the importance index calculated during the decoding process based on each Token vector in the second sequence, perform pruning and decoding processing on the Token vector, and obtain the inference output corresponding to the text to be processed;

[0172] In response to not obtaining the inference acceleration flag set by the user, or the length of the second sequence not satisfying the first length threshold, perform decoding processing based on the attention mechanism on each Token vector in the second sequence, and obtain the inference output corresponding to the text to be processed. Among them, the first length threshold can be set according to the specific application scenario.

[0173] The text inference device disclosed in the embodiments of the present application is used to implement the above text inference method. For the specific implementation manners of each module of the device, refer to the specific implementation manners of the corresponding steps in the foregoing method embodiments, which will not be elaborated here.

[0174] In summary, the text inference device disclosed in the embodiments of the present application obtains the first sequence composed of Token encodings corresponding to the text to be processed by performing input processing on the text to be processed; performs vectorization processing on the Token encodings in the first sequence to obtain the second sequence composed of Token vectors corresponding to the text to be processed; based on the attention scores corresponding to each Token vector in the second sequence, dynamically selects partial hidden layer vectors corresponding to the Token vectors to perform decoding processing based on the attention mechanism, prunes the Tokens input to the feed-forward network module, and performs inference through the feed-forward network module to obtain the inference output, effectively reducing the time required for the large language model to generate the first Token while ensuring less loss of model accuracy, and bringing a better large language model usage experience to users.

[0175] Based on the above embodiments, this embodiment further provides a question answering device, which is applied to the process of inferring a question text based on a large language model to obtain an answer text. The device includes:

[0176] A first sequence acquisition module, configured to perform input processing on a question text to obtain a first sequence composed of Token encodings corresponding to the question text.

[0177] A second sequence acquisition module, configured to perform vectorization processing on the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the question text.

[0178] The answer text acquisition module is used to perform pruning and decoding processing on the Token vectors based on the importance metrics calculated during the decoding process of each Token vector in the second sequence, and acquire the answer text corresponding to the question text.

[0179] For the specific implementation of performing input processing on the question text to obtain the first sequence composed of the Token encodings corresponding to the question text, refer to the specific implementation of performing input processing on the text to be processed in the foregoing embodiments to obtain the first sequence composed of the Token encodings corresponding to the text to be processed, which will not be elaborated here.

[0180] For the specific implementation of performing vectorization processing on the Token encodings in the first sequence to obtain the second sequence composed of the Token vectors corresponding to the question text, refer to the specific implementation of performing vectorization processing on the Token encodings in the first sequence in the foregoing embodiments to obtain the second sequence composed of the Token vectors corresponding to the question text, which will not be elaborated here.

[0181] Optionally, the answer text acquisition module is further configured to:

[0182] Based on the importance metrics calculated during the decoding process of each second Token vector and each first Token vector in the second sequence, select the hidden layer vectors corresponding to the first Token vectors and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing, and acquire the answer text corresponding to the question text, where the first Token vectors are the Token vectors within a specified observation window, and the second Token vectors are the Token vectors outside the specified observation window.

[0183] Optionally, the importance metrics include attention scores. Based on the importance metrics calculated during the decoding process of each second Token vector and each first Token vector in the second sequence, selecting the hidden layer vectors corresponding to the first Token vectors and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing to acquire the answer text corresponding to the question text includes:

[0184] Based on the attention scores respectively calculated for each second Token vector and each first Token vector in the second sequence in a preset multi-layer decoding layer, and the attention score threshold corresponding to the corresponding decoding layer, select the hidden layer vectors corresponding to the first Token vectors and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing, and acquire the answer text corresponding to the question text.

[0185] Optionally, the specified observation window is: the hidden layer vectors corresponding to the last N Tokens in the second sequence, where N is a positive integer greater than or equal to 32.

[0186] Optionally, the decoding layer includes: a self-attention module and a feed-forward network module. Based on the attention scores respectively calculated between each second Token vector and each first Token vector in the second sequence in a preset multi-layer decoding layer, and the attention score threshold corresponding to the corresponding decoding layer, selecting the hidden layer vectors corresponding to the first Token vector and some of the hidden layer vectors corresponding to the second Token vectors for decoding processing to obtain the answer text corresponding to the question text, including:

[0187] Successively taking each decoding layer as the current decoding layer, and performing the following decoding operations through the current decoding layer to obtain the decoding output corresponding to the current decoding layer;

[0188] Obtaining the answer text corresponding to the question text based at least on the decoding output corresponding to the last decoding layer;

[0189] Among them, the decoding operations include:

[0190] Obtaining the hidden layer vector output by the self-attention module in the current decoding layer;

[0191] Based on the attention scores calculated between each second Token vector and each first Token vector in the second sequence in the current decoding layer, and the attention score threshold corresponding to the current decoding layer, dynamically selecting some of the hidden layer vectors as the retained hidden layer vectors;

[0192] Inputting the retained hidden layer vectors and the hidden layer vectors corresponding to the first Token vector into the feed-forward network module in the current decoding layer for processing to obtain a processing result;

[0193] Integrating the processing result and the hidden layer vectors other than the retained hidden layer vectors to obtain the decoding output corresponding to the current decoding layer.

[0194] Optionally, the successively taking each decoding layer as the current decoding layer includes:

[0195] Successively taking each decoding layer after the L-th layer as the current decoding layer, where L is an integer greater than or equal to 10.

[0196] Optionally, the attention score threshold corresponding to the decoding layer is preset according to the importance of the decoding layer.

[0197] Optionally, after vectorizing the Token encoding in the first sequence to obtain a second sequence composed of Token vectors corresponding to the question text, the apparatus includes: a judgment module.

[0198] The judgment module is configured to execute the answer text acquisition module in response to obtaining an inference acceleration flag set by the user and the length of the second sequence satisfying a first length threshold, or in response to the second sequence satisfying the first length threshold;

[0199] The judgment module is further configured to, in response to not obtaining the inference acceleration flag set by the user, or the length of the second sequence not satisfying the first length threshold, perform decoding processing based on an attention mechanism on each Token vector in the second sequence to obtain the answer text corresponding to the question text.

[0200] For the specific implementation of pruning and decoding the Token vectors based on the importance metrics calculated during the decoding process of each Token vector in the second sequence to obtain the answer text corresponding to the question text, refer to the specific implementation of pruning and decoding the Token vectors based on the importance metrics calculated during the decoding process of each Token vector in the second sequence to obtain the inference output corresponding to the text to be processed in the foregoing embodiments.

[0201] The question answering apparatus disclosed in the embodiments of the present application is used to implement the above question answering method. For the specific implementation of each module of the apparatus, refer to the specific implementation of the corresponding steps in the foregoing method embodiments, which will not be elaborated here.

[0202] In summary, the question answering apparatus disclosed in the embodiments of the present application processes the question text as input to obtain a first sequence composed of Token encodings corresponding to the question text; vectorizes the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the question text; dynamically selects hidden layer vectors corresponding to some of the Token vectors for decoding processing based on an attention mechanism based on the attention scores corresponding to each Token vector in the second sequence, prunes the Tokens input to the feed-forward network module, and performs inference through the feed-forward network module to obtain the answer text corresponding to the question text. This apparatus effectively reduces the time required for the large language model to generate the first Token while ensuring less accuracy loss of the large language model, bringing a better large language model usage experience to users.

[0203] Based on the above embodiments, this embodiment further provides a translation apparatus applied to a scenario of text translation based on a large language model. The apparatus includes:

[0204] The first sequence acquisition module is used to perform input processing on the text to be translated to obtain a first sequence composed of Token encodings corresponding to the text to be translated.

[0205] The second sequence acquisition module is used to perform vectorization processing on the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be translated.

[0206] The translation result acquisition module is used to perform pruning and decoding processing on the Token vectors based on the importance indicators calculated during the decoding process for each Token vector in the second sequence, and obtain the translation result corresponding to the text to be translated.

[0207] For the specific implementation manner of performing input processing on the text to be translated to obtain a first sequence composed of Token encodings corresponding to the text to be translated, refer to the specific implementation manner of performing input processing on the text to be processed to obtain a first sequence composed of Token encodings corresponding to the text to be processed in the foregoing embodiments, which will not be elaborated here.

[0208] For the specific implementation manner of performing vectorization processing on the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be translated, refer to the specific implementation manner of performing vectorization processing on the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be processed in the foregoing embodiments, which will not be elaborated here.

[0209] Optionally, the performing pruning and decoding processing on the Token vectors based on the importance indicators calculated during the decoding process for each Token vector in the second sequence and obtaining the translation result corresponding to the text to be translated includes:

[0210] Based on the importance indicators calculated during the decoding process for each second Token vector and each first Token vector in the second sequence, select the hidden layer vectors corresponding to the first Token vectors and partial hidden layer vectors corresponding to the second Token vectors for decoding processing, and obtain the translation result corresponding to the text to be translated, where the first Token vector is: the Token vector within a specified observation window, and the second Token vector is: the Token vector outside the specified observation window.

[0211] Optionally, the importance metric includes: an attention score. Based on the importance metric calculated during the decoding process between each second Token vector and each first Token vector in the second sequence, selecting the hidden layer vector corresponding to the first Token vector and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing to obtain the translation result corresponding to the text to be translated, includes:

[0212] Based on the attention scores respectively calculated between each second Token vector and each first Token vector in the second sequence in a preset multi-layer decoding layer, and the attention score threshold corresponding to the corresponding decoding layer, selecting the hidden layer vector corresponding to the first Token vector and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing to obtain the translation result corresponding to the text to be translated.

[0213] Optionally, the specified observation window is: the hidden layer vectors corresponding to the last N Tokens in the second sequence, where N is a positive integer greater than or equal to 32.

[0214] Optionally, the decoding layer includes: a self-attention module and a feed-forward network module. Based on the attention scores respectively calculated between each second Token vector and each first Token vector in the second sequence in a preset multi-layer decoding layer, and the attention score threshold corresponding to the corresponding decoding layer, selecting the hidden layer vector corresponding to the first Token vector and the hidden layer vectors corresponding to some of the second Token vectors for decoding processing to obtain the translation result corresponding to the text to be translated, includes:

[0215] Successively taking each decoding layer as the current decoding layer, and performing the following decoding operations through the current decoding layer to obtain the decoding output corresponding to the current decoding layer;

[0216] Obtaining the translation result corresponding to the text to be translated at least based on the decoding output corresponding to the last decoding layer;

[0217] Wherein, the decoding operation includes:

[0218] Obtaining the hidden layer vector output by the self-attention module in the current decoding layer;

[0219] Based on the attention scores calculated between each second Token vector and each first Token vector in the second sequence in the current decoding layer, and the attention score threshold corresponding to the current decoding layer, dynamically selecting some of the hidden layer vectors as the retained hidden layer vectors;

[0220] Input the reserved hidden layer vector and the hidden layer vector corresponding to the first Token vector into the feed-forward network module in the current decoding layer for processing to obtain a processing result;

[0221] Integrate the processing result and the hidden layer vectors other than the reserved hidden layer vector to obtain the decoding output corresponding to the current decoding layer.

[0222] Optionally, the step of sequentially using each decoding layer as the current decoding layer includes:

[0223] Sequentially use each decoding layer after the L-th layer as the current decoding layer, where L is an integer greater than or equal to 10.

[0224] Optionally, the attention score threshold corresponding to the decoding layer is preset according to the importance of the decoding layer.

[0225] Optionally, after vectorizing the Token encoding in the first sequence to obtain the second sequence composed of Token vectors corresponding to the text to be translated, the apparatus includes:

[0226] A judgment module, configured to execute the translation result acquisition module in response to obtaining an inference acceleration flag set by the user and the length of the second sequence satisfying a first length threshold, or in response to the second sequence satisfying the first length threshold;

[0227] The judgment module is further configured to perform decoding processing based on the attention mechanism on each Token vector in the second sequence to obtain the translation result corresponding to the text to be translated in response to not obtaining the inference acceleration flag set by the user, or the length of the second sequence not satisfying the first length threshold.

[0228] For the specific implementation of pruning and decoding the Token vectors based on the importance metrics calculated during the decoding process of each Token vector in the second sequence to obtain the translation result corresponding to the text to be translated, refer to the specific implementation of pruning and decoding the Token vectors based on the importance metrics calculated during the decoding process of each Token vector in the second sequence to obtain the inference output corresponding to the text to be processed in the foregoing embodiments.

[0229] The translation apparatus disclosed in the embodiments of the present application is used to implement the above translation method. For the specific implementation of each module of the apparatus, refer to the specific implementation of the corresponding steps in the foregoing method embodiments, which will not be elaborated here.

[0230] In summary, the translation device disclosed in the embodiments of the present application processes the input of the text to be translated to obtain a first sequence composed of Token encodings corresponding to the text to be translated; vectorizes the Token encodings in the first sequence to obtain a second sequence composed of Token vectors corresponding to the text to be translated; based on the attention scores corresponding to each Token vector in the second sequence, dynamically selects some of the hidden layer vectors corresponding to the Token vectors for decoding processing based on the attention mechanism, prunes the Tokens input to the feed-forward network module, and performs inference through the feed-forward network module to obtain the translation result corresponding to the text to be translated. On the basis of ensuring less accuracy loss of the large language model, this device effectively reduces the time required for the large language model to generate the first Token during the translation inference process, helps to improve the translation speed, and brings a better large language model usage experience to users.

[0231] The embodiments of the present application further provide a non-volatile readable storage medium, in which one or more modules (programs) are stored. When the one or more modules are applied to a device, the device can be caused to execute instructions (instructions) for each method step in the embodiments of the present application.

[0232] The embodiments of the present application further provide a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method as described in the embodiments of the present application.

[0233] The embodiments of the present application further provide an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method as described in the embodiments of the present application. In the embodiments of the present application, the electronic device includes devices such as servers and terminal devices.

[0234] The embodiments of the present application further disclose a computer program product, including a computer program / computer-executable instructions, characterized in that when the computer program / computer-executable instructions are executed by a processor in an electronic device, the method as described in the embodiments of the present application is implemented.

[0235] The embodiments of the present disclosure can be implemented as a device configured as desired using any suitable hardware, firmware, software, or any combination thereof. The device may include electronic devices such as servers (clusters) and terminals. Figure 6 Exemplary device 600 that can be used to implement the various embodiments described in the present application is schematically shown.

[0236] For one embodiment, Figure 6An exemplary device 600 is shown, which has one or more processors 602, a control module (chipset) 604 coupled to at least one of the processors 602, a memory 606 coupled to the control module 604, a non-volatile memory (NVM) / storage device 608 coupled to the control module 604, one or more input / output devices 610 coupled to the control module 604, and a network interface 612 coupled to the control module 604.

[0237] The processor 602 may include one or more single-core or multi-core processors, and the processor 602 may include any combination of general-purpose processors or dedicated processors (such as a graphics processor, an application processor, a baseband processor, etc.). In some embodiments, the device 600 can act as a server, a terminal, or other devices described in the embodiments of the present application.

[0238] In some embodiments, the device 600 may include one or more computer-readable media (such as the memory 606 or the NVM / storage device 608) having instructions 614, and one or more processors 602 combined with the one or more computer-readable media and configured to execute the instructions 614 to implement modules so as to perform the actions described in the present disclosure.

[0239] For one embodiment, the control module 604 may include any suitable interface controller to provide any suitable interface to at least one of the processors 602 and / or any suitable device or component communicating with the control module 604.

[0240] The control module 604 may include a memory controller module to provide an interface to the memory 606. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0241] The memory 606 can be used, for example, to load and store data and / or instructions 614 for the device 600. For one embodiment, the memory 606 may include any suitable volatile memory, such as a suitable DRAM. In some embodiments, the memory 606 may include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).

[0242] For one embodiment, the control module 604 may include one or more input / output controllers to provide an interface to the NVM / storage device 608 and the input / output devices 610.

[0243] For example, the NVM / storage device 608 can be used to store data and / or instructions 614. The NVM / storage device 608 can include any suitable non-volatile memory (e.g., flash memory) and / or can include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more compact disc (CD) drives, and / or one or more digital versatile disc (DVD) drives).

[0244] The NVM / storage device 608 can include storage resources that are part of a device installed as device 600, or it can be accessible by the device without necessarily being part of the device. For example, the NVM / storage device 608 can be accessed via a network through the (one or more) input / output devices 610.

[0245] (One or more) input / output devices 610 can provide an interface for device 600 to communicate with any other suitable devices. The input / output devices 610 can include communication components, audio components, sensor components, etc. The network interface 612 can provide an interface for device 600 to communicate through one or more networks. Device 600 can wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing a wireless network based on communication standards like Bluetooth, WiFi, 2G, 3G, 4G, 5G, etc., or a combination thereof for wireless communication.

[0246] For one embodiment, at least one of the (one or more) processors 602 can be logically packaged together with one or more controllers of the control module 604 (e.g., a memory controller module). For one embodiment, at least one of the (one or more) processors 602 can be logically packaged together with one or more controllers of the control module 604 to form a system-in-package (SiP). For one embodiment, at least one of the (one or more) processors 602 can be logically integrated on the same die with one or more controllers of the control module 604. For one embodiment, at least one of the (one or more) processors 602 can be logically integrated on the same die with one or more controllers of the control module 604 to form a system-on-chip (SoC).

[0247] In various embodiments, the device 600 may be, but is not limited to, a terminal device such as a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the device 600 may have more or fewer components and / or a different architecture. For example, in some embodiments, the device 600 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touch screen display), a non-volatile memory port, multiple antennas, a graphics chip, an application specific integrated circuit (ASIC), and a speaker.

[0248] Among them, a main control chip can be used as a processor or a control module in the detection device, sensor data, location information, etc. are stored in a memory or an NVM / storage device, the sensor group can be used as an input / output device, and the communication interface can include a network interface.

[0249] Embodiments of the present application also provide an electronic device, including: a processor; and a memory having executable code stored thereon, which when executed, causes the processor to execute one or more of the methods as in the embodiments of the present application. In the embodiments of the present application, various data such as target files, file-application association data, etc. can be stored in the memory, and user behavior data, etc. can also be included, thereby providing a data basis for various processes.

[0250] Embodiments of the present application also provide one or more machine-readable media having executable code stored thereon, which when executed, causes a processor to execute one or more of the methods as in the embodiments of the present application.

[0251] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0252] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.

[0253] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate for implementation in the process Figure 1one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks.

[0254] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks.

[0255] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal device, such that a series of operational steps are performed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks.

[0256] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.

[0257] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal device including the said element.

[0258] The above has introduced in detail a text inference method, a question answering method, a translation method, an electronic device, a storage medium, and a computer program product provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A text reasoning method, characterized in that: The method comprises: Performing input processing on the text to be processed to obtain a first sequence consisting of Token codes corresponding to the text to be processed; Vectorizing the Token codes in the first sequence to obtain a second sequence consisting of Token vectors corresponding to the text to be processed; Based on the importance index of each Token vector in the second sequence calculated during the decoding process, the Token vector is pruned and decoded to obtain the inference output corresponding to the text to be processed.

2. The method according to claim 1, characterized in that The step of pruning and decoding the Token vectors based on the importance index calculated during the decoding process of each Token vector in the second sequence to obtain the inference output corresponding to the text to be processed includes: Based on the importance indexes calculated during the decoding process of each second Token vector and each first Token vector in the second sequence, the hidden layer vector corresponding to the first Token vector and some of the hidden layer vectors corresponding to the second Token vector are selected for decoding processing to obtain the inference output corresponding to the text to be processed, wherein the first Token vector is: a Token vector within a specified observation window, and the second Token vector is: a Token vector outside the specified observation window.

3. The method according to claim 2, characterized in that The importance index includes: an attention score, the importance index calculated based on each second Token vector and each first Token vector in the second sequence during the decoding process, selecting a hidden layer vector corresponding to the first Token vector and a portion of hidden layer vectors corresponding to the second Token vector for decoding processing, and obtaining an inference output corresponding to the text to be processed, including: Based on the attention scores respectively calculated for each second Token vector and each first Token vector in the second sequence in the preset multi-layer decoding layers, and the attention score threshold corresponding to the corresponding decoding layer, the hidden layer vector corresponding to the first Token vector and part of the hidden layer vectors corresponding to the second Token vector are selected for decoding processing to obtain the inference output corresponding to the text to be processed.

4. The method according to claim 2, characterized in that: The designated observation window is: the hidden layer vectors corresponding to the last N tokens in the second sequence, where N is a positive integer greater than or equal to 32.

5. The method according to claim 3, characterized in that: The decoding layer includes: a self-attention module and a feedforward network module. The attention scores calculated in the preset multi-layer decoding layers respectively based on each second Token vector and each first Token vector in the second sequence, and the attention score threshold corresponding to the corresponding decoding layer, select the hidden layer vector corresponding to the first Token vector and the hidden layer vector corresponding to part of the second Token vector for decoding processing, and obtain the inference output corresponding to the text to be processed, including: Taking each decoding layer as the current decoding layer in turn, and performing the following decoding operations through the current decoding layer to obtain a decoding output corresponding to the current decoding layer; Based on at least the decoding output corresponding to the last decoding layer, obtaining the inference output corresponding to the text to be processed; The decoding operation includes: Obtain the hidden layer vector output by the self-attention module in the current decoding layer; Dynamically select part of the hidden layer vectors as reserved hidden layer vectors based on the attention scores calculated in the current decoding layer by each second Token vector and each first Token vector in the second sequence, and the attention score threshold corresponding to the current decoding layer; Input the retained hidden layer vector and the hidden layer vector corresponding to the first Token vector into the feedforward network module in the current decoding layer for processing to obtain a processing result; The processing result and the hidden layer vector other than the retained hidden layer vector are integrated to obtain a decoding output corresponding to the current decoding layer.

6. The method according to claim 5, characterized in that The method dynamically selects part of the hidden layer vectors as reserved hidden layer vectors based on the attention scores calculated in the current decoding layer by each second Token vector and each first Token vector in the second sequence, and the attention score threshold corresponding to the current decoding layer, including: Obtain the attention score corresponding to each of the second Token vectors as the attention score corresponding to the corresponding hidden layer vector; In descending order of the corresponding attention scores, K hidden layer vectors are selected as retained hidden layer vectors, wherein the sum of the attention scores corresponding to the K hidden layer vectors is greater than or equal to the attention score threshold corresponding to the current decoding layer, and K is an integer greater than 1.

7. The method according to claim 5, characterized in that The sequentially taking each decoding layer as the current decoding layer includes: The decoding layers after the Lth layer are sequentially used as current decoding layers, where L is an integer greater than or equal to 10.

8. The method according to claim 3, characterized in that The attention score threshold corresponding to the decoding layer is pre-set according to the importance of the decoding layer.

9. The method according to claim 1, characterized in that: After vectorizing the token codes in the first sequence to obtain a second sequence consisting of token vectors corresponding to the text to be processed, the method includes: In response to obtaining the inference acceleration flag set by the user and the length of the second sequence meeting the first length threshold, or in response to the second sequence meeting the first length threshold, executing the step of pruning and decoding the Token vector based on the importance index calculated during the decoding process of each Token vector in the second sequence, and obtaining the inference output corresponding to the text to be processed; In response to not obtaining the inference acceleration flag set by the user, or the length of the second sequence does not meet the first length threshold, each Token vector in the second sequence is decoded based on the attention mechanism to obtain the inference output corresponding to the text to be processed.

10. A problem-solving method, characterized in that: The method comprises: Input processing is performed on the question text to obtain a first sequence consisting of Token codes corresponding to the question text; Vectorizing the Token codes in the first sequence to obtain a second sequence consisting of Token vectors corresponding to the question text; Based on the importance index of each Token vector in the second sequence calculated during the decoding process, the Token vector is pruned and decoded to obtain the answer text corresponding to the question text.

11. A text translation method, characterized in that: The method comprises: Performing input processing on the text to be translated to obtain a first sequence consisting of Token codes corresponding to the text to be translated; Vectorizing the Token codes in the first sequence to obtain a second sequence consisting of Token vectors corresponding to the text to be translated; Based on the importance index of each Token vector in the second sequence calculated during the decoding process, the Token vector is pruned and decoded to obtain a translation result corresponding to the text to be translated.

12. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 11 when executed by a processor.

14. A computer program product comprising a computer program / computer executable instructions, characterized in that: When the computer program / computer executable instructions are executed by a processor in an electronic device, the method according to any one of claims 1 to 11 is implemented.

Citation Information

Cited By

  • Large language model reasoning method, device, equipment, medium and program product

    CN121615762A

  • Large language model reasoning method, device, equipment, medium and program product

    CN121615762B