Context understanding-oriented generative reading understanding method and device, and medium

Through the parallel decoding strategy of pipeline decoder, the problem of the generation speed bottleneck of generative reading comprehension model is solved, faster inference speed and lower memory usage are achieved, while maintaining higher generation quality, which is suitable for generative reading comprehension tasks in natural language processing.

CN120493936APending Publication Date: 2025-08-15NANJING UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510526896.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

There is a bottleneck in the generation speed of existing generative reading comprehension models under the autoregression paradigm. The gradual generation method limits the generation speed of the model, especially when dealing with long input contexts.

Method used

The pipeline decoder is used to obtain vector representations using the encoder of the pre-trained language model, and generate understanding results through parallel decoding, set the delay time Δt and the maximum number of subsequences submax. The pipeline decoder generates a new subsequence every Δt, and uses a three-dimensional mask matrix to limit the attention of the subsequence. The encoder part is frozen during training, and the pipeline decoder is trained.

Benefits of technology

The rapid inference speed and low GPU memory usage of generative reading comprehension are realized. The generation quality performs well on phrase-level and sentence-level datasets. Compared with the sequential decoder, the inference speed is 2.0-5.2 times. The GPU memory usage is reduced, and the generation quality is only slightly reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493936A_ABST
    Figure CN120493936A_ABST
Patent Text Reader

Abstract

According to the context-understanding-oriented generative reading understanding method and device and the medium, vector representation of a to-be-read understanding text is obtained by using an encoder of a pre-training language model, a pipeline decoder is used during decoding, an understanding result is generated for the vector representation in a parallel decoding mode, the pipeline decoder is a stacked Transformer decoder, and the pipeline decoder is used for decoding the to-be-read understanding text. And starting to generate a new sub-sequence every delay time delta t, generating the first lexical element of the sub-sequence by depending on the lexical element generated by the previous sub-sequence until the maximum sub-sequence number is reached or all the previous sub-sequences are subjected to lexical element generation, and forming a final understanding result by the sub-sequences obtained by decoding. The pipeline decoder has good performance in reading and understanding of phrase-level and sentence-level data sets, and compared with a sequence decoder, the pipeline decoder has higher reasoning speed and lower GPU (Graphics Processing Unit) memory usage amount, and the generation quality can meet the requirement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer artificial intelligence technology, and relates to natural language processing, machine reading comprehension and generative artificial intelligence. It is a generative reading comprehension method, device and medium for contextual understanding. Background Art

[0002] With the rapid development of artificial intelligence, natural language processing (NLP), as one of its key branches, has made significant progress. NLP aims to enable computers to understand, interpret, and generate natural language, enabling them to interact with humans more effectively. Machine reading comprehension (MRC) is a core task in NLP, aiming to enable computers to automatically extract, understand, and answer questions from text. MRC tasks are generally divided into two categories: extractive and generative. Extractive MRC requires the model to extract a fragment of an answer from a given text, typically used when the answer is explicitly stated. Generative MRC, on the other hand, requires the model to generate a complete answer, capable of handling open-ended questions and reasoning. Traditional machine reading comprehension methods rely heavily on extractive MRC techniques. While this approach has achieved some success in specific scenarios, it is limited in that it can only process direct information within a text and lacks the flexibility to handle complex semantic reasoning and abstract problems. In contrast, generative reading comprehension requires models to not only generate complete answers but also handle open-ended questions and flexibly reason and generate across diverse contexts. Therefore, generative reading comprehension is considered an important task in natural language processing and is currently a hot topic of research.

[0003] In recent years, with the development of deep learning and large-scale pre-trained language models such as BERT and GPT, generative reading comprehension has gained significant momentum. In particular, techniques based on attention mechanisms and self-supervised learning have significantly improved the models' contextual understanding and generation capabilities. Despite this, generative reading comprehension still faces many challenges, such as contextual understanding, reasoning capabilities, and generation quality, which limit the models' performance in practical applications. Therefore, improving the accuracy, generation capabilities, and reasoning performance of generative models remains a critical issue in this field.

[0004] The remarkable capabilities of generative models for text generation have ushered in a new era for artificial intelligence, significantly improving productivity across various fields. As the applications of these models continue to expand and user demand increases, optimizing their generation efficiency has become a key research direction. A key challenge is the quadratic time complexity of the self-attention mechanism, which scales with the length of the input text, leading to a significant efficiency bottleneck. This issue of generation efficiency is particularly prominent in context-aware text generation tasks, such as enhanced retrieval generation, text summarization, and keyword generation. Models often need to process a long input context when generating new text, making the efficiency issue even more significant. Summary of the Invention

[0005] The problem addressed by this invention is that the existing autoregressive paradigm faces a generation speed bottleneck. Because many generative models require a stepwise autoregressive generation based on all previously generated labels, the generation process is often slow. While this stepwise generation approach can ensure accurate answers, it also significantly limits the model's generation speed.

[0006] The technical solution of the present invention is: a generative reading comprehension method for contextual understanding, which uses a pre-trained language model encoder to obtain a vector representation of the text to be read and understood, and uses a pipeline decoder during decoding to generate an understanding result by parallel decoding based on the vector representation. The pipeline decoder is a stacked Transformer decoder, and hyperparameters are set, including delay time Δt, maximum subsequence number sub max and the maximum time step time max , the pipeline decoder starts to generate a new subsequence every delay time Δt. Each subsequence contains the predicted word. In the first time step, subsequence 1 is generated. After an interval of Δt, subsequence 2 is generated in parallel. The first word of subsequence 2 depends on the word generated by the previous subsequence, and so on, until the maximum number of subsequences is reached or the previous subsequences have completed word generation. In the generation of each subsequence, when the decoding generates an end mark or the maximum time step time is reached max When , the subsequence generation stops, and the subsequences obtained by decoding constitute the final understanding result.

[0007] Furthermore, when the decoder input exceeds 2048 words, a sliding window mechanism is used for batch decoding, and the window overlap rate is set to 15%.

[0008] Furthermore, an attention mask matrix is used to restrict each subsequence to access only the preceding subsequence. A three-dimensional mask consisting of three orthogonal dimensions is used: combining batch, subsequence, and word-unit mask information. Each dimension corresponds to a different semantic control objective, allowing each subsequence to be decoded and output using the BeamSearch decoding strategy in the stacked transformers.

[0009] Furthermore, if no new token is generated for 5 consecutive time steps, the decoding is terminated early.

[0010] Furthermore, when training the pipeline decoder, for the answers in the training data, the answer text is divided into several subsequence targets by periods, and in the i-th target subsequence Y i Add the start tag before <bos>, in Y i Add the closing tag after <eos>, thereby constructing the target subsequence ^Yi for training ^ =[ <bos> ,Y i , <eos>], add an empty subsequence Y to the end of the subsequence F =[ <bos> , <eos>] to learn the generation of the terminator sequence and obtain the target sequence Y of the answer ^ =[^Y1 ^ ,…,^Yn ^ ,Y F ] as training samples, freeze the encoder part of the pre-trained language model during training, and train the pipeline decoder.

[0011] The text to be read and understood in the present invention includes a text composed of a text and questions and a text to be summarized.

[0012] The present invention also provides an electronic device, including a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the above-mentioned generative reading comprehension method for context understanding.

[0013] The present invention also provides a computer-readable storage medium storing a computer program, which, when directly executed or called for execution, implements the above-mentioned generative reading comprehension method for contextual understanding.

[0014] The beneficial effects of the present invention are as follows: The pipeline decoder proposed in the generative reading comprehension method of the present invention can perform subsequence decoding in parallel. Compared with the prior art method of relying on the previous word to decode step by step, the present invention proposes to generate a new decoding subsequence based on the word predicted by the existing decoding subsequence after a set delay time. Finally, multiple subsequences are decoded in parallel, achieving rapid understanding of the question and text vector and generating answers.

[0015] The pipeline decoder proposed in the present invention performs well in reading comprehension of both phrase-level and sentence-level datasets. When the decoder of the present invention is replaced with existing generative reading comprehension models such as T5base and T5-large, compared with the sequential decoders of these existing reading comprehension models under the traditional autoregressive paradigm, the pipeline decoder of the present invention not only has higher inference speed and lower GPU memory usage, but also meets the generation quality requirements.

[0016] 1) Experimental results on phrase-level dataset

[0017] Inference Speed: On the MSQA multi-answer question-answering dataset, the pipelined decoder of our invention achieved a 1.7x speedup compared to the sequential decoder. On the KP20K and KPTimes datasets, using the T5-Base model, the acceleration effect of our method was even more significant, increasing the generation speed by 2.0x and 2.5x, respectively, compared to the original decoder. This is because the number of target subsequences in KP20K and KPTimes is larger than that in MSQA, allowing our pipelined decoder to build more parallel subsequences, resulting in a more significant speedup.

[0018] Generation Quality: Although the pipeline decoder of the present invention only partially relies on the word at the initial position when generating subsequences, its generation quality is not inferior to that of the sequential decoder with complete word dependencies according to most indicators. In the MSQA question answering task, the EM score of the pipeline decoder of the present invention on T5-Base only slightly decreased by 0.8% on the test set compared with the sequential decoder. In the experimental results, the pipeline decoder of the present invention surpassed the sequential decoder in 8 of the 16 indicators. For example, in the KP20K dataset, the F1@5 and F1@M indicators of the sequential decoder in the missing keyword prediction task were 2.2 and 4.2, while the indicators of the pipeline decoder of the present invention were 4.0 and 7.6. In the missing keyword prediction task of the KP20K dataset, the improvement of the quality accuracy indicator of the present invention ranged from 1.8% to 3.4%.

[0019] GPU memory usage: Compared to the sequential decoder, the pipelined decoder does not add additional computational overhead, and due to only partial token dependencies, the pipelined decoder uses even less memory.

[0020] The pipelined decoder of our invention performs well on three phrase-level datasets, including generation quality, inference speed (i.e., throughput), and GPU memory usage. Compared to the sequential decoder, the pipelined decoder not only has higher inference speed and lower GPU memory usage, but also meets the generation quality requirements.

[0021] 2) Experimental results on sentence-level dataset

[0022] Inference Speed: On the WikiHowQA and PubMed datasets, the pipelined decoder achieves at least 5.2x and 6.5x higher throughput than the sequential decoder, respectively.

[0023] Generation Quality: Unlike phrase-level tasks in short text generation, sentence-level subsequence generation faces a higher risk of missing dependencies. In this regard, our pipelined decoder maintains a close relationship with the sequential decoder on most metrics, with the largest differences being only 0.5%, 1.4%, and 2.1% on WikiHowQA, CNN / DMail, and PubMed, respectively.

[0024] GPU memory usage: The pipelined decoder still uses less memory compared to the sequential decoder. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 Schematic diagram comparing the pipeline decoder of the present invention and the sequential decoder of the prior art in terms of reading comprehension decoding.

[0026] Figure 2 Schematic diagram of time step iteration in the pipeline decoder of the present invention.

[0027] Figure 3 This is a network architecture diagram of the pipeline decoder of the present invention.

[0028] Figure 4 The difference in attention mechanism distribution between the pipeline decoder of the present invention and the existing sequential decoder, (a) is the sequential decoder, and (b) is the pipeline decoder.

[0029] Figure 5 This is an example of the dependency pattern of the pipeline decoder of the present invention. DETAILED DESCRIPTION

[0030] The present invention proposes a generative reading comprehension method for contextual understanding. The method uses a pre-trained language model encoder to obtain a vector representation of the text to be read and understood. A pipeline decoder is used for decoding. The understanding result is generated by parallel decoding based on the vector representation. The pipeline decoder is a stacked Transformer decoder with hyperparameters, including delay time Δt, maximum subsequence number sub max and the maximum time step time max , the pipeline decoder starts to generate a new subsequence every delay time Δt. Each subsequence contains the predicted word unit. At the first time step, subsequence 1 is generated. After an interval of Δt, subsequence 2 is generated in parallel. The first word unit of subsequence 2 depends on the word unit generated by the previous subsequence, and so on, until the maximum number of subsequences is reached or the previous subsequences have completed word unit generation. Each subsequence is decoded by a layer of transformer decoder. When the decoding generates an end mark or the maximum time step time is reached max When , subsequence decoding stops, and the decoded subsequences form the final comprehension result. This paper proposes an accelerated weak dependency model, in which each token in a subsequence only partially depends on tokens in the subsequence generated before that time step. Subsequences are then decoded in parallel. By decoding multiple subsequences of the reading comprehension text in parallel, inference speed is accelerated.

[0031] The hidden state of a word can encode information about the current and future tags. The present invention has found that the generation of each word does not necessarily need to strictly rely on all previous words. The generation quality is mainly affected by a few key words, such as the first word. The method of the present invention generates parallel subsequences based on the vector representation of the question and text, and decodes to generate the answer. The decoder only relies on the first few words that have been generated in the previous subsequence to start the generation of a new subsequence. The pipeline decoder proposed in the present invention starts the generation of multiple subsequences in sequence in a pipeline manner. Under a fixed delay time, the pipeline decoder first generates a new subsequence, and generates multiple new subsequences in parallel under the delay time. Each word pays attention to all context words and words generated in the previous subsequence. Figure 1 The following shows an example of context-aware question answering based on the WikiHowQA dataset. The context and question obtain the vector representation of the text and question. The decoding process of the traditional sequential decoder and the pipeline decoder of the present invention is as follows: Figure 1 As shown in , unlike a sequential decoder that relies on the entire preceding subsequence "Go to Start Menu" to generate tokens step by step, the pipeline decoder of the present invention relies solely on the first token "Go" to generate the first token "Check" of the second subsequence, with a delay of one time step. In this embodiment, the pipeline decoder generates a maximum of four tokens in parallel, which is the maximum number of predicted tokens in the subsequence. Therefore, whereas a traditional sequential decoder generates tokens sequentially, requiring 24 time steps to generate a complete answer, the pipeline decoder, by generating multiple subsequences in parallel, requires only eight time steps.

[0032] Figure 2 The pipeline decoding process of the present invention is shown. Assume that the answer to a complete data set is "ABCDE", where A, B, C, D, and E represent different word units. The sequential decoder of the existing technology needs to generate word units step by step in sequence. The method of the present invention uses a pipeline decoder to split the answer sequence into multiple subsequences and generate them in parallel. For example, if the answer is split into the sequence "ABCDE", <sep>DE", where the two subsequences "ABC" and "DE" are separated by the delimiter <sep>Separated, the pipeline decoder generates them in parallel with a delay of one time step. Unlike the sequential decoder that generates one word unit at each time step, the pipeline decoder generates 1, 2, and 2 words units in the 1st to 3rd time steps respectively. Specifically, in the 1st time step, the pipeline decoder generates word unit "A" in subsequence 1; after a delay of one time step, in the 2nd time step, it generates subsequence 2, and synchronously generates word units "B" and "D" in subsequence 1 and subsequence 2, with "D" as the first word unit of subsequence 2; in the 3rd time step, it generates "C" and "E" in subsequence 1 and subsequence 2 respectively, and tries to create a new subsequence 3. The third subsequence generates an end marker based on the generation results of the previous subsequence. <eos>, the creation of this new subsequence fails. Finally, at the 4th time step, when the semantic generation of a subsequence is complete, a terminator eos is generated to indicate that this subsequence does not need to be generated anymore. Subsequence 1, subsequence 2, and the new subsequence 3 are all generated. <eos>, marking the end of pipeline decoding. The splitting of the answer sequence and the number of subsequence tokens are learned dynamically through training.

[0033] The implementation of the present invention is described in detail below.

[0034] (1) The word unit representation module is used to obtain the vector representation of the text, also known as token. A pre-trained language model such as BERT is used to obtain the text representation, refine the vector representation of words, capture richer semantic information, and improve the representation ability of the text representation. The present invention also uses a pre-trained language model to obtain the vector representation of the text to be read and understood. The text to be read and understood includes text composed of text and questions and text to be summarized. That is, the method of the present invention can be used for reading and understanding summaries and generating answers based on question reading text.

[0035] (2) As a hyperparameter, the delay time Δt is set to control the pipeline decoder to start generating a new subsequence every Δt.

[0036] (3) Parallel decoding is performed through a stacked Transformer decoder. Specifically, unlike the sequential decoder, where the new token generated by the sequential decoder at time step t+1 depends on the previous t consecutive tokens, the token in the i-th subsequence generated by the pipeline decoder depends on the previous i-1 subsequences generated up to time step t, where the subsequence may not be completely generated. Figure 4 This paper demonstrates the differences in the dependencies between the sequential and pipeline decoders, comparing their attention distribution in the generated sequence. As shown by the dependencies marked by the solid black line, the pipeline decoder only needs to focus on the three previously generated tokens when generating the token "D," while the sequential decoder needs to focus on five. This significantly reduces the computational effort of the pipeline decoder. The dashed black line marks the generation dependency of token "E." Similarly, compared to the sequential decoder, the pipeline decoder of this invention reduces its reliance on token "C" when generating token "E," eliminating the need to wait for the entire sequence "ABC" to be decoded, thus speeding up computation time.

[0037] (4) Pipeline decoding strategy

[0038] 4.1) Initialization

[0039] Input: Input sequence X of question and text concatenation, delay time Δt, maximum time step time max , maximum number of subsequences sub max .

[0040] G is initialized to an empty list and is used to store the generated subsequences.

[0041] i is initialized to 0, indicating the number of subsequences currently generated.

[0042] 4.2) Coding stage

[0043] Use the Encoder function of the pre-trained language model to encode the input sequence X and obtain the encoded representation H E , which is the vector representation of the text to be read and understood.

[0044] 4.3) Decoding stage

[0045] For each time step t from 1 to time max :

[0046] If (t-1) mod i = 0, mod represents the modulo operation, that is, the previous time step corresponds to the generation of a subsequence, and the subsequence number i is less than sub max If so, it will indicate the start of a special symbol <bos>Add to G and increase the value of i to start a new subsequence G i+1 . Use the decoder function to decode the current G and H by time step E Decoding is performed, that is, decoding each subsequence stored in G in parallel.

[0047] Loop until i = sub max .

[0048] 4.4) Return results

[0049] Return the final G and get all subsequence results to form the final answer.

[0050] The corresponding decoding algorithm code is as follows:

[0051]

[0052]

[0053] (5) The stacked transformer decoder of the present invention constructs a parallel mask system. The transformer decoder uses a masked self-attention mechanism to ensure that when generating each symbol, the model can only see the symbols before it, but not the future symbols. The mask is a matrix used to shield the attention weights of future symbols. The decoder of the present invention uses a three-dimensional mask, which is composed of three orthogonal dimensions: the mask information of the three dimensions of batch, subsequence, and token. Each dimension corresponds to a different semantic control target. The mask system tensor shape is [batch_size, max_answer_num, max_length], which shields invalid filling positions and automatically ignores the output of non-target positions in the loss calculation. Figure 2 As shown, for the parallel decoding part, the present invention allows several subsequences to execute the BeamSearch decoding strategy in the stacked transformer, and generates word unit tokens of the subsequences in parallel on several branches.

[0054] (6) Training strategy. The training goal of the pipeline decoder is to enable it to start and end the generation of subsequences. The present invention uses the existing public training data set and then divides the answer text into several subsequence targets by periods to train the decoder. Specifically, the present invention trains the decoder in the i-th target subsequence Y i Add the start tag before <bos>, in Y i Add the closing tag after <eos>, thereby constructing the training target subsequence ^Yi ^ =[ <bos> ,Y i , <eos>]. When a sufficient number of subsequences are generated, in order to help the pipeline decoder learn to stop generating new subsequences, an empty subsequence Y is added at the end of the sequence F =[ <bos> , <eos>], indicating that no new subsequences will be generated. The complete training target sequence is Y ^ =[^Y1 ^ ,…,^Yn ^ ,Y F ].

[0055] The implementation of the present invention is demonstrated below through a specific embodiment.

[0056] 1. System Architecture

[0057] The generative reading comprehension system of this embodiment includes the following core modules:

[0058] The word representation module uses the BERT-base pre-trained model to vectorize the input text C and question Q. In specific implementation, the text is segmented into segments of 512 words or less, and a 12-layer Transformer encoder is used to obtain a 768-dimensional dynamic word vector. The formula is expressed as:

[0059] H E =BERT([CLS]⊕Q⊕[SEP]⊕C⊕[SEP])

[0060] Pipeline decoder: It consists of 6 layers of stacked Transformer decoders, each layer is equipped with self-attention mechanism and cross-attention mechanism. Compared with the traditional sequential decoder, this decoder can process k subsequences (k≤sub max ), since the decoder is trained to learn when to stop generating new subsequences, k here is determined dynamically based on the input text fragment.

[0061] Figure 3 The network layers included in the pipeline decoder of the present invention are demonstrated. Multiple subsequences realize dynamic information interaction and generation through cross-layer attention, and are subsequently converted into instance word units through a feedforward network and saved.

[0062] 2. Pipeline decoding implementation example

[0063] Take a piece of text as the reading comprehension object and generate a text summary for the text.

[0064] Set hyperparameters: delay time Δt = 2, maximum number of subsequences sub max =5, maximum time step time max =30, the implementation process is as follows:

[0065] Initialization phase: Create an empty subsequence list G and load pre-trained weights.

[0066] Time step iteration: decode existing subsequences in parallel, when all subsequences contain <eos>or reach time max Stop when

[0067] Implementation example:

[0068] Input text: Scientists recently achieved a major breakthrough in materials research, successfully synthesizing a new material. Experiments have shown that this material exhibits remarkable superconducting properties at extremely low temperatures, enabling lossless transmission of electric current. This discovery is considered a milestone in the development of quantum computing, as superconducting materials play a key role in the stability of quantum bits and the efficiency of information transmission. The relevant research results have been published in the journal Nature and have attracted widespread attention in the academic community.

[0069] Summary answer: Scientists have discovered a new type of material. This material has superconducting properties and is suitable for use in quantum computers.

[0070] Separating the answer by the period in the text will generate 3 subsequences:

[0071] Subsequence 1: [ <bos>Scientists discover new material <eos>]

[0072] Child order 2: <bos>The material has superconducting properties <eos>]

[0073] Child order 3: <bos>Application to quantum computers <eos>]

[0074] Figure 5 The microscopic decoding process of each subsequence in the above example is shown at the word-gram level. w1-w5, v1-v5, and u1-u3 represent the word-grams in each subsequence, respectively. This embodiment completes the decoding of all answers in 6 time steps, and multiple word-grams can be decoded in parallel in each time step.

[0075] 3. Exception Handling Example

[0076] For the actual text to be read and understood, the present invention further performs the following optimization processing.

[0077] 1) Long text processing: When the input exceeds 2048 words, a sliding window mechanism is used with a window overlap ratio of 15%;

[0078] 2) Subsequence conflict: The attention mask matrix is used to restrict each subsequence to only access the previous subsequence. Each token in the subsequent subsequence can depend on the previous subsequence, but not vice versa.

[0079] 3) Early stopping mechanism: If no new token is generated for 5 consecutive time steps, the decoding is terminated early.< / eos> < / bos> < / eos> < / bos> < / eos> < / bos> < / eos> < / eos> < / bos> < / eos> < / bos> < / eos> < / bos> < / bos> < / eos> < / eos> < / sep> < / sep> < / eos> < / bos> < / eos> < / bos> < / eos> < / bos>

Claims

1. A generative reading comprehension method for contextual understanding, characterized by The encoder of the pre-trained language model is used to obtain the vector representation of the text to be read and understood. The pipeline decoder is used for decoding to generate the understanding result by parallel decoding based on the vector representation. The pipeline decoder is a stacked Transformer decoder with hyperparameters set, including the delay time Δt, the maximum number of subsequences sub max and the maximum time step time max , the pipeline decoder starts to generate a new subsequence every delay time Δt. Each subsequence contains the predicted word. In the first time step, subsequence 1 is generated. After an interval of Δt, subsequence 2 is generated in parallel. The first word of subsequence 2 depends on the word generated by the previous subsequence, and so on, until the maximum number of subsequences is reached or the previous subsequences have completed word generation. In the generation of each subsequence, when the decoding generates an end mark or the maximum time step time is reached max When , the subsequence generation stops, and the subsequences obtained by decoding constitute the final understanding result.

2. The context-oriented generative reading comprehension method according to claim 1 is characterized by: When the decoder input exceeds 2048 words, a sliding window mechanism is used for batch decoding, and the window overlap rate is set to 15%.

3. The context-oriented generative reading comprehension method according to claim 1 is characterized by The attention mask matrix is used to restrict each subsequence to only access the previous subsequence. A three-dimensional mask is used, consisting of three orthogonal dimensions: batch, subsequence, and word-unit mask information. Each dimension corresponds to a different semantic control target, so that each subsequence can be decoded and output by executing the BeamSearch decoding strategy in the stacked transformer.

4. The context-oriented generative reading comprehension method according to claim 1 is characterized by: If no new word is generated for 5 consecutive time steps, the decoding is terminated early.

5. The context-oriented generative reading comprehension method according to claim 1 is characterized by: When training the pipeline decoder, for the answers in the training data, the answer text is divided into several subsequence targets by periods, and in the i-th target subsequence Y i Add the start tag before <bos>, in Y i Add the closing tag after <eos>, thereby constructing the target subsequence for training ^Yi^=[ <bos> ,Y i , <eos>], add an empty subsequence Y to the end of the subsequence F =[ <bos> , <eos>] to learn the generation of terminator sequences and obtain the target sequence Y^=[^Y1^,…,^Yn^,Y F ] as training samples, freeze the encoder part of the pre-trained language model during training, and train the pipeline decoder.< / eos> < / bos> < / eos> < / bos> < / eos> < / bos> 6. The context-oriented generative reading comprehension method according to claim 1 is characterized by: The text to be read and understood includes a text composed of text and questions and a text to be summarized.

7. An electronic device comprising a processor and a memory, characterized in that The memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the generative reading comprehension method for context understanding as described in any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that A computer program is stored, and when the computer program is directly executed or called for execution, the context-oriented generative reading comprehension method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Instruction processing device and method, electronic equipment and storage medium

    CN121387373A