Automatic Generation Method for Literature Related Work Based on Integrated Causal Intervention
The causal intervention method derived by the causal graph and do operator, combined with the Transformer model, solves the problem of pseudo-correlation relationship in related work generation, and generates higher quality related work texts.
Patent Information
- Application Number
- CN202211496491.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-24
AI Technical Summary
The existing related work generation technology ignores causality, resulting in the correlation learned by the model being pseudo-correlation, which affects the generalization ability of the model, and the generated related work is poor.
Through the causal graph, the sentence order, text relationship and transition content in the related work generation process are modeled, the causal intervention steps are derived using the do operator and backdoor criterion, the neural network module is designed for causal intervention, and fused with the Transformer model to generate high-quality related work text.
It effectively weakens the negative impact of pseudo-correlation relationships, improves the logical coherence and quality of related work generation, and improves the model generation ability.
Smart Images

Figure CN115827853B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for automatically generating the Related Work of literature integrating causal intervention, belonging to the technical field of artificial intelligence and natural language processing. Background Art
[0002] Academic research is usually based on the work of predecessors for further exploration and innovation. Therefore, when introducing new viewpoints and new methods to the public, it is necessary to introduce other relevant researches involved in the current research in this field at the same time, which can enable readers to quickly grasp the key points of the current research. This also indicates the importance of the related work section in academic papers.
[0003] A comprehensive related work necessarily includes a large number of reference documents, which will take a lot of time for the author to read and summarize, and even continuously pay attention to the latest research results. However, in recent years, with more and more scientific publications being digitized and publicly available on the Internet, it has become possible to establish large-scale datasets. Therefore, the task of automatically generating related work has received more and more attention from researchers.
[0004] Related work can be regarded as evolved from the multi-document summarization task. The biggest difference between it and the multi-document summarization task is that in the process of generating related work, not only the core content of different researches needs to be summarized, but also these researches need to be compared and the similarities and differences between them need to be obtained. Currently, most of the latest researches focus on generative methods. Xing et al. made the initial attempt in this task. They used the context of the cited sentences and the abstracts of the cited documents as the input of the model, and made the model output the related work text. Ge et al. used the graph attention network (GAT) to encode the citation network and introduced it as external knowledge into the generation process. Chen et al. proposed a Target-aware related work generation model, which captures the relationship between the cited documents and the target documents through the target-aware attention mechanism. Through more advanced encoding methods, introducing external knowledge, new training strategies and other ways, these researches can better ensure that the generated related work has good coherence and logic.
[0005] However, most of the existing related work generation techniques ignore the inherent causal relationships in the related work generation process, which may lead to the association relationships learned by the model being actually pseudo-correlation relationships and ultimately having a side effect on the generalization ability of the model. Therefore, we believe it is necessary to further explore the more reliable causal relationships in related work generation and learn them. For example, if certain preferences are introduced into a dataset intentionally or unintentionally during the production stage or the samples in it do not conform to independent and identically distributed (such as related work texts always continuously using certain conjunctions), then in these cases, most existing models will regard this preference as a reliable correlation relationship and ultimately result in good performance on the training set but poor performance on the test set.
[0006] Therefore, we propose to use causal intervention to eliminate the pseudo-correlation relationships in the related work generation process and prompt the model to learn more reliable causal relationships, ultimately obtaining related work texts of better quality. Summary of the Invention
[0007] The object of the present invention is to perform causal intervention on the related work generation process of the literature, weaken the negative impact of the existing pseudo-correlation relationships, and prompt the model to learn the correct causal relationships between related elements, so as to generate related work text information of higher quality.
[0008] The innovation of the present invention lies in: combining a causal graph to model the causal relationships among the three key elements in related work generation, namely sentence order, text relationship, and transitional content, and regarding the sentence order as a confounding factor, which establishes a pseudo-correlation relationship with the transitional content through the text relationship. Further, using the do operator and the backdoor criterion, the specific steps for implementing causal intervention are derived and a neural network module is designed based on these steps to implement causal intervention. Finally, the intervention module is integrated with the existing generation model to obtain a complete end-to-end related work generation model with causal intervention.
[0009] The present invention is implemented based on the following technical solutions.
[0010] A method for automatically generating literature Related Work based on integrated causal intervention, comprising the following steps:
[0011] Step 1: Model the causal relationships in the related work generation process.
[0012] Model the causal relationship among the three key elements in the generation of related work: sentence order, text relationship, and transitional content. Consider the sentence order as a confounding factor, which establishes a spurious correlation with the transitional content through the text relationship. Among them, the sentence order and the document relationship independently guide the generation of the transitional content.
[0013] To generate a more fluent and coherent related work text in terms of text organization and logical relationship, three elements play a key role and there is a causal relationship among them, namely: sentence order, document relationship, and transitional content.
[0014] Among them, the sentence order and the document relationship independently guide the generation of the transitional content. The relationship between the sentence order and the transitional content is more of a writing experience. For example, the author is used to using words such as "firstly" at the beginning of a paragraph and "finally" at the end of a paragraph. The relationship between the document relationship and the transitional content guides the text to compare the relationships between different studies.
[0015] Since the sentence order information is easier to capture than the document relationship information, and the samples in the dataset are prone to non - independent and identically distributed situations, it is very likely that the sentence order information participates in the process of the document relationship guiding the transitional content, which weakens the guiding role of the document relationship and makes the model prefer the sentence order information. Therefore, this method considers the sentence order as a confounding factor in this process, and a causal graph is constructed.
[0016] Step 2: Based on the causal relationship modeling results, use the do - operator to derive the intervention process.
[0017] To specifically implement the causal intervention process, based on the causal graph modeling results and introducing the do - operator and the backdoor criterion, the intervention process is derived to provide a theoretical basis for the implementation of the intervention. When the do - operator is introduced, the traditional conditional probability model will be modified, and the observed samples are no longer limited to the existing data, but the range of observed samples will be expanded by intentional data assignment.
[0018] Step 3: Design a causal intervention module to implement the specific intervention process.
[0019] Step 3.1: In the causal intervention module, first perform the original intervention. According to the derivation result of the do - operator, fuse the sentence order information into the process of the document relationship guiding the generation of the transitional content, and initially eliminate the negative impact of the spurious correlation relationship caused by the sentence order.
[0020] Step 3.2: Since the word representations after the original intervention may not be smoothly distributed, and the intervention will cause the destruction of the original context information it contains, therefore, through the context-aware remapping process, the representations after the intervention are recalculated. During this process, the context information is captured and fused, making the final representation distribution smoother and containing context information.
[0021] Step 3.3: In order to obtain the optimal intervention intensity, through the optimal intensity learning process, the output intensities corresponding to the existing original, intervened, and remapped representations are learned respectively. By fusing these three intensities, the optimal intervention intensity is obtained, and the final representation is calculated and output.
[0022] Step 4: The causal intervention module is fused with the Transformer model to obtain the final end-to-end related work generation model with causal intervention ability.
[0023] Since both the encoder and decoder of the Transformer include multiple Transformer Blocks, there are different choices for the fusion position and number of the causal intervention module. In the present invention, the intervention module is placed in the Transformer decoder, and tests are carried out at intervals of different numbers of Transformer Blocks.
[0024] Step 5: Generate related work through causal intervention.
[0025] After passing the final output of the Transformer decoder through the linear layer and the softmax layer, the specific word content at each position is obtained.
[0026] Step 6: Train the neural network, use cross-entropy as the loss function, calculate the loss value, and realize the self-generation of high-quality literature related work content.
[0027] Within the set number of training rounds, use cross-entropy to calculate the gap loss between the predicted generated related work and the real related work, and optimize and adjust the model parameters through backpropagation, finally making the model performance reach the optimal.
[0028] So far, the generation process of related work through causal intervention is completed, the negative impact of the pseudo-correlation relationship in this process is weakened, and the model realizes the self-generation of high-quality literature related work content by learning more reliable causal relationships.
[0029] Beneficial effects
[0030] The method of the present invention has the following advantages compared with the existing related work generation methods.
[0031] 1. This method explicitly models the causal relationships among the sentence order, document relationships, and transitional content in the related work generation process through a causal graph, and analyzes and points out that the sentence order, as a confounding factor, can cause spurious correlation relationships, laying a theoretical foundation for the subsequent implementation of causal intervention.
[0032] 2. Based on the causal graph modeling and the derivation results of the do-operator, this method uses a deep learning model to specifically implement causal intervention. For the original intervention, a context-aware remapping process and an optimal intervention strength learning process are further designed to enable the overall causal intervention effect to reach the optimal.
[0033] 3. This method effectively integrates the intervention module with the Transformer model, enabling the final related work generation model to not only have the good generation ability of the Transformer but also be able to perform causal intervention on the generation process on this basis, thereby further improving the model performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a flowchart of the method of the present invention.
[0035] Figure 2 is an implementation framework diagram of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0036] According to the above technical solutions, the method of the present invention will be described in detail below through specific embodiments.
[0037] Embodiment
[0038] The website address of the dataset selected in this embodiment is:
[0039] https: / / github.com / iriscxy / relatedworkgeneration
[0040] Select a set of samples from the embodiment, and each sample contains a set of cited documents D = {ref1, ref2,..., ref D} and the corresponding related work text Y = (w1, w2,..., w M ). The related work generation task receives D as the input and outputs the predicted related work text Due to considerations such as training costs, the abstract part of each cited document is selected to represent the content of the entire document.
[0041] As Figure 1 shown, the method for automatically generating related work of literature based on fused causal intervention includes the following steps:
[0042] Step 1: Model the causal relationships in the process of generating related work. Through the analysis of the generation process, it is considered that three elements play key roles, namely sentence order, document relationship, and transitional content, denoted as c, x, and y respectively.
[0043] Among them, there is a direct causal relationship between the sentence order and the transitional content, which is interpreted as a writing experience. For example, the author tends to use "firstly", "secondly", "finally" continuously. In addition, there is also a direct causal relationship between the document relationship and the transitional content, because related work often needs to compare and explain the relationships between different references. Thus, in the initially obtained causal graph, there are two paths: c→y and x→y. The sentence order is often confused with the document relationship and becomes a confounding factor involved in the guiding process of the document relationship, which leads to the establishment of the pseudo-correlation relationship c→x→y. This pseudo-correlation relationship will cause the model to fail to correctly learn the association between the document relationship and the transitional content. Thus, the causal graph G is completed.
[0044] Step 2: Based on the constructed causal graph G, use the do-operator to derive the causal intervention process:
[0045] P(y|do(x)) = ∑ c P(y|do(x), c)P(c|do(x))
[0046] = ∑ c P(y|x, c)P(c|do(x))
[0047] = ∑ c P(y|x, c)P(c)
[0048] where P(y|do(x)) represents deriving y based on the intervened x, c is the sentence order, x is the document relationship, and y is the transitional content.
[0049] Through the derivation of the do-operator, the path c→x in the causal graph G is cut off, weakening the influence of the pseudo-correlation relationship.
[0050] Step 3: Design a causal intervention module to implement the specific intervention process. In the design process, combine the already constructed causal graph G and the intervention derivation process.
[0051] Specifically, it includes the following steps:
[0052] Step 3.1: Perform the original causal intervention process.
[0053] During the decoding process, denote the original input as representing a single word representation in the decoded sequence, being the length of the decoded sequence; denote the original intervention output as representing the representation of a single word in the sequence after the original intervention.
[0054] First, implement the estimation process of P(y|x, c), fusing the sentence order information with the original representation:
[0055]
[0056] where, o j represents the position information of the j-th sentence, specifically as follows:
[0057]
[0058] O represents the set of s different sentence position information, represents the representation after fusing the i-th word with the j-th position information, s is the number of all possible sentence positions, which appears as a hyperparameter in the experiment. Therefore, The original has been enhanced for sentence order information at different j positions.
[0059] After being mapped by the linear layer, the sentence order information is fused with the original document relationship information, that is, P(y|x, c): = e odr , e odr generally refers to
[0060] Using the currently decoded subsequence Predict the possibility that the currently generated sentence is at different positions:
[0061]
[0062] where, h i represents the set of position distributions, represents the probability that the order of the current sentence i is the j-th sentence, Softmax is the normalized exponential function, FFN represents the feed-forward neural network, and ReLU represents the activation function.
[0063] So far, the estimation process of P(c): = h in the intervention derivation has been completed.
[0064] When the estimations of P(y|x, c) and P(c) are completed, multiply the two and sum over different sentence orders:
[0065]
[0066] Among them, represents the representation of the i-th word in the sequence after the original intervention process.
[0067] The multiplication with achieves the estimation of P(y|x, c)P(c). Finally, after summation, the original intervention process is completed.
[0068] Step 3.2: Since there are insufficient trainable modules in the original intervention process and the intervened representations may lose their original context information, therefore, through a context-aware remapping process, the original intervention representation E itv is remapped.
[0069] Specifically, use a context window of a fixed size n w to gradually scan the E itv sequence:
[0070]
[0071] B i represents the subsequence of length n w captured and returned by the context window WIN(·), which is the final output of Step 3.1.
[0072] The focus of remapping will be on the intermediate representation i of the subsequence B . Update this subsequence through the multi-head attention mechanism:
[0073]
[0074] Among them, represents the set of updated subsequence representations of B i , MultiHead represents the multi-head attention mechanism, represents the updated single-word representation.
[0075] Although all the representations in the entire output will be updated, only the intermediate representation is retained and used to replace the original intermediate representation itv of the input sequence E
[0076] Through the gradual advancement of the window, finally all the representations will be updated. Denote the output of the context-aware remapping as which represents the single-word representation of the output sequence.
[0077] Step 3.3: Since it is impossible to guarantee that the complete causal intervention intensity will definitely enable the model to achieve optimal performance, the final output intensity of different representations is adjusted through the optimal intervention intensity learning process.
[0078] In the foregoing steps, there are three representations: E ori , E itv and E rmp . Calculate their respective output intensities:
[0079]
[0080]
[0081]
[0082]
[0083] Among them, W ori , W itv , W rmp correspond to the output intensity calculation matrices of the original representation the original intervention representation output in Step 3.1 and the remapped representation output in Step 3.2 respectively. σ(·) is the sigmoid function, and Softmax is the normalized exponential function. Through the Softmax operation, the original output intensity is finally normalized into the final output intensity
[0084] By combining the final causal intervention representation
[0085]
[0086] Denote the final output as which represents the representation of a single word in the final output sequence.
[0087] Step 4: Integrate the causal intervention module with the Transformer model to obtain a final end-to-end related work generation model with causal intervention capabilities.
[0088] Although the causal intervention module has the ability to implement causal intervention, it still needs to be integrated with the generation model to obtain an end-to-end complete model. Considering the advantages of the Transformer model in the field of natural language processing in recent years, the causal intervention module is selected to be integrated with it.
[0089] During the fusion process, two problems need to be solved: First, in Transformer, through matrix operations, the representations of all words in the sequence are calculated in parallel, which makes it impossible to judge whether each word needs to be intervened one by one as in RNN and affect the subsequent generation process; Second, Transformer is composed of multiple Transformer Blocks, which makes the fusion strategy not unique.
[0090] Step 4.1: To solve the first problem, before each intervention, project the word representation onto the vocabulary and calculate the mask preds for its corresponding sentence start symbol [CLS]:
[0091]
[0092] Mask CLS = δ(preds, CLS_ID)
[0093] Among them, the argmax function will return the subscript of the maximum value in the sequence, and the linear layer Linear vocab will project the word representation to the length of the entire vocabulary, and preds contains the IDs corresponding to all words after projection. The δ(·) function will compare whether the two parameters are the same, return 1 if they are the same, and return 0 if they are different; CLS_ID is the ID value of the sentence start symbol in the vocabulary; Mask CLS represents whether each word is at the sentence start position through 0-1 values. The final E opm is calculated as:
[0094] E opm = E opm ⊙ Mask CLS + E ori ⊙ (~Mask CLS )
[0095] Among them, the ⊙ operation multiplies the representation in each E opm by the 0-1 value, and the ~ operation negates the values in Mask CLS . The final E opm only retains the word information after intervention and at the sentence start position, and the representations of other words are restored using the information in E ori .
[0096] Step 4.2: Place the causal intervention module in the middle of the Transformer Block to achieve its fusion with Transformer:
[0097] Fused-CaM(·) = CaM(TF-Block[·]) ×n
[0098]
[0099] Among them, Fused-CaM(·) represents the final model obtained by fusing the intervention module with the Transformer, TF-Block[·] represents a single Transformer Block, and CaM(·) is the causal intervention module. The output of the final fusion model passes through the linear layer Linear vocab The word probability distribution result obtained after mapping and Softmax normalization, where D is the task input, that is, the set of cited documents.
[0100] Step 5: Generate related work through causal intervention.
[0101] In Step 4.2, the normalized model prediction word probability distribution is finally obtained Among them is the length of the generated sequence, and d voc is the length of the vocabulary. The values in reveal the likelihood of each word at each position corresponding to each word. Using the argmax function to select the index of the maximum value at each word position, that is, the index of the word with the highest probability in the vocabulary VOC, denoted as I, and then using the index and the vocabulary to obtain the specific word, the finally generated related work text can be obtained
[0102]
[0103]
[0104] Among them, represents the set of real numbers of length , and I i represents the index value of each specific word in the vocabulary.
[0105] Step 6: After obtaining , use the cross-entropy loss function to calculate the gap between the true word probability distribution P i (Y) and the predicted word probability distribution
[0106]
[0107] So far, the generation process of related work through causal intervention is completed, weakening the negative impact of the pseudo-correlation relationship in this process. The model realizes the self-generation of high-quality literature related work content by learning more reliable causal relationships.
[0108] Figure 2 This is the implementation framework diagram of this method. The performance of the related work generated by this method in terms of F1 of the three metrics of ROUGE-1, ROUGE-2, and ROUGE-L is shown in the last row of Table 1.
[0109] Table 1 Comparison of the effects of 9 generation methods - 2 datasets
[0110]
Claims
1. A method for automatically generating literature Related Work based on fused causal intervention, characterized in that It includes the following steps: Step 1: Model the causal relationships in the generation process of related work; Model the causal relationships among the three elements in the generation of related work: sentence order, text relationship, and transitional content. Consider the sentence order as a confounding factor, which establishes a pseudo-correlation relationship with the transitional content through the text relationship. Among them, the sentence order and the document relationship independently guide the generation of the transitional content; Among them, the sentence order and the document relationship independently guide the generation of the transitional content. The relationship between the document relationship and the transitional content guides the text to compare the relationships between different studies; Step 2: Derive the intervention process using the do-operator based on the causal relationship modeling results; Step 3: Design a causal intervention module to implement the specific intervention process, as follows: Step 3.1: In the causal intervention module, first perform the original intervention. According to the derivation result of the do-operator, fuse the sentence order information into the process of guiding the generation of the transitional content by the document relationship, and initially eliminate the negative impact of the pseudo-correlation relationship caused by the sentence order; Step 3.2: Recalculate the intervened representation through the context-aware remapping process. During this process, capture and fuse its context information to make the final representation distribution smoother and contain context information; Step 3.3: Through the optimal strength learning process, learn the output strengths corresponding to the existing original, intervened, and remapped representations respectively. By fusing these three strengths, obtain the optimal intervention strength and calculate the output of the final representation; Step 4: Integrate the causal intervention module with the Transformer model to obtain a final end-to-end related work generation model with causal intervention capabilities; Step 5: Generate related work through causal intervention; After passing the final output of the Transformer decoder through the linear layer and the softmax layer, obtain the specific word content at each position; Step 6: Train the neural network, use cross-entropy as the loss function, calculate the loss value, and realize the self-generation of high-quality literature related work content; Within the set number of training rounds, use cross-entropy to calculate the gap loss between the predicted generated related work and the true related work, and optimize and adjust the model parameters through backpropagation, finally making the model performance reach the optimal; 2. The method for automatically generating literature Related Work based on fusion causal intervention according to claim 1, wherein, In Step 2, based on the causal graph constructed in Step 1, the derivation of the causal intervention process using the do-operator is as follows: P(y|do(x)) = ∑ c P(y|do(x), c)P(c|do(x)) = ∑ c P(y|x,c)P(c|do(x)) = ∑ c P(y|x,c)P(c) Among them, P(y|do(x)) represents deriving y based on the intervened x, c is the sentence order, x is the document relationship, and y is the transitional content. Through the derivation of the do-operator, the path c→x in the causal graph is cut off, weakening the influence of the pseudo-correlation relationship; 3. The method for automatically generating literature Related Work based on fusion causal intervention according to claim 1, wherein In step 3.1, during the decoding process, denote the original input as which represents a single word representation in the decoding sequence, and let \(L\) be the length of the decoding sequence; denote the original intervention output as which represents a single word representation in the sequence after the original intervention; First, implement the estimation process of P(y|x,c), and fuse the sentence order information with the original representation: where, o j represents the position information of the j-th sentence, specifically as follows: $O$ represents a set of $s$ different sentence position information, represents the representation after the $i$-th word is fused with the $j$-th position information, and $s$ is the number of all possible sentence positions, The original Enhance the sentence order information at different $j$ positions; After being mapped by the linear layer, the sentence order information is fused with the original document relationship information, i.e., P(y|x,c) ∶= e odr , e odr Generally referring to the above text Utilize the currently decoded subsequence Predict the probabilities of the currently generated sentence at different positions: Among them, h i represents a set of position distributions, represents the probability that the order of the current sentence i is the j-th sentence, Softmax is the normalized exponential function, FFN represents the feed-forward neural network, and ReLU represents the activation function; So far, the estimation process of P(c)∶=h in the intervention derivation is completed; After estimating P(y|x,c) and P(c), multiply the two and sum over different sentence orders: Among them, represents the representation of the i-th word in the sequence after the original intervention process; Multiplication with achieves the estimation of P(y|x,c)P(c), and finally, after summation, the original intervention process is completed; In step 3.2, through a context-aware remapping process, the original intervention representation E itv is remapped; Use a context window of a fixed size n w to gradually scan the E itv sequence: B i represents a subsequence of length n captured and returned by the context window WIN(·) w and is the final output of step 3.1; The focus of the remapping is on subsequence B i 's intermediate representation and updates this subsequence through the multi-head attention mechanism: Among them, represents B i the updated subsequence representation set, and MultiHead represents the multi-head attention mechanism, represents the updated single-word representation; Although all the representations in the output will be updated, only the intermediate representation is retained and used to replace the original intermediate representation of the input sequence E itv Denote the context-aware remapping output as representing a single word representation of the output sequence; In step 3.3, adjust the final output strength of different representations through the optimal intervention intensity learning process; Calculate the output intensities corresponding to three characterizations E ori , E itv and E rmp respectively: Among them, W ori , W itv , W rmp correspond to the original representation the original intervention representation output in Step 3.1 and the remapping representation output in Step 3.2 respectively; σ(·) is the sigmoid function, and Softmax is the normalized exponential function; through the Softmax operation, the original output intensity is finally normalized into the final output intensity By combining to obtain the final causal intervention representation Denote the final output as which represents the representation of a single word in the final output sequence.
4. The method for automatically generating literature Related Work based on fusion causal intervention according to claim 1, wherein Step 4 includes the following steps: Step 4.1: Before each intervention, project the word representation onto the vocabulary and calculate the mask preds for its corresponding sentence start token [CLS]: Mask CLS = δ(preds, CLS_ID) Among them, the argmax function returns the index of the maximum value in the sequence, and the linear layer Linear vocab projects the word representation to the length of the entire vocabulary. preds contains the IDs corresponding to all words after projection; the δ(·) function compares whether the two parameters are the same. If they are the same, it returns 1, and if they are different, it returns 0; CLS_ID is the ID value of the sentence start symbol in the vocabulary; Mask CLS represents whether each word is in the sentence start position through 0-1 values; the final E opm The calculation method is as follows: E opm = E opm ⊙ Mask CLS + E ori ⊙ (~Mask CLS ) Among them, the ⊙ operation multiplies each representation in E opm by a 0-1 value, and the ~ operation performs a negation operation on the values in Mask CLS ; the final E opm only retains the word information of the words after intervention and at the beginning of the sentence, and the representations of other words are restored using the information in E ori ; Step 4.2: Place the causal intervention module in the middle of the Transformer Block to achieve its integration with the Transformer: Fused-CaM(·) = CaM(TF-Block[·]) ×n Among them, Fused-CaM(·) represents the final model obtained by fusing the intervention module with the Transformer, TF-Block[·] represents a single Transformer Block, and CaM(·) is the causal intervention module; is the result of the model prediction word probability distribution obtained after the output of the final fusion model passes through the linear layer Linear vocab mapping and Softmax normalization, and D is the task input.
5. The method for automatically generating literature Related Work based on fused causal intervention according to claim 1, wherein, In step 5, the argmax function is used to select the index of the maximum value at each word position, that is, the index of the word with the highest probability in the vocabulary VOC, denoted as I. Then, the specific word is obtained by using the index and the vocabulary, and the finally generated related work text is obtained. Among them, is the result of the model's predicted word probability distribution obtained after normalization, is the length of the generated sequence, d voc is the length of the vocabulary, represents the set of real numbers of length I i represents the index value of each specific word in the vocabulary; In step 6, when obtaining use the cross-entropy loss function to calculate the gap between the true word probability distribution P i (Y) and the predicted word probability distribution Complete the generation of related work with causal intervention.