A discourse structure analysis system based on a multi-expert mechanism
By introducing the labeled data generated by multi-expert mechanisms and large language models, the structure and training data of the chapter structure analysis system are optimized, and the problem of insufficient generalization ability of existing methods in non-news texts is solved, and efficient utilization of corpus in different fields is achieved and generalization ability is improved.
Patent Information
- Application Number
- CN202510653624.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing deep learning-based chapter structure analysis method lacks generalization ability when dealing with non-news texts, and it is difficult to effectively utilize corpus with large domain differences. The scale and field breadth of available chapter data are insufficient, which limits its promotion and application in diverse text scenarios.
The chapter structure analysis system based on a multi-expert mechanism is adopted, and the automatic matching and efficient utilization of corpus corpus in different fields is achieved through word-level coding module, EDU-level coding module, decoding context learning module, multi-expert pointer network decoding module and pre-training module for large-model generation corpus, combined with fine-tuning of manual annotation chapter corpus, the system structure and training data are optimized to achieve automatic matching and efficient utilization of corpus in different fields.
It significantly improves the generalization ability of the system, enables it to better adapt to heterogeneous corpus, improves the performance of chapter structure analysis, and enhances the application effect in diverse text scenarios.
Smart Images

Figure CN120180242B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly to a discourse structure analysis system based on a multi-expert mechanism. Background Art
[0002] Today, with the rapid development of digitalization and informatization, text information is everywhere, whether in social media interactions or news report releases. However, in the face of a vast amount of data, how to efficiently utilize and accurately understand text information has become a major challenge. Against this background, discourse structure analysis aims to uncover the semantic connections and hierarchical organization among various clauses, sentences, and even paragraphs in a text. Its core tasks include: identifying elementary discourse units and constructing a discourse structure tree, etc. For example, for the given text "Technology is constantly evolving, intelligent devices have penetrated people's lives, and they are changing our way of working", the discourse structure analysis method first needs to identify three elementary discourse units (EDUs): EDU1 (Technology is constantly evolving,), EDU2 (Intelligent devices have penetrated people's lives,), and EDU3 (They are changing our way of working.). Subsequently, the model needs to further analyze the relationships between these units and construct the corresponding discourse structure tree ((EDU1, EDU2), EDU3). Improving the performance of discourse structure analysis can indirectly enhance the effects of various high-level natural language processing tasks such as sentiment analysis, machine translation, dialogue systems, and abstract generation. In recent years, numerous researchers have conducted in-depth research on discourse structure analysis and made certain progress, and their results have been effectively verified in multiple practical application scenarios such as sentiment analysis.
[0003] Currently, existing deep learning-based discourse structure analysis methods are mainly divided into two categories: 1) Bottom-up methods: This method starts from elementary discourse units and merges two units each time to gradually construct the entire discourse structure tree. It can be divided into transition-based models and graph-based models. The former focuses on what operation should be performed at each step, while the latter focuses on how to calculate the merging probability of two adjacent units. 2) Top-down methods: This method starts from the entire text and divides it into two parts layer by layer, and continuously repeats the division until it can no longer be divided. This type of method is closer to the human way of text understanding and is currently actively researched. However, limited by the small scale and limited domain coverage of manually annotated discourse data, these methods often face the problem of insufficient generalization ability in practical applications. Due to the limitations in scale and domain coverage of such corpora, the models trained based on them show a significant decline in performance when processing non-news texts, and the generalization ability is relatively limited. This problem is particularly prominent in practical applications and restricts the popularization and application of discourse structure analysis technology in diverse text scenarios.
[0004] To improve the generalization performance, researchers have tried to introduce various data augmentation methods, including: integrating various passage-related corpora using a multi-task learning framework, cross-lingual modeling using annotated data in different languages, and pre-training the system through passage-related tasks. Although these methods have improved the performance to a certain extent, there are still two major problems: 1) The current system is difficult to effectively utilize corpora with large domain differences; 2) The scale and domain breadth of available passage data are still insufficient. These problems limit the effectiveness of existing data augmentation techniques in enhancing the system's generalization ability. Summary of the Invention
[0005] Therefore, an embodiment of the present invention proposes a discourse structure analysis system based on a multi-expert mechanism to improve the system's generalization ability.
[0006] A discourse structure analysis system based on a multi-expert mechanism according to an embodiment of the present invention includes a character-level encoding module, an EDU-level encoding module based on a multi-expert mechanism, a decoding context learning module, a decoding module based on a multi-expert pointer network, a pre-training module based on a large model-generated corpus, and a fine-tuning module based on an artificially annotated discourse corpus;
[0007] The character-level encoding module learns the initial semantic representation of basic discourse units based on a pre-trained language model, and the basic discourse units are obtained by segmenting the input text using a basic discourse unit recognition tool;
[0008] The EDU-level encoding module based on a multi-expert mechanism takes the initial semantic representation of basic discourse units as input, and uses a Transformer layer integrating a multi-expert mechanism to learn the final semantics of basic discourse units. The Transformer layer integrating a multi-expert mechanism includes at least a multi-expert mechanism layer. In the multi-expert mechanism layer, the weights corresponding to all domain-private experts are calculated, and then the domain-private expert with the highest weight is selected. Then, based on the shared expert and the selected k domain-private experts, the context representation of basic discourse units is calculated, and finally, the final semantic representation of basic discourse units is calculated with the context representation of basic discourse units as input; k The decoding context learning module takes the final semantic representation of basic discourse units as input, calculates the semantic representation of the text segment to be segmented, and then learns the decoding context representation based on a long short-term memory network;
[0009] The decoding module based on a multi-expert pointer network takes the decoding context representation as input, uses the shared pointer network and domain-private pointer networks in the multi-expert pointer network to predict the optimal segmentation position of the text segment to be segmented, and uses the optimal segmentation position to segment the text segment to be segmented to construct a complete discourse structure tree;
[0010] The decoding module based on a multi-expert pointer network takes the decoding context representation as input, uses the shared pointer network and domain-private pointer networks in the multi-expert pointer network to predict the optimal segmentation position of the text segment to be segmented, and uses the optimal segmentation position to segment the text segment to be segmented to construct a complete discourse structure tree;
[0011] The pre-training module based on large model generated corpus uses a large language model to generate automatically annotated passage corpora in multiple fields, and pre-trains the passage structure analysis system to obtain a pre-trained passage structure analysis system;
[0012] The fine-tuning module based on manually annotated passage corpus uses the manually annotated passage corpus to train the pre-trained passage structure analysis system in the fine-tuning stage to obtain a trained passage structure analysis system.
[0013] The passage structure analysis system based on the multi-expert mechanism according to the embodiment of the present invention is optimized from two aspects: system structure and training data. In terms of system structure, a multi-expert mechanism is introduced in the encoding and decoding stages, enabling the system to automatically select a matching expert for processing according to the field to which the input text belongs, thereby more effectively utilizing corpora in different fields and enhancing the adaptability to heterogeneous corpora, which can improve the generalization ability of the system; in terms of training data, a large number of passage annotation data in multiple fields are automatically generated by means of a large language model, significantly enhancing the domain diversity of the corpus, providing data support for the effective training of multiple experts, further improving the generalization ability of the system, and having good practical prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The above and / or additional aspects and advantages of the embodiments of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0015] Figure 1 is a schematic structural diagram of a passage structure analysis system based on a multi-expert mechanism according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] Please refer to Figure 1 , an embodiment of the present invention provides a passage structure analysis system based on a multi-expert mechanism, including a character-level encoding module, an EDU-level encoding module based on a multi-expert mechanism, a decoding context learning module, a decoding module based on a multi-expert pointer network, a pre-training module based on large model generated corpus, and a fine-tuning module based on manually annotated passage corpus.
[0018] The word-level encoding module learns the preliminary semantic representation of basic discourse units based on a pre-trained language model, where the basic discourse units are obtained by segmenting the input text using a basic discourse unit recognition tool.
[0019] Among them, the pre-trained language model (Pretrained Language Model) obtains general semantic and syntactic information through unsupervised or self-supervised tasks (such as masked prediction, next-word prediction) on a large-scale corpus and is used as a basic encoder, which can significantly improve the performance of many downstream tasks. Specifically, given the input text (which can be a passage or a paragraph), first use an existing basic discourse unit recognition tool to segment it into , where , , are the 1st, th, and th basic discourse units in the input text respectively. Based on the pre-trained language model, learn the preliminary semantic representations of as follows:
[0020] ;
[0021] Among them, is the preliminary semantic representation of the th basic discourse unit, is the pre-trained language model (such as Bert or Roberta), is the concatenated global placeholder, , , are the 1st, th, and th, th words in the
[0022] The EDU-level encoding module based on the multi-expert mechanism takes the preliminary semantic representation of the basic discourse unit as input and uses a Transformer layer integrating the multi-expert mechanism to learn the final semantics of the basic discourse unit. The Transformer layer integrating the multi-expert mechanism includes at least a multi-expert mechanism layer. In the multi-expert mechanism layer, calculate the weights corresponding to all domain-private experts, then select the k domain-private experts with the highest weights, and then calculate the context representation of the basic discourse unit based on the shared expert and the selected k domain-private experts. Finally, calculate the final semantic representation of the basic discourse unit with the context representation of the basic discourse unit as input.
[0023] Among them, the Mixture of Experts (MoE) is an architectural idea to improve the expressiveness and generalization ability of the model. It introduces multiple expert sub-networks, each expert is good at processing data of different types or fields, and a routing network dynamically selects the most appropriate expert or expert combination to participate in the calculation according to the input. The present invention regards each head of the multi-head attention mechanism in the classic Transformer layer as an expert, and divides all experts into two parts: shared experts and domain-private experts. The multi-expert mechanism dynamically activates some domain-private experts according to the input, so that it can more effectively utilize multi-domain chapter corpora with large differences and effectively improve the generalization performance of the model.
[0024] The EDU-level encoding module based on the multi-expert mechanism satisfies the following formula:
[0025] ;
[0026] in, , , The first and , The final semantic representation of a basic chapter unit, , The first and The preliminary semantic representation of the basic chapter unit, Represents the Transformer layer that integrates multiple expert mechanisms. The Transformer layer that integrates multiple expert mechanisms includes a multiple expert mechanism layer, a feedforward fully connected layer, a normalization layer, and a residual connection operation layer.
[0027] Specifically, the weights of all domain private experts are calculated in the multi-expert mechanism layer, and then the one with the highest weight is selected. k In the process of finding private experts in a certain field, the following equation is satisfied:
[0028] ;
[0029] ;
[0030] ;
[0031] in, , The first and second k The weights corresponding to private experts in each field, is the normalization function, , The first and second k The weight vector components corresponding to the private experts in the field, , are the indexes corresponding to and respectively, denotes a function of the largest k values and the corresponding indexes in the output vector, is the weight vector corresponding to all domain - specific experts, is the first parameter matrix to be learned.
[0032] During the process of calculating the context representation of the basic discourse unit based on the shared experts and the selected k domain - specific experts, the following formula is satisfied:
[0033] + ;
[0034] ;
[0035] ;
[0036] where is the context representation of the th basic discourse unit, is the number of shared experts, is the th attention mechanism used as a shared expert, is the th attention mechanism used as a domain - specific expert, is the query in the attention mechanism, is the key and value in the attention mechanism, is the context representation of the th basic discourse unit calculated through , is the context representation of the th basic discourse unit calculated through , is the weight corresponding to the selected th domain - specific expert. In specific implementation, the number of shared experts can be set to 2 - 3; the total number of domain - specific experts can be more, such as 20 - 30; the number of selected domain - specific experts k can be set to 2 - 3.
[0037] During the process of calculating the final semantic representation of the basic discourse unit with the context representation of the basic discourse unit as the input, the following formula is satisfied:
[0038] ;
[0039] where is the final semantic representation of the th basic passage unit, is the residual connection operation, is the normalization layer, is the feed-forward fully connected layer. The calculations of the feed-forward fully connected layer, the normalization layer, and the residual connection operation are the same as those in the classical Transformer layer, and will not be elaborated here. In specific implementation, multiple layers can be stacked to calculate the final semantic representation of the basic passage unit. By introducing a multi-expert mechanism in the encoding layer, different experts can focus on input data in different fields, avoiding forced global fitting of all data by a single model, reducing the risk of overfitting, and thus improving the generalization ability of the model.
[0040] The decoding context learning module takes the final semantic representation of the basic passage unit as input, calculates the semantic representation of the text segment to be segmented, and then learns the decoding context representation based on the long short-term memory network.
[0041] The present invention adopts a top-down decoding strategy to obtain the passage structure tree of the input text. This strategy starts from the whole text, performs binary segmentation on it layer by layer, and continues to segment each sub-part according to the principle of depth-first or breadth-first until no further segmentation is possible (i.e., this part only includes one basic passage unit), thereby constructing a complete passage structure tree.
[0042] The decoding context learning module is based on the common long short-term memory network , takes the final semantic representation of the basic passage unit as input, and learns the context representation at each segmentation moment.
[0043] Specifically, given the text segment to be segmented composed of the th to the th consecutive basic passage units , being the th basic passage unit, first calculate the semantic representation, and the formula is as follows:
[0044] ;
[0045] where is the semantic representation of the text segment to be segmented composed of the th to the th consecutive basic passage units is the sub-network that fuses multiple vector representations, which can be an attention mechanism or a simple pooling operation, etc.; is the The final semantic representation of a basic text segment.
[0046] Then, based on learning the decoding context representation at the -th moment:
[0047] ;
[0048] is the decoding context representation at the -th moment, is the long short-term memory network, is the decoding context representation at the -(th) moment.
[0049] During the decoding process, given an input text containing basic text segments, the first segment to be segmented is (i.e., the entire input text). Assuming that the text segments obtained after segmentation are and , the second segment to be segmented is . Assuming that the text segments obtained after segmentation are and ; According to the depth-first principle, the third segment to be segmented is ; And so on, until no further segmentation is possible and then backtracking. In the implementation process, segmentation can also be carried out according to the breadth-first principle. Then, the segments to be segmented in the above example are , , , and etc.
[0050] The decoding module based on the multi-expert pointer network takes the decoding context representation as input, uses the shared pointer network and domain-private pointer network in the multi-expert pointer network to predict the optimal segmentation position of the text segment to be segmented, and uses the optimal segmentation position to segment the text segment to be segmented to construct a complete text structure tree.
[0051] Among them, the decoding module based on the multi-expert pointer network predicts the optimal segmentation position of each text segment to be segmented based on the multi-expert pointer network. Similarly, multiple expert pointer networks are divided into two groups: the shared pointer network and the domain-private pointer network.
[0052] In the decoding module based on the multi-expert pointer network, for the text segment to be segmented, its candidate segmentation position is defined as , and the segmentation position is in front of , and the segmentation position is After that, other segmentation positions are between two adjacent basic discourse units.
[0053] First, calculate the semantic representation at the segmentation position , and the formula is as follows: , the formula is as follows:
[0054] ;
[0055] Among them, is the semantic representation at the segmentation position , is a multi-layer feedforward neural network or a simple vector addition operation, , are respectively th, th final semantic representations of the basic discourse units;
[0056] Then, calculate the weights corresponding to all domain-private pointer networks, and select the domain-private pointer network with the highest weight. The expression is as follows:
[0057] ;
[0058] ;
[0059] ;
[0060] Among them, , are respectively the weights corresponding to the selected 1st and q th domain-private pointer networks, , are the weight vector components corresponding to the selected 1st and q th domain-private pointer networks, , are respectively the indices corresponding to , , is the weight vector corresponding to all domain-private pointer networks, is the second parameter matrix to be learned;
[0061] Next, based on the shared pointer network and the selected q domain-private pointer networks, calculate the final weight vector of the candidate segmentation position. The expression is as follows:
[0062] + ;
[0063] ;
[0064] ;
[0065] Among them, represents the final weight vector of the candidate segmentation positions, is the number of shared pointer networks, is the th shared pointer network, is the th domain-private pointer network, is the weight vector of the candidate segmentation positions calculated through , is the weight vector of the candidate segmentation positions calculated through , is 's semantic representation at the segmentation position , is 's semantic representation at the segmentation position ; In specific implementation, the number of shared pointer networks can be set to 2 - 3; The total number of domain pointer networks can be more, such as 20 - 30; The number of selected domain-private pointer networks can be set to 2 - 3.
[0066] Finally, select the candidate segmentation position corresponding to the largest component in as the optimal segmentation position (i.e., the segmentation position at the th moment), and the expression is as follows:
[0067] ;
[0068] Among them, is the optimal segmentation position, is the function to obtain the index corresponding to the largest component in the vector.
[0069] The pre-training module for generating corpus based on large models uses large language models to generate automatically annotated passage corpora in multiple domains and pre-trains the passage structure analysis system to obtain a pre-trained passage structure analysis system.
[0070] Among them, although the corpora automatically annotated by large language models inevitably have certain noise, they have characteristics such as large quantity and wide coverage of domains and can be used for model pre-training to improve its generalization ability.
[0071] Given domain description and manually annotated passage corpus , using the imitation function of the large model, generate automatically annotated discourse corpora in a given domain, specifically as follows:
[0072] ;
[0073] Among them, is a large language model, is a function for constructing prompt text, is the text describing the imitation task, are multiple examples, is the th sentence with manually annotated discourse structure in , are respectively the 1st and th sentences with automatically annotated discourse structure in the given domain obtained by imitation. Given the manually annotated discourse corpus and multiple domain descriptions, multiple automatically annotated discourse corpora can be obtained according to the above method. It should be noted that directly giving the whole article to the large model for imitation results in low-quality automatically annotated corpora. Therefore, in the present invention, sentences containing multiple basic discourse units are used as the basic unit for imitation.
[0074] Given the automatically annotated discourse corpus sets in multiple domains , define the pre-training cost as:
[0075] ;
[0076] Among them, is the pre-training cost, is the given automatically annotated discourse corpus set in multiple domains, is a corpus in is the automatically annotated sentence in is the th segmentation position automatically annotated in is the cross-entropy cost function. During the implementation of the invention, after pre-training for a certain number of rounds (for example, 20) based on the corpus set , it can be ended to obtain the pre-trained discourse structure analysis system.
[0077] The fine-tuning module based on the manually annotated discourse corpus uses the manually annotated discourse corpus to train the pre-trained discourse structure analysis system in the fine-tuning stage to obtain the trained discourse structure analysis system.
[0078] Among them, define the training cost in the fine-tuning stage as:
[0079] ;
[0080] Among them, is the training cost in the fine-tuning stage, is the given manually annotated passage corpus, is the passage in the th predicted segmentation position, is the passage in the th manually annotated segmentation position.
[0081] By minimizing until convergence, that is, the training in the fine-tuning stage is completed, and a trained passage structure analysis system is obtained.
[0082] In summary, the passage structure analysis system based on the multi-expert mechanism according to this embodiment is optimized from two aspects: system structure and training data. In terms of system structure, a multi-expert mechanism is introduced in the encoding and decoding stages, enabling the system to automatically select a matching expert for processing according to the field to which the input text belongs, thereby more effectively utilizing corpora in different fields and enhancing the adaptability to heterogeneous corpora, and improving the generalization ability of the system. In terms of training data, a large number of passage annotation data in multiple fields are automatically generated with the help of a large language model, significantly enhancing the domain diversity of the corpus, providing data support for the effective training of multiple experts, further improving the generalization ability of the system, and having good practical prospects.
[0083] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0084] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A discourse structure analysis system based on a multi-expert mechanism, characterized in that, It includes a character-level encoding module, an EDU-level encoding module based on a multi-expert mechanism, a decoding context learning module, a decoding module based on a multi-expert pointer network, a pre-training module based on a large model-generated corpus, and a fine-tuning module based on an artificially annotated text corpus; The character-level encoding module learns the initial semantic representation of basic text units based on a pre-trained language model, where the basic text units are obtained by segmenting the input text using a basic text unit recognition tool; The EDU-level encoding module based on the multi-expert mechanism takes the initial semantic representation of the basic discourse unit as input, and uses the Transformer layer integrating the multi-expert mechanism to learn the final semantics of the basic discourse unit. The Transformer layer integrating the multi-expert mechanism includes at least a multi-expert mechanism layer. In the multi-expert mechanism layer, the weights corresponding to all domain-specific experts are calculated, and then the k domain-specific expert with the highest weight is selected. Then, based on the shared expert and the selected k domain-specific experts, the context representation of the basic discourse unit is calculated. Finally, the final semantic representation of the basic discourse unit is calculated with the context representation of the basic discourse unit as input; The decoding context learning module takes the final semantic representation of the basic text units as input, calculates the semantic representation of the text segment to be segmented, and then learns the decoding context representation based on a long short-term memory network; The decoding module based on the multi-expert pointer network takes the decoding context representation as input, uses the shared pointer network and the domain-private pointer network in the multi-expert pointer network to predict the optimal segmentation position of the text segment to be segmented, and uses the optimal segmentation position to segment the text segment to be segmented to construct a complete text structure tree; The pre-training module based on the large model-generated corpus uses a large language model to generate automatically annotated text corpora in multiple domains, and pre-trains the text structure analysis system to obtain a pre-trained text structure analysis system; The fine-tuning module based on the artificially annotated text corpus uses the artificially annotated text corpus to train the pre-trained text structure analysis system in the fine-tuning stage to obtain a trained text structure analysis system.
2. The discourse structure analysis system based on a multi-expert mechanism according to claim 1, wherein The character-level encoding module satisfies the following formula: ; Among them, is the preliminary semantic representation of the th basic passage unit, is the pre-trained language model, is the spliced global placeholder, , , are respectively the 1st, th, and th, and th words in the th basic passage unit.
3. The discourse structure analysis system based on a multi-expert mechanism according to claim 2, characterized in that The EDU-level encoding module based on the multi-expert mechanism satisfies the following formula: ; Among them, , , are respectively the final semantic representations of the 1st, th, th basic discourse units, , are respectively the preliminary semantic representations of the 1st, th basic discourse units, represents the Transformer layer that fuses the multi-expert mechanism.
4. The discourse structure analysis system based on the multi-expert mechanism according to claim 3, wherein Calculate the weights corresponding to all domain-specific private experts in the multi-expert mechanism layer, and then select the k domain-specific private experts with the highest weights. The following equation is satisfied during this process: ; ; ; Among them, and are the weights corresponding to the first and the k th domain private experts selected respectively, is the normalization function, and are the weight vector components corresponding to the first and the k th domain private experts selected respectively, and are the indexes corresponding to and respectively, represents the function of the largest k values and the corresponding indexes in the output vector, is the weight vector corresponding to all domain private experts, is the first parameter matrix to be learned.
5. The discourse structure analysis system based on a multi-expert mechanism according to claim 4, wherein Based on shared experts and selected k private experts in each domain, during the process of calculating the context representation of the basic text units, the following formula is satisfied: + ; ; ; Among them, is the context representation of the th basic passage unit, is the number of shared experts, is the th attention mechanism used as a shared expert, is the th attention mechanism used as a domain - private expert, is the context representation of the th basic passage unit calculated through , is the context representation of the th basic passage unit calculated through , is the weight corresponding to the th selected domain - private expert.
6. The discourse structure analysis system based on a multi-expert mechanism according to claim 5, wherein In the process of calculating the final semantic representation of the basic text units with the context representation of the basic text units as input, the following formula is satisfied: ; Among them, is the final semantic representation of the th basic passage unit, is the residual connection operation, is the normalization layer, is the feed-forward fully connected layer.
7. The discourse structure analysis system based on a multi-expert mechanism according to claim 6, characterized in that The decoding context learning module satisfies the following formula: ; ; Among them, is the text segment to be segmented composed of the th to the th consecutive basic text units, is a sub-network that fuses multiple vector representations, is the final semantic representation of the th basic text unit, is the decoding context representation at the th moment, is a long short-term memory network, is the decoding context representation at the -1 th moment.
8. The discourse structure analysis system based on a multi-expert mechanism according to claim 7, characterized in that In the decoding module based on the multi-expert pointer network, for the text segment to be segmented , first calculate the semantic representation at the segmentation position . The formula is as follows: ; Among them, is the segmentation position and the semantic representation at this position, is a multi-layer feedforward neural network, and are respectively the final semantic representations of the Then, calculate the weights corresponding to all domain private pointer networks, and select the domain private pointer network with the highest weight. The expression is as follows: ones. The expression is as follows: ; ; ; Among them, , are respectively the weights corresponding to the first and the q -th domain private pointer networks selected, , are the weight vector components corresponding to the first and the q -th domain private pointer networks selected, , are respectively the indexes corresponding to , , is the weight vector corresponding to all domain private pointer networks, is the second parameter matrix to be learned; Next, based on the shared pointer network and the selected q final weight vectors of the candidate segmentation positions are calculated by the domain private pointer network, and the expression is as follows: + ; ; ; Among them, represents the final weight vector of the candidate segmentation position, is the number of shared pointer networks, is the th shared pointer network, is the th domain-private pointer network, is the weight vector of the candidate segmentation position obtained through calculation, is the weight vector of the candidate segmentation position obtained through calculation, is the semantic representation at the segmentation position of, is the semantic representation at the segmentation position of; Finally, select the candidate segmentation position corresponding to the largest component as the optimal segmentation position, and the expression is as follows: ; Among them, is the optimal segmentation position, is a function that obtains the index corresponding to the maximum component in the vector.
9. The discourse structure analysis system based on a multi-expert mechanism according to claim 8, characterized in that The pre-training module based on the large model-generated corpus satisfies the following formula: ; Among them, is the pre-training cost, is the corpus of texts in multiple fields with given automatic annotation, is a corpus in is the automatically annotated sentence in is the th segmentation position automatically annotated in is the cross-entropy cost function.
10. The discourse structure analysis system based on a multi-expert mechanism according to claim 9, characterized in that The fine-tuning module based on the artificially annotated text corpus satisfies the following formula: ; Among them, is the training cost in the fine-tuning stage, is the given manually annotated passage corpus, is the passage the th predicted segmentation position in, is the passage the th manually annotated segmentation position in.
Citation Information
Patent Citations
Multi-feature bidirectional gating field expert entity extraction method and system
CN112101028A
NER-oriented Chinese clinical text data enhancement method and device
CN114861600A