A text processing method, system, storage medium and terminal device
By pre-training the cross-language summary model, using feature extraction and coding modules, and combining multi-task post-coding modules, the problem of poor cross-language summary effect in the existing technology is solved, and more accurate cross-language summary information extraction is achieved.
Patent Information
- Application Number
- CN202110902041.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-08-06
AI Technical Summary
Existing cross-language summary technology is poor in text processing, resulting in poor user experience.
The pre-trained cross-language summary model is adopted to extract text features through the feature extraction module and the feature encoding module, and the three branch post-coding modules are used to process the translation information, single-language summary information and cross-language summary information respectively. Finally, a more accurate cross-language summary model is trained by integrating this information.
Improve the accuracy of cross-language summary information, especially when there are fewer cross-language summary data sets, ensuring the accuracy and effectiveness of extracting cross-language summary information.
Smart Images

Figure CN114328805B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information processing based on artificial intelligence, and particularly relates to a text processing method, system, storage medium and terminal device. Background Art
[0002] Cross-lingual summarization technology is a task of summarizing the core information of the source language text and organizing it into a summary in the form of the target language. The research on cross-lingual summarization technology is of great significance for application scenarios such as cross-border e-commerce (assisting users in making decisions), public opinion analysis (helping analysts filter redundant information), and content recommendation (recommending foreign language news to users).
[0003] However, the effect of the cross-lingual summary information obtained during the existing text processing is relatively poor. If the obtained cross-lingual summary information is further applied to multiple scenarios, it will cause a bad user experience. Summary of the Invention
[0004] Embodiments of the present invention provide a text processing method, system, storage medium and terminal device, which realize cross-lingual extraction of a summary by using a relatively accurate cross-lingual summary model.
[0005] On the one hand, an embodiment of the present invention provides a text processing method, including:
[0006] Obtain a target object and call a pre-trained cross-lingual summary model;
[0007] Extract cross-lingual summary information of the target object through the cross-lingual summary model;
[0008] Wherein, the cross-lingual summary model is pre-trained through the following steps:
[0009] Determine an initial training model, where the initial training model includes a feature extraction module, a feature encoding module, and an encoded post-processing module with three branches. The feature extraction module is used to extract feature information of a sample object, the feature encoding module is used to encode the feature information of the sample object to obtain encoded features, the encoded post-processing module of the first branch in the three branches is used to determine translation information of the sample object according to the encoded features, the encoded post-processing module of the second branch in the three branches is used to determine monolingual summary information of the sample object according to the encoded features, and the encoded post-processing module of the third branch in the three branches is used to determine cross-lingual summary information of the sample object according to the encoded features;
[0010] Determine training samples, where the training samples include multiple first sample objects and their translation annotations, multiple second sample objects and their monolingual summary annotations, and multiple third sample objects and their cross-lingual summary annotations;
[0011] Train the cross - language summarization model according to the initial training model and training samples.
[0012] Another aspect of the embodiments of the present invention provides a text processing system, including:
[0013] A calling unit, configured to obtain a target object and call a pre - trained cross - language summarization model;
[0014] A summary extraction unit, configured to extract cross - language summary information of the target object through the cross - language summarization model;
[0015] The text processing system further includes:
[0016] A training unit, configured to determine an initial training model. The initial training model includes a feature extraction module, a feature encoding module, and an encoded post - processing module with three branches. The feature extraction module is configured to extract feature information of a sample object. The feature encoding module is configured to encode the feature information of the sample object to obtain encoded features. The encoded post - processing module of the first branch in the three branches is configured to determine translation information of the sample object according to the encoded features. The encoded post - processing module of the second branch in the three branches is configured to determine monolingual summary information of the sample object according to the encoded features. The encoded post - processing module of the third branch in the three branches is configured to determine cross - language summary information of the sample object according to the encoded features; determine training samples, where the training samples include multiple first sample objects and their translation annotations, multiple second sample objects and their monolingual summary annotations, and multiple third sample objects and their cross - language summary annotations; train the cross - language summarization model according to the initial training model and training samples.
[0017] Another aspect of the embodiments of the present invention further provides a computer - readable storage medium, which stores multiple computer programs, and the computer programs are adapted to be loaded and executed by a processor to perform the text processing method as described in one aspect of the embodiments of the present invention.
[0018] Another aspect of the embodiments of the present invention further provides a terminal device, including a processor and a memory;
[0019] The memory is configured to store multiple computer programs, and the computer programs are used to be loaded and executed by the processor to perform the text processing method as described in one aspect of the embodiments of the present invention; the processor is configured to implement each of the multiple computer programs.
[0020] It can be seen that in the method of this embodiment, the text processing system uses a pre-trained cross-lingual summarization model to extract cross-lingual summary information of the target object. Among them, in the initial training model determined during the process of pre-training the cross-lingual summarization model, there are three branches of encoded post-processing modules, corresponding to three different tasks, namely, determining translation information, monolingual summary information, and cross-lingual summary information. These three tasks share the same feature encoding module and the same feature extraction module. Since the task of determining cross-lingual summary information can be the integration of the two subtasks of determining translation information and monolingual summary information, the information for implementing these three tasks is utilized during the process of training the cross-lingual summarization model, and the overall task of determining cross-lingual summary information and its included subtasks are taken into account. In this way, even when the cross-lingual summary dataset (i.e., the third sample object in the training samples) is relatively small, the trained cross-lingual summarization model is also relatively accurate in extracting cross-lingual summary information. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 It is a schematic diagram of a text processing method provided by an embodiment of the present invention;
[0023] Figure 2 It is a flowchart of a text processing method provided by an embodiment of the present invention;
[0024] Figure 3 It is a schematic diagram of the initial training model determined in an embodiment of the present invention;
[0025] Figure 4 It is a flowchart of a method for training a cross-lingual summarization model in an embodiment of the present invention;
[0026] Figure 5 It is a flowchart of a method for training a cross-lingual summarization model in an application embodiment of the present invention;
[0027] Figure 6 It is a schematic diagram of the initial training model determined in an application embodiment of the present invention;
[0028] Figure 7 It is a schematic diagram of the distributed system to which the text processing method in another application embodiment of the present invention is applied;
[0029] Figure 8It is a schematic diagram of the block structure in another application embodiment of the present invention;
[0030] Figure 9 It is a schematic diagram of the logical structure of a text processing system provided by an embodiment of the present invention;
[0031] Figure 10 It is a schematic diagram of the logical structure of a terminal device provided by an embodiment of the present invention. Detailed implementation manners
[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0033] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0034] An embodiment of the present invention provides a text processing method, which is mainly applied to a text processing system to extract cross-language summary information of a target object. For example, Figure 1 As shown, the text processing system can extract cross-language summary information through the following steps:
[0035] Obtain a target object and call a pre-trained cross-language summary model;
[0036] Extract the cross-language summary information of the target object through the cross-language summary model;
[0037] Among them, the cross-language summary model is pre-trained through the following steps:
[0038] Determine an initial training model, where the initial training model includes a feature extraction module, a feature encoding module, and post - processing encoding modules for three branches. The feature extraction module is used to extract feature information of a sample object, the feature encoding module is used to encode the feature information of the sample object to obtain encoded features, the post - processing encoding module of the first branch among the three branches is used to determine translation information of the sample object according to the encoded features, the post - processing encoding module of the second branch among the three branches is used to determine monolingual summary information of the sample object according to the encoded features, and the post - processing encoding module of the third branch among the three branches is used to determine cross - language summary information of the sample object according to the encoded features;
[0039] Determine training samples, where the training samples include multiple first sample objects and their translation annotations, multiple second sample objects and their monolingual summary annotations, and multiple third sample objects and their cross - language summary annotations;
[0040] Train the cross - language summary model according to the initial training model and the training samples.
[0041] In practical applications, the text processing system can be applied to application terminals or servers, such as servers for data recommendation, dialogue summary servers, or search servers, etc.
[0042] It should be noted that the above - mentioned cross - language summary model is a machine - learning model based on artificial intelligence. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning, and decision - making.
[0043] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware - level technologies and software - level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0044] Machine learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0045] In this way, the initial training model determined during the process of pre-training the cross-lingual summarization model includes post-processing modules for encoding in three branches, corresponding to three different tasks, namely determining translation information, monolingual summarization information, and cross-lingual summarization information. These three tasks share the same feature encoding module and the same feature extraction module. Since the task of determining cross-lingual summarization information can be the integration of the two subtasks of determining translation information and monolingual summarization information, the information for implementing these three tasks is utilized during the process of training the cross-lingual summarization model, and the overall task of determining cross-lingual summarization information and its included subtasks are taken into account. Thus, even when the dataset for cross-lingual summarization (i.e., the third sample object in the training samples) is relatively small, the trained cross-lingual summarization model can extract cross-lingual summarization information more accurately.
[0046] An embodiment of the present invention provides a text processing method, mainly a method executed by a text processing system. The flowchart is as Figure 2 shown, including:
[0047] Step 101, obtain a target object and call a pre-trained cross-lingual summarization model.
[0048] It can be understood that in one case, the text processing system provides a user interface, so that the user can input the target object through the user interface, thereby initiating the cross-lingual summarization process of this embodiment. In another case, the text processing system can actively use a text information as the target object and initiate the cross-lingual summarization process of this embodiment. Among them, the target object can refer to text information in a certain language form.
[0049] After the text processing system obtains the target object, it will call the cross-lingual summarization model pre-set in the system. This cross-lingual summarization model is a machine learning model, which can be trained by a certain method, and its operation logic is set in the text processing system in advance.
[0050] Step 102, extract the cross-lingual summarization information of the target object through the cross-lingual summarization model.
[0051] Among them, before executing the above steps 101 and 102, the text processing system can pre-train a cross-lingual summarization model through the following steps:
[0052] Step 201, determine an initial training model.
[0053] It can be understood that when the text processing system determines the initial training model, it will determine the multi-layer structure included in the initial training model and the initial values of the parameters in each layer mechanism. Among them, the parameters of the initial training model refer to the fixed parameters used in the calculation process of each layer structure in the initial training model, which do not need to be assigned values at any time, such as parameters such as parameter scale, number of network layers, and user vector length.
[0054] The structure of the initial training model can be as Figure 3 shown. Specifically, the initial training model includes: a feature extraction module, a feature encoding module, and a post-processing module for encoding with three branches. The feature extraction module is used to extract the feature information of the sample object; the feature encoding module is used to encode the feature information of the sample object to obtain encoded features; the post-processing module for encoding in the first branch of the three branches is used to determine the translation information of the sample object according to the encoded features, the post-processing module for encoding in the second branch of the three branches is used to determine the monolingual summarization information of the sample object according to the encoded features, and the post-processing module for encoding in the third branch of the three branches is used to determine the cross-lingual summarization information of the sample object according to the encoded features.
[0055] Among them, the translation information refers to the text information in another language corresponding to the text information in one language in the sample object. The monolingual summarization information refers to the summarization information in the same language corresponding to the text information in one language in the sample object. The cross-lingual summarization information refers to the summarization information in another language corresponding to the text information in one language in the sample object. For example, if the sample object is text information in Chinese, the translation information can be the corresponding English text information, the monolingual summarization information is the corresponding Chinese summarization information, and the cross-lingual summarization information can be the corresponding English summarization information.
[0056] Furthermore, as Figure 3 shown, the initial training model can also include attention models corresponding to the three branches respectively. Specifically:
[0057] (1) If the translation information of the sample object includes multiple translation words, these multiple translation words are determined sequentially in the determination process. In this embodiment, when determining any one of the multiple translation words, it is necessary to combine the features of the already determined translation words to determine. In this way, in the process of determining multiple translation words, the relationship between adjacent translation words, that is, the context information of the translation words, is considered, so that the finally determined translation information is more accurate.
[0058] The attention module of the first branch is used to encode the determined translation words based on the attention mechanism to obtain the first post-encoded historical features, and the post-encoding processing module of the first branch is used to determine another translation word after the determined translation words according to the encoded features and the first post-encoded historical features.
[0059] (2) If the monolingual summary information of the sample object includes multiple monolingual words, these multiple monolingual words are determined sequentially during the determination process. In this embodiment, when determining any one of the multiple monolingual words, it is necessary to combine the features of the determined monolingual words for determination. In this way, during the determination process of the multiple monolingual words, the relationship between adjacent monolingual words, that is, the context information of the monolingual words, is considered, making the finally determined monolingual summary information more accurate.
[0060] The attention module of the second branch is used to encode the determined monolingual words based on the attention mechanism to obtain the second post-encoded historical features, and the post-encoding processing module of the second branch is used to determine another monolingual word after the determined monolingual words according to the encoded features and the second post-encoded historical features.
[0061] (3) If the cross-lingual summary information of the sample object includes multiple cross-lingual words, these multiple cross-lingual words are determined sequentially during the determination process. In this embodiment, when determining any one of the multiple cross-lingual words, it is necessary to combine the features of the determined cross-lingual words for determination. In this way, during the determination process of the multiple cross-lingual words, the relationship between adjacent cross-lingual words, that is, the context information of the cross-lingual words, is considered, making the finally determined cross-lingual summary information more accurate.
[0062] During the determination process of these cross-lingual words, the attention module of the third branch is used to encode the determined cross-lingual words based on the attention mechanism to obtain the third post-encoded historical features, and the post-encoding processing module of the third branch is used to determine another cross-lingual word after the determined cross-lingual words according to the encoded features and the third post-encoded historical features.
[0063] Among them, when the attention modules corresponding to the three branches perform encoding based on the attention mechanism, they can perform encoding based on the single-head self-attention mechanism (Self-Attention) or the multi-head self-attention mechanism (MultiHead). Among them, the attention mechanism mainly processes the target object through the attention function, and the essence of the attention function is a mapping from a query (Q) to a series of key (K)-value (V) pairs. Through the attention function, the attention features of the target object can be obtained, which are used to represent the relatively important and notable features in the target object. Specifically, when obtaining the above-mentioned first post-historical encoding features, the input of the attention function used is specifically the features of the determined translation words, and the obtained first post-historical encoding features are mainly used to represent the features of the relatively important translation words in the determined translation words; when obtaining the above-mentioned second post-historical encoding features, the input of the attention function used is specifically the features of the determined monolingual words, and the obtained second post-historical encoding features are mainly used to represent the features of the relatively important monolingual words in the determined monolingual words; when obtaining the above-mentioned third post-historical encoding features, the input of the attention function used is specifically the features of the determined cross-lingual words, and the obtained third post-historical encoding features are mainly used to represent the features of the relatively important cross-lingual words in the determined cross-lingual words.
[0064] Among them, in the similar single-head self-attention mechanism, the final attention features are directly obtained based on its input (i.e., Q, K, and V); the multi-head self-attention mechanism means that after performing multiple linear transformations on its input (i.e., Q, K, and V), the corresponding attention features are respectively obtained based on the input after multiple transformations, and then the final attention features are determined by integrating the attention features obtained multiple times.
[0065] In the specific implementation process, the initial training model can specifically be a neural network in the form of a multi-task convolutional neural network (Multi-task convolutional neural network, MTCNN), such as the Output Network (ONet) of MTCNN.
[0066] Step 202: Determine the training samples, where the training samples include multiple first sample objects and their translation annotations, multiple second sample objects and their monolingual summary annotations, and multiple third sample objects and their cross-lingual summary annotations.
[0067] Step 203: Train the cross-lingual summary model according to the initial training model and the training samples.
[0068] Specifically, as Figure 4 shown, the text processing system can train the cross-lingual summary model through the following steps:
[0069] Step 2031: Use the initial training model to determine the translation information of each first sample object, the monolingual summary information of the second sample object, and the cross - language summary information of the third sample object respectively.
[0070] Step 2032: Adjust the initial training model according to the translation information, monolingual summary information, and cross - language summary information determined by the initial training model, as well as the translation annotation, monolingual summary annotation, and cross - language summary annotation of the corresponding sample objects in the training samples, to obtain the final training model.
[0071] Specifically, the text processing system will first calculate the overall loss function related to the initial training model according to the translation information, monolingual summary information, and cross - language summary information of each sample object obtained in the above Step 2031, as well as the translation annotation, monolingual summary annotation, and cross - language summary annotation of the corresponding sample objects in the training samples. This overall loss function is used to indicate the error between the translation information, monolingual summary information, and cross - language summary information of each sample object detected by the initial training model and the actual translation information, monolingual summary information, and cross - language summary information in the corresponding sample objects, such as the cross - entropy loss function, etc.; then adjust the parameter values of the parameters in the initial training model according to the calculated overall loss function.
[0072] Among them, the process of training the training model is to minimize the value of the above - mentioned error as much as possible. This training process continuously optimizes the parameter values of the parameters in the initial training model determined in the above Step 201 through a series of mathematical optimization means such as backpropagation derivative and gradient descent, and makes the calculated value of the above - mentioned overall loss function drop to the lowest.
[0073] In this embodiment, when calculating the overall loss function related to the initial training model, it mainly includes the following parts:
[0074] Calculate the first loss function related to the encoded post - processing module of the first branch according to the translation information of the first sample object determined by the initial training model and the translation annotation of the corresponding first sample object in the training samples; calculate the second loss function related to the encoded post - processing module of the second branch according to the monolingual summary information of the second sample object determined by the initial training model and the monolingual summary annotation of the corresponding second sample object in the training samples; calculate the third loss function related to the encoded post - processing module of the third branch according to the cross - language summary information of the third sample object determined by the initial training model and the cross - language summary annotation of the corresponding third sample object in the training samples; calculate the overall loss function according to the first loss function, the second loss function, and the third loss function.
[0075] Among them, when calculating the overall loss function, the first weight value and the second weight value can be determined first, and then the sum of the first product of the first weight value and the first loss function, the second product of the second weight value and the second loss function, and the third loss function is used as the overall loss function. Among them, the first weight value and the second weight value can be adjusted dynamically. Specifically, when adjusting the first weight value, the adjusted first weight value is the function calculation value among the first weight value before adjustment, the total number of training steps of the encoding post-processing module of the first branch, and the current count of the training steps; when adjusting the second weight value, the adjusted second weight value is the function calculation value among the second weight value before adjustment, the total number of training steps of the encoding post-processing module of the second branch, and the current count of the training steps.
[0076] It should be noted that the above steps 2031 to 2032 are an adjustment of the parameter values in the initial training model based on the translation information, monolingual summary information, and cross-lingual summary information determined by the initial training model. In actual applications, the above steps 2031 to 2032 need to be continuously looped until the adjustment of the parameter values meets certain stopping conditions.
[0077] Therefore, after the text processing system executes the above steps 2031 to 2032 of the embodiment, it is also necessary to determine whether the current adjustment of the parameter values meets the preset stopping conditions. When it is met, the parameter values adjusted in the above step 2032 are used as the parameter values in the finally trained training model; when it is not met, the initial training model after adjusting the parameter values is returned to execute the above steps 2031 to 2032. Among them, the preset stopping conditions include but are not limited to any one of the following conditions: the difference between the currently adjusted parameter values and the parameter values adjusted last time is less than a threshold, that is, the adjusted parameter values reach convergence; and the number of adjustments of the parameter values is equal to the preset number, etc.
[0078] Step 2033, determining that the pre-trained cross-lingual summary model may include the feature extraction module, the feature encoding module, and the encoding post-processing module of the third branch in the finally trained training model.
[0079] It can be seen that in the method of this embodiment, the text processing system uses a pre-trained cross-lingual summarization model to extract cross-lingual summary information of the target object. Among them, the initial training model determined during the process of pre-training the cross-lingual summarization model includes encoding post-processing modules with three branches, corresponding to three different tasks, namely determining translation information, monolingual summary information, and cross-lingual summary information. These three tasks share the same feature encoding module and the same feature extraction module. Since the task of determining cross-lingual summary information can be the integration of the two subtasks of determining translation information and monolingual summary information, the information for implementing these three tasks is utilized during the process of training the cross-lingual summarization model, and at the same time, the overall task of determining cross-lingual summary information and its included subtasks are considered. In this way, even when the cross-lingual summary dataset (i.e., the third sample object in the training samples) is small, the trained cross-lingual summarization model is relatively accurate in extracting cross-lingual summary information.
[0080] The following is a specific application example to illustrate the text processing method of the present invention. The method of this embodiment can include the following two parts:
[0081] (1) As Figure 5 shown, the training of the cross-lingual summarization model can be achieved through the following steps:
[0082] Step 301, determine the training samples, which can specifically include three parts. The first part of the training samples includes multiple first sample objects and their translation annotations. The second part of the training samples includes multiple second sample objects and their monolingual summary annotations. The third part of the training samples includes multiple third sample objects and their cross-lingual summary annotations.
[0083] Step 302, determine the initial training model, whose structure is as Figure 6 shown, and can include: a feature extraction module, the feature encoding module is specifically an encoder, and the encoding post-processing modules of the three branches respectively specifically include a decoder and a probability output (such as a softmax function), and the attention modules of the three branches.
[0084] Among them, the feature extraction module is used to extract the feature information of the sample objects in each part, and specifically can be embedding features (Embedding), such as token embeddings and positional embeddings.
[0085] The encoder is used to encode the embedding features output by the feature extraction module to obtain encoded features.
[0086] The attention module of the first branch is used to process the determined translation word Y mt,<tPerform encoding based on the attention mechanism to obtain the first post-historical encoding feature Specifically, it can be represented by the following formula 1-1. Among them, the attention mechanism is specifically the multi-head self-attention mechanism (in other embodiments, a single-head self-attention mechanism can be used), and its input is the determined translation word Y mt,<t The embedding feature y1 of; The decoder MT of the first branch is used to determine an interaction representation feature O according to the first post-historical encoding feature And the encoded feature obtained for the first sample object Determine an interaction representation feature O mt,t , specifically, it can be represented by the following formula 1-2. Among them, the multi-head self-attention mechanism can be used in the feedforward neural network (FFN) to determine the interaction representation feature O mt,t , where t is a certain translation word; The probability output of the first branch is used to output the probability p of the translation information (including multiple translation words) of the first sample object through the softmax function mt , specifically, it can be represented by the following formula 1-3. Among them, W mt And b mt Are parameters to be learned and need to determine the corresponding initial values first. X mt Represents the embedding feature of the first sample object, and Y mt,t Represents the determined translation word Y mt,<t The probability of another translation word afterwards.
[0087]
[0088]
[0089]
[0090] The attention module of the second branch is used to perform encoding based on the attention mechanism on the determined monolingual word Y ms,<t To obtain the second post-historical encoding feature Specifically, it can be represented by the following formula 2-1. Among them, the attention mechanism is specifically the multi-head self-attention mechanism, and its input is the determined monolingual word Y ms,<t The embedding feature y2 of; The decoder MS of the second branch is used to determine an interaction representation feature O according to the second post-historical encoding feature And the encoded feature obtained for the second sample object Determine an interaction representation feature O ms,t , specifically, it can be represented by the following formula 2-2. Among them, the multi-head self-attention mechanism can be used in the FFN network to determine the interaction representation feature O ms,t, where t is a single-language word; the probability output of the second branch is used to output the probability p of the single-language summary information (including multiple single-language words) of the second sample object through the softmax function mt , which can be specifically represented by the following formula 2-3, where W ms and b ms are parameters to be learned and their corresponding initial values need to be determined first. X ms represents the embedding feature of the second sample object, and Y ms,t represents the determined single-language word Y ms,<t and the probability of another single-language word after that.
[0091]
[0092]
[0093]
[0094] The attention module of the third branch is used to perform encoding based on the attention mechanism on the determined cross-language word Y cls,<t to obtain the feature of the third historical encoding Specifically, it can be represented by the following formula 3-1, where the attention mechanism is specifically the multi-head self-attention mechanism, and its input is the embedding feature y3 of the determined cross-language word Y cls,<t The decoder CLS of the third branch is used to determine an interactive representation feature O based on the feature of the third historical encoding and the encoded feature obtained for the third sample object cls,t Specifically, it can be represented by the following formula 3-2, where the multi-head self-attention mechanism can be used in the FFN network to determine the interactive representation feature O cls,t , where t is a cross-language word; the probability output of the third branch is used to output the probability p of the cross-language summary information (including multiple cross-language words) of the third sample object through the softmax function cls , which can be specifically represented by the following formula 3-3, where W cls and b cls are parameters to be learned and their corresponding initial values need to be determined first. X cls represents the embedding feature of the third sample object, and Y cls,t represents the determined cross-language word Y cls,<t and the probability of another cross-language word after that.
[0095]
[0096]
[0097]
[0098] Step 303: Determine the translation information of the first sample object, the monolingual summary information of the second sample object, and the cross-lingual summary information of the third sample object respectively through the initial training model.
[0099] Step 304: Calculate the first loss function according to the translation information of the first sample object determined by the initial training model and the translation annotation in the training sample, which can be specifically represented by the following formula 4-1; calculate the second loss function according to the monolingual summary information of the second sample object determined by the initial training model and the monolingual summary annotation in the training sample, which can be specifically represented by the following formula 4-2; calculate the third loss function according to the cross-lingual summary information of the third sample object determined by the initial training model and the cross-lingual summary annotation in the training sample, which can be specifically represented by the following formula 4-3.
[0100]
[0101]
[0102]
[0103] Step 305: Calculate the overall loss function related to the initial training model according to the first loss function L MT 、the second loss function L MS and the third loss function L CLS Specifically, it can be represented by the following formula 5:
[0104] L = L CLS + αL MT + βL MS (5)
[0105] Where α and β are the first weight value and the second weight value respectively, and can be adjusted dynamically. Specifically, when adjusting the first weight value, it can be adjusted through the following formula 6, and when adjusting the second weight value, it can be adjusted through the following formula 7:
[0106] α2 = max(0, α1 - d1), d1 = α1 * t2 / T3 (6)
[0107] β2 = max(0, β1 - d2), d2 = β1 * t2 / T4 (7)
[0108] Among them, α1 and α2 are the first weight values before and after adjustment respectively, β1 and β2 are the second weight values before and after adjustment respectively, t2 is the current count of the training steps, and T3 and T4 are the total training steps of the monolingual summarization task and the cross-lingual summarization task respectively. Among them, during the training process, the training steps need to be counted starting from 1. As the count of the training steps increases, the first weight value and the second weight value will gradually decrease until the current count of the training steps is equal to the total training steps, that is, when t2 = T3, both the first weight value and the second weight value are zero. In this way, during the process of training the cross-lingual summarization model, while using the information of the translation task and the monolingual summarization task, the cross-lingual summarization task can be made to dominate, making the finally obtained cross-lingual summarization model more accurate.
[0109] Step 306: Adjust the parameter values of the parameters in the initial training model according to the overall loss function to obtain the final training model.
[0110] Step 307: Determine whether the parameter values adjusted in Step 306 meet the preset stop condition. If they meet, use the parameter values adjusted in the above Step 306 as the parameter values in the finally trained training model, and continue to execute Step 308; if they do not meet, return to execute the above Step 303.
[0111] Step 308: Determine that the pre-trained cross-lingual summarization model may include the feature extraction module, the encoder, and the decoder CLS, probability output, and attention module in the finally trained training model, and preset the trained cross-lingual summarization model into the text processing system.
[0112] (2) When the text processing system initiates the process of obtaining the cross-lingual summary information of the target object, the pre-trained cross-lingual summarization model can be called, and the cross-lingual summary information of the target object can be directly obtained through the cross-lingual summarization model.
[0113] In actual tests, first, two cross-lingual summarization models are trained by using the existing method and the method of the embodiment of the present invention respectively. One of them is a cross-lingual summarization model from Chinese to English, and the other is a cross-lingual summarization model from English to Chinese. After using the two cross-lingual summarization models to obtain the cross-lingual summary information of the corresponding sample objects respectively, the matching rate (abbreviated as the correct matching rate) of the cross-lingual summary information calculated with the actual cross-lingual summary information of the sample objects is shown in Table 1 below. It can be seen that the cross-lingual summary information obtained by the cross-lingual summarization model trained by the method of the embodiment of the present invention is more accurate:
[0114]
[0115] Table 1
[0116] It can be seen that in the process of training the cross - language summarization model in this embodiment, the implementation information of three tasks (i.e., cross - language translation task, monolingual summarization task, and cross - language summarization task) is integrated, enhancing the performance of the trained cross - language summarization model. In particular, in the scenario where the training samples for cross - language summarization annotation are scarce, the trained cross - language summarization model is also relatively accurate.
[0117] The following uses another specific application example to illustrate the text processing method of the present invention. The text processing system in the embodiment of the present invention is mainly a distributed system 100. The distributed system may include a client 300 and multiple nodes 200 (any form of computing device connected to the network, such as a server, a user terminal). The client 300 and the nodes 200 are connected in the form of network communication.
[0118] Taking the distributed system as a blockchain system as an example, refer to Figure 7 FIG. is an optional structural schematic diagram of the distributed system 100 provided by the embodiment of the present invention applied to a blockchain system. It is formed by multiple nodes 200 (any form of computing device connected to the network, such as a server, a user terminal) and a client 300. A peer - to - peer (P2P) network is formed among the nodes. The P2P protocol is an application - layer protocol running on top of the Transmission Control Protocol (TCP). In a distributed system, any machine such as a server or a terminal can join and become a node. A node includes a hardware layer, an intermediate layer, an operating system layer, and an application layer.
[0119] Refer to Figure 7 FIG. shows the functions of each node in the blockchain system. The functions involved include:
[0120] 1) Routing, a basic function of a node, used to support communication between nodes.
[0121] In addition to the routing function, a node may also have the following functions:
[0122] 2) Application, used to be deployed in the blockchain, to implement specific services according to actual business needs, record the data related to the implemented functions to form record data, carry a digital signature in the record data to indicate the source of the task data, and send the record data to other nodes in the blockchain system. When other nodes verify the source and integrity of the record data successfully, the record data is added to the temporary block.
[0123] For example, the services implemented by the application include the code for implementing the cross - language summarization function. The cross - language summarization function mainly includes:
[0124] Obtain a target object and call a pre-trained cross-lingual summarization model; extract cross-lingual summary information of the target object through the cross-lingual summarization model; wherein, the cross-lingual summarization model is pre-trained through the following steps: determine an initial training model, the initial training model includes a feature extraction module, a feature encoding module, and an encoded post-processing module with three branches. The feature extraction module is used to extract feature information of a sample object, the feature encoding module is used to encode the feature information of the sample object to obtain encoded features, and the encoded post-processing module of the first branch in the three branches is used to determine translation information of the sample object according to the encoded features. The encoded post-processing module of the second branch in the three branches is used to determine monolingual summary information of the sample object according to the encoded features. The encoded post-processing module of the third branch in the three branches is used to determine cross-lingual summary information of the sample object according to the encoded features; determine training samples, the training samples include multiple first sample objects and their translation annotations, multiple second sample objects and their monolingual summary annotations, and multiple third sample objects and their cross-lingual summary annotations; train the cross-lingual summarization model according to the initial training model and the training samples.
[0125] 3) A blockchain includes a series of blocks (Blocks) that are sequentially connected in the order of generation. Once a new block is added to the blockchain, it will not be removed again. The block records the record data submitted by nodes in the blockchain system.
[0126] See Figure 8 This is an optional schematic diagram of the block structure provided by the embodiments of the present invention. Each block includes the hash value of the transaction records stored in this block (the hash value of this block), and the hash value of the previous block. Each block is connected through the hash value to form a blockchain. In addition, the block may also include information such as the timestamp when the block is generated. A blockchain is essentially a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains relevant information for verifying the validity of its information (anti-counterfeiting) and generating the next block.
[0127] The embodiments of the present invention also provide a text processing system, and its structural schematic diagram is as Figure 9 shown, and specifically may include:
[0128] A calling unit 10, configured to obtain a target object and call a pre-trained cross-lingual summarization model;
[0129] A summary extraction unit 11, configured to extract cross-lingual summary information of the target object through the cross-lingual summarization model called by the calling unit 10;
[0130] The text processing system further includes:
[0131] A training unit 12, configured to determine an initial training model, where the initial training model includes a feature extraction module, a feature encoding module, and post-encoding processing modules for three branches. The feature extraction module is configured to extract feature information of a sample object, the feature encoding module is configured to encode the feature information of the sample object to obtain encoded features, the post-encoding processing module of the first branch in the three branches is configured to determine translation information of the sample object according to the encoded features, the post-encoding processing module of the second branch in the three branches is configured to determine monolingual summary information of the sample object according to the encoded features, and the post-encoding processing module of the third branch in the three branches is configured to determine cross-lingual summary information of the sample object according to the encoded features; determine training samples, where the training samples include a plurality of first sample objects and their translation annotations, a plurality of second sample objects and their monolingual summary annotations, and a plurality of third sample objects and their cross-lingual summary annotations; and train the cross-lingual summary model according to the initial training model and the training samples. Then, the above-mentioned calling unit 10 will call the cross-lingual summary model trained by the training unit 12.
[0132] Wherein, the initial training model determined by the training unit 12 further includes attention modules corresponding to the three branches respectively; specifically, the translation information of the sample object includes a plurality of translation words, and the attention module of the first branch is configured to perform attention mechanism-based encoding on the determined translation words to obtain a first historical encoded feature, and the post-encoding processing module of the first branch is configured to determine another translation word after the determined translation word according to the encoded features and the first historical encoded feature; the monolingual summary information of the sample object includes a plurality of monolingual words, and the attention module of the second branch is configured to perform attention mechanism-based encoding on the determined monolingual words to obtain a second historical encoded feature, and the post-encoding processing module of the second branch is configured to determine another monolingual word after the determined monolingual word according to the encoded features and the second historical encoded feature; the cross-lingual summary information of the sample object includes a plurality of cross-lingual words, and the attention module of the third branch is configured to perform attention mechanism-based encoding on the determined cross-lingual words to obtain a third historical encoded feature, and the post-encoding processing module of the third branch is configured to determine another cross-lingual word after the determined cross-lingual word according to the encoded features and the third historical encoded feature.
[0133] Wherein, the attention modules corresponding to the three branches respectively are configured to perform encoding based on a single-head self-attention mechanism or a multi-head self-attention mechanism.
[0134] Further, when training the cross - language summarization model according to the initial training model and training samples, the training unit 12 is specifically configured to respectively determine, through the initial training model, the translation information of each first sample object, the monolingual summarization information of the second sample object, and the cross - language summarization information of the third sample object; adjust the initial training model according to the translation information, monolingual summarization information, and cross - language summarization information determined by the initial training model, and the translation annotation, monolingual summarization annotation, and cross - language summarization annotation of the corresponding sample objects in the training samples, so as to obtain the final training model; determine that the pre - trained cross - language summarization model includes the feature extraction module, the feature encoding module, and the encoded post - processing module of the third branch in the final training model.
[0135] Among them, when the training unit 12 adjusts the initial training model according to the translation information, monolingual summarization information, and cross - language summarization information determined by the initial training model, and the translation annotation, monolingual summarization annotation, and cross - language summarization annotation of the corresponding sample objects in the training samples, it is specifically configured to calculate an overall loss function related to the initial training model according to the translation information, monolingual summarization information, and cross - language summarization information determined by the initial training model, and the translation annotation, monolingual summarization annotation, and cross - language summarization annotation of the corresponding sample objects in the training samples; adjust the parameter values of the parameters in the initial training model according to the overall loss function.
[0136] Among them, when the training unit 12 calculates an overall loss function related to the initial training model according to the translation information, monolingual summarization information, and cross - language summarization information determined by the initial training model, and the translation annotation, monolingual summarization annotation, and cross - language summarization annotation of the corresponding sample objects in the training samples, it is specifically configured to calculate a first loss function related to the encoded post - processing module of the first branch according to the translation information of the first sample object determined by the initial training model and the translation annotation of the corresponding first sample object in the training samples; calculate a second loss function related to the encoded post - processing module of the second branch according to the monolingual summarization information of the second sample object determined by the initial training model and the monolingual summarization annotation of the corresponding second sample object in the training samples; calculate a third loss function related to the encoded post - processing module of the third branch according to the cross - language summarization information of the third sample object determined by the initial training model and the cross - language summarization annotation of the corresponding third sample object in the training samples; calculate the overall loss function according to the first loss function, the second loss function, and the third loss function.
[0137] Among them, the training unit 12 calculates the overall loss function according to the first loss function, the second loss function and the third loss function, specifically for determining a first weight value and a second weight value; and takes the sum of the first product of the first weight value and the first loss function, the second product of the second weight value and the second loss function, and the third loss function as the overall loss function. In this case, the text processing system may further include an adjustment unit 13, configured to determine that the adjusted first weight value is a function calculation value among the first weight value before adjustment, the total number of training steps of the encoding post-processing module of the first branch, and the current count of the training steps; and determine that the adjusted second weight value is a function calculation value among the second weight value before adjustment, the number of training steps, the total number of training steps of the encoding post-processing module of the second branch, and the current count of the training steps.
[0138] Further, the training unit 12 is further configured to stop adjusting the parameter value when the number of adjustments of the parameter value in the initial training model is equal to a preset number, or when the difference between the currently adjusted parameter value and the parameter value adjusted last time is less than a threshold.
[0139] It can be seen that in the text processing system of this embodiment, the abstract extraction unit 11 uses a pre-trained cross-lingual abstract model to extract cross-lingual abstract information of the target object. Among them, in the initial training model determined during the process of pre-training the cross-lingual abstract model, there are encoding post-processing modules for three branches, corresponding to three different tasks, namely determining translation information, monolingual abstract information, and cross-lingual abstract information. And these three tasks share the same feature encoding module and the same feature extraction module. Since the task of determining cross-lingual abstract information can be the integration of the two subtasks of determining translation information and monolingual abstract information, the information for implementing these three tasks is utilized during the process of training the cross-lingual abstract model, and at the same time, the overall task of determining cross-lingual abstract information and its included subtasks are considered. In this way, even when the cross-lingual abstract data set (i.e., the third sample object in the training samples) is small, the trained cross-lingual abstract model is relatively accurate when extracting cross-lingual abstract information.
[0140] The embodiment of the present invention further provides a terminal device, and its structural schematic diagram is as Figure 10As shown, the terminal device may vary significantly due to different configurations or performances, and may include one or more central processing units (CPUs) 20 (e.g., one or more processors) and a memory 21, and one or more storage media 22 for storing application programs 221 or data 222 (e.g., one or more mass storage devices). Among them, the memory 21 and the storage media 22 may be transient storage or persistent storage. The programs stored in the storage media 22 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations for the terminal device. Further, the central processing unit 20 may be configured to communicate with the storage media 22 and execute a series of instruction operations in the storage media 22 on the terminal device.
[0141] Specifically, the application program 221 stored in the storage media 22 includes an application program for cross-language summarization, and this program may include the calling unit 10, the summary extraction unit 11, the training unit 12, and the adjustment unit 13 in the above text processing system, which will not be elaborated here. Further, the central processing unit 20 may be configured to communicate with the storage media 22 and execute a series of operations corresponding to the cross-language summarization application program stored in the storage media 22 on the terminal device.
[0142] The terminal device may further include one or more power supplies 23, one or more wired or wireless network interfaces 24, one or more input / output interfaces 25, and / or one or more operating systems 223, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM, etc.
[0143] The steps performed by the text processing system in the above method embodiments may be based on the Figure 10 structure of the shown terminal device.
[0144] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium, which stores a plurality of computer programs, and the computer programs are suitable for being loaded and executed by a processor to perform the text processing method executed by the above text processing system.
[0145] Another aspect of the embodiments of the present invention further provides a terminal device, including a processor and a memory; the memory is used to store a plurality of computer programs, and the computer programs are used to be loaded and executed by the processor to perform the text processing method executed by the above text processing system; the processor is used to implement each of the plurality of computer programs.
[0146] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the text processing method provided in the above various optional implementation manners.
[0147] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0148] The above has introduced in detail a text processing method, system, storage medium and terminal device provided by an embodiment of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A text processing method, characterized in that, it includes: obtaining a target object and calling a pre-trained cross-lingual summarization model; extracting cross-lingual summary information of the target object through the cross-lingual summarization model; wherein, the cross-lingual summarization model is pre-trained through the following steps: determining an initial training model, the initial training model includes a feature extraction module, a feature encoding module, and an encoded post-processing module with three branches. The feature extraction module is used to extract feature information of a sample object, the feature encoding module is used to encode the feature information of the sample object to obtain encoded features, the encoded post-processing module of the first branch in the three branches is used to determine translation information of the sample object according to the encoded features, the encoded post-processing module of the second branch in the three branches is used to determine monolingual summary information of the sample object according to the encoded features, and the encoded post-processing module of the third branch in the three branches is used to determine cross-lingual summary information of the sample object according to the encoded features; determining training samples, the training samples include multiple first sample objects and their translation annotations, multiple second sample objects and their monolingual summary annotations, and multiple third sample objects and their cross-lingual summary annotations; training the cross-lingual summarization model according to the initial training model and the training samples.
2. The method according to claim 1, characterized in that, the initial training model further includes attention modules respectively corresponding to the three branches; the translation information of the sample object includes multiple translation words, and the attention module of the first branch is used to perform attention mechanism-based encoding on the determined translation words to obtain a first historical encoded feature, and the encoded post-processing module of the first branch is used to determine another translation word after the determined translation word according to the encoded feature and the first historical encoded feature; the monolingual summary information of the sample object includes multiple monolingual words, and the attention module of the second branch is used to perform attention mechanism-based encoding on the determined monolingual words to obtain a second historical encoded feature, and the encoded post-processing module of the second branch is used to determine another monolingual word after the determined monolingual word according to the encoded feature and the second historical encoded feature; the cross-lingual summary information of the sample object includes multiple cross-lingual words, and the attention module of the third branch is used to perform attention mechanism-based encoding on the determined cross-lingual words to obtain a third historical encoded feature, and the encoded post-processing module of the third branch is used to determine another cross-lingual word after the determined cross-lingual word according to the encoded feature and the third historical encoded feature.
3. The method according to claim 2, characterized in that, the attention modules respectively corresponding to the three branches are used to perform encoding based on a single-head self-attention mechanism or a multi-head self-attention mechanism.
4. The method according to any one of claims 1 to 3, characterized in that, the training of the cross-lingual summarization model according to the initial training model and the training samples specifically includes: Determine the translation information of the first sample object, the monolingual summary information of the second sample object, and the cross - lingual summary information of the third sample object respectively through the initial training model; Adjust the initial training model according to the translation information, monolingual summary information, and cross - lingual summary information determined by the initial training model, and the translation annotation, monolingual summary annotation, and cross - lingual summary annotation of the corresponding sample object in the training sample, so as to obtain the final training model; Determine that the pre - trained cross - lingual summary model includes the feature extraction module, feature encoding module, and the encoded post - processing module of the third branch in the final training model.
5. The method according to claim 4, wherein, The adjusting the initial training model according to the translation information, monolingual summary information, and cross - lingual summary information determined by the initial training model, and the translation annotation, monolingual summary annotation, and cross - lingual summary annotation of the corresponding sample object in the training sample specifically includes: Calculate the overall loss function related to the initial training model according to the translation information, monolingual summary information, and cross - lingual summary information determined by the initial training model, and the translation annotation, monolingual summary annotation, and cross - lingual summary annotation of the corresponding sample object in the training sample; Adjust the parameter values of the parameters in the initial training model according to the overall loss function.
6. The method according to claim 5, wherein, The calculating the overall loss function related to the initial training model according to the translation information, monolingual summary information, and cross - lingual summary information determined by the initial training model, and the translation annotation, monolingual summary annotation, and cross - lingual summary annotation of the corresponding sample object in the training sample specifically includes: Calculate the first loss function related to the encoded post - processing module of the first branch according to the translation information of the first sample object determined by the initial training model and the translation annotation of the corresponding first sample object in the training sample; Calculate the second loss function related to the encoded post - processing module of the second branch according to the monolingual summary information of the second sample object determined by the initial training model and the monolingual summary annotation of the corresponding second sample object in the training sample; Calculate the third loss function related to the encoded post - processing module of the third branch according to the cross - lingual summary information of the third sample object determined by the initial training model and the cross - lingual summary annotation of the corresponding third sample object in the training sample; Calculate the overall loss function according to the first loss function, the second loss function, and the third loss function.
7. The method according to claim 6, wherein, The calculating the overall loss function according to the first loss function, the second loss function, and the third loss function specifically includes: Determine the first weight value and the second weight value; Take the sum of the first product of the first weight value and the first loss function, the second product of the second weight value and the second loss function, and the third loss function as the overall loss function.
8. The method according to claim 7, wherein, further includes: Determine that the adjusted first weight value is the function calculation value among the first weight value before adjustment, the total number of training steps of the encoding post-processing module of the first branch, and the current count of the training steps; Determine that the adjusted second weight value is the function calculation value among the second weight value before adjustment, the number of training steps, the total number of training steps of the encoding post-processing module of the second branch, and the current count of the training steps.
9. The method according to claim 5, wherein, when the number of adjustments to the parameter value in the initial training model is equal to a preset number, or if the difference between the currently adjusted parameter value and the parameter value adjusted last time is less than a threshold, then stop adjusting the parameter value.
10. A text processing system, wherein, comprising: a calling unit, configured to obtain a target object and call a pre-trained cross-lingual summarization model; a summarization extraction unit, configured to extract cross-lingual summary information of the target object through the cross-lingual summarization model; The text processing system further comprises: a training unit, configured to determine an initial training model, where the initial training model includes a feature extraction module, a feature encoding module, and encoding post-processing modules of three branches. The feature extraction module is configured to extract feature information of a sample object, the feature encoding module is configured to encode the feature information of the sample object to obtain encoded features, the encoding post-processing module of the first branch among the three branches is configured to determine translation information of the sample object according to the encoded features, the encoding post-processing module of the second branch among the three branches is configured to determine monolingual summary information of the sample object according to the encoded features, and the encoding post-processing module of the third branch among the three branches is configured to determine cross-lingual summary information of the sample object according to the encoded features; determine training samples, where the training samples include multiple first sample objects and their translation annotations, multiple second sample objects and their monolingual summary annotations, and multiple third sample objects and their cross-lingual summary annotations; and train the cross-lingual summarization model according to the initial training model and the training samples.
11. A computer-readable storage medium, wherein, the computer-readable storage medium stores multiple computer programs, and the computer programs are adapted to be loaded and executed by a processor to perform the text processing method according to any one of claims 1 to 9.
12. A terminal device, wherein, comprising a processor and a memory; the memory is configured to store multiple computer programs, and the computer programs are used to be loaded and executed by the processor to perform the text processing method according to any one of claims 1 to 9; the processor is configured to implement each of the multiple computer programs.
Citation Information
Patent Citations
Natural language processing model training method and device
CN110188358A
Abstraction generation method and device, electronic equipment and storage medium
CN111382261A