A method, related apparatus and device for text processing

Through the first prior network and the second prior network synchronously converting the word embedding vector as an implicit vector, and combining with the third prior network to generate the target embedding vector, the problem of error accumulation in the translation and summary generalization process in the prior art is solved, and efficient text processing is achieved.

CN115705473BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110881719.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-02
Publication Date
2025-07-18
Estimated Expiration
2041-08-02

AI Technical Summary

Technical Problem

Existing end-to-end text processing methods are prone to error accumulation during the translation and summary generalization process, resulting in high complexity and low efficiency, and it is impossible to summarize the summary after the translation is completed or translate it after the generalization is completed.

Method used

The word embedding vector is synchronized through the first prior network and the second prior network to convert the word embedding vector into an implicit vector, and the target embedding vector is generated in combination with the third prior network to realize synchronous translation and summary of the text, avoiding step-by-step processing.

Benefits of technology

Reduces the complexity of text processing, improves the efficiency of text processing, and ensures the accuracy of translation and summary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115705473B_ABST
    Figure CN115705473B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method, related device, and equipment for text processing, which are used to reduce the complexity of text processing and improve the efficiency of text processing. The method in the embodiments of the present application includes: obtaining N word embedding vectors corresponding to the text to be processed, where the text to be processed belongs to the text corresponding to the first language; generating a target embedding vector according to the N word embedding vectors; inputting the target embedding vector into a first prior network, and outputting a first hidden vector through the first prior network; inputting the target embedding vector into a second prior network, and outputting a second hidden vector through the second prior network; inputting the first hidden vector, the second hidden vector, and the target embedding vector into a third prior network, and outputting a third hidden vector through the third prior network, where the third hidden vector is the vector representation of the text summary belonging to the second language; and generating a summary text of the text to be processed according to the third hidden vector and the target embedding vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a text processing method, related devices and equipment. Background Art

[0002] With the rapid development of Internet technology and digital technology, as well as the continuous improvement of people's living needs, the development of various intelligent translations is rapidly expanding. However, there is a huge amount of text information data in various languages on the Internet, which requires users to spend a long time reading, screening or translating texts to better understand the relevant information or content in the text. The translation of text summaries can make it easier for users to quickly understand the main content of large texts. The current processing method of cross-language summaries is usually to process summaries through end-to-end training methods.

[0003] However, the end-to-end summary processing method generally requires that the text to be processed be translated first before the translated text can be summarized, which may easily result in summary errors due to inaccurate translation, or the text to be processed is summarized first before the summarized summary text is translated, and finally the summary of the text to be processed is translated, which may also result in inaccurate translated summary due to inaccurate summary. Moreover, this processing method can only summarize the summary after the translation is completed, or can only translate after the summary is completed. A series of cumbersome step-by-step processing procedures are required to complete the summary translation of the text to be processed, which may not only easily cause problems of error propagation or error accumulation, but also increase the complexity of text processing, thereby reducing the efficiency of text processing. Summary of the invention

[0004] The embodiments of the present application provide a method, related apparatus and device for text processing, which are used to realize the synchronous conversion of a first latent vector and a second latent vector of a target embedding vector through a first priori network and a second priori network, and to use the first latent vector, the second latent vector and the target embedding vector as inputs to a third priori network, so as to obtain a third latent vector indicating that the text summary belongs to a second language. Then, by decoding the third latent vector and the target embedding vector, it is possible to synchronously realize the summary and translation of the text to be processed corresponding to the target embedding vector without going through a series of cumbersome step-by-step processing procedures, thereby reducing the complexity of text processing and improving the efficiency of text processing.

[0005] In view of this, the present application provides a text processing method, comprising:

[0006] Obtain N word embedding vectors corresponding to the text to be processed, where the text to be processed belongs to the text corresponding to the first language, and N is an integer greater than 1;

[0007] Generate a target embedding vector based on N word embedding vectors;

[0008] Input the target embedding vector into a first prior network, and output a first hidden vector through the first prior network. The first hidden vector represents the vector representation of the target embedding vector belonging to the second language;

[0009] Input the target embedding vector into a second prior network, and output a second hidden vector through the second prior network. The second hidden vector is the vector representation of the text summary, and the text summary is the summary expression of the text to be processed;

[0010] Input the first hidden vector, the second hidden vector, and the target embedding vector into a third prior network, and output a third hidden vector through the third prior network. The third hidden vector is the vector representation of the text summary belonging to the second language;

[0011] Generate a summary text of the text to be processed based on the third hidden vector and the target embedding vector, where the summary text belongs to the text corresponding to the second language.

[0012] Another aspect of the present application provides a text processing device, including:

[0013] An acquisition unit, configured to acquire N word embedding vectors corresponding to the text to be processed. The text to be processed belongs to the text corresponding to the first language, and N is an integer greater than 1;

[0014] A generation unit, configured to generate a target embedding vector based on the N word embedding vectors;

[0015] A processing unit, configured to input the target embedding vector into a first prior network, and output a first hidden vector through the first prior network. The first hidden vector represents the vector representation of the target embedding vector belonging to the second language;

[0016] The processing unit is further configured to input the target embedding vector into a second prior network, and output a second hidden vector through the second prior network. The second hidden vector is the vector representation of the text summary, and the text summary is the summary expression of the text to be processed;

[0017] The processing unit is further configured to input the first hidden vector, the second hidden vector, and the target embedding vector into a third prior network, and output a third hidden vector through the third prior network. The third hidden vector is the vector representation of the text summary belonging to the second language;

[0018] The generation unit is further configured to generate a summary text of the text to be processed based on the third hidden vector and the target embedding vector, where the summary text belongs to the text corresponding to the second language.

[0019] In a possible design, in an implementation manner of another aspect of the embodiments of the present application, the generation unit may specifically be configured to:

[0020] Convert the N word embedding vectors into an N×d dimensional vector matrix, where d is an integer greater than 1;

[0021] Perform dimensionality reduction on the N×d dimensional vector matrix to obtain a target embedding vector, which is a 1×d dimensional vector.

[0022] In a possible design, in an implementation manner on the other hand of the embodiment of the present application,

[0023] The obtaining unit is further configured to obtain a sample training set, which includes a first original sample, a first target sample, a second original sample, a second target sample, a third original sample, and a third target sample. The first target sample is a translated text of the first original sample, the second target sample is an abstract text of the second original sample, and the third target sample is an abstract text after translation of the third original sample;

[0024] The processing unit is further configured to output a first predicted probability distribution through a first prior network according to the first original sample, and output a second predicted probability distribution through a second prior network according to the second original sample, and output a third predicted probability distribution through a third prior network according to the third original sample;

[0025] The processing unit is further configured to output a first true probability distribution through a first recognition network according to the first target sample, and output a second true probability distribution through a second recognition network according to the second target sample, and output a third true probability distribution through a third recognition network according to the third target sample, the first true probability distribution, and the second true probability distribution;

[0026] The processing unit is further configured to update the model parameters of the first prior network according to the divergence between the first predicted probability distribution and the first true probability distribution;

[0027] The processing unit is further configured to update the model parameters of the second prior network according to the divergence between the second predicted probability distribution and the second true probability distribution;

[0028] The processing unit is further configured to update the model parameters of the third prior network according to the divergence between the third predicted probability distribution and the third true probability distribution.

[0029] In a possible design, in an implementation manner on the other hand of the embodiment of the present application, the processing unit may specifically be configured to:

[0030] Obtain at least two word embedding vectors corresponding to the first original sample, and the first original sample belongs to the text corresponding to the first language;

[0031] Generate an embedding vector of the first original sample according to at least two word embedding vectors corresponding to the first original sample;

[0032] Input the embedding vector of the first original sample into the first prior network, and output a first predicted probability distribution through the first prior network.

[0033] In a possible design, in an implementation manner of another aspect of the embodiments of the present application, the processing unit may specifically be used for:

[0034] Obtain at least two word embedding vectors corresponding to the second original sample, where the second original sample belongs to the abstract text corresponding to the first language;

[0035] Generate an embedding vector of the second original sample according to at least two word embedding vectors corresponding to the second original sample;

[0036] Input the embedding vector of the second original sample into the second prior network, and output a second predicted probability distribution through the second prior network.

[0037] In a possible design, in an implementation manner of another aspect of the embodiments of the present application, the processing unit may specifically be used for:

[0038] Obtain at least two word embedding vectors corresponding to the third original sample, where the third original sample belongs to the text corresponding to the first language;

[0039] Generate an embedding vector of the third original sample according to at least two word embedding vectors corresponding to the third original sample;

[0040] Input the embedding vector of the third original sample into the third prior network, and output a third predicted probability distribution through the third prior network.

[0041] In a possible design, in an implementation manner of another aspect of the embodiments of the present application, the processing unit may specifically be used for:

[0042] Obtain the divergence between the first predicted probability distribution and the first true probability distribution;

[0043] Take the divergence between the first predicted probability distribution and the first true probability distribution as the first loss value;

[0044] Update the model parameters of the first prior network according to the first loss value;

[0045] The processing unit may specifically be used for:

[0046] Obtain the divergence between the second predicted probability distribution and the second true probability distribution;

[0047] Take the divergence between the second predicted probability distribution and the second true probability distribution as the second loss value;

[0048] Update the model parameters of the second prior network according to the second loss value;

[0049] The processing unit can specifically be used for:

[0050] Obtain the divergence between the third predicted probability distribution and the third true probability distribution;

[0051] Take the divergence between the third predicted probability distribution and the third true probability distribution as the third loss value;

[0052] Update the model parameters of the third prior network according to the third loss value.

[0053] On the other hand, this application provides a computer device, including: a memory, a transceiver, a processor, and a bus system;

[0054] Among them, the memory is used to store programs;

[0055] When the processor is used to execute the programs in the memory, it implements the methods in the above aspects;

[0056] The bus system is used to connect the memory and the processor to enable the memory and the processor to communicate.

[0057] On the other hand, this application provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it enables the computer to execute the methods in the above aspects.

[0058] In another aspect of this application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the network device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the network device to execute the methods provided in the above aspects.

[0059] From the above technical solutions, it can be seen that the embodiments of this application have the following advantages:

[0060] By obtaining N word embedding vectors corresponding to the text to be processed in the first language, generating a target embedding vector based on the N word embedding vectors, and then inputting the target embedding vector into a first prior network, a first hidden vector capable of representing that the target embedding vector belongs to the second language is output through the first prior network. At the same time, the target embedding vector can be input into a second prior network, and a second hidden vector capable of representing the vector representation of the text summary is output through the second prior network. Then, the first hidden vector, the second hidden vector, and the target embedding vector can be input into a third prior network, and a third hidden vector capable of representing the vector representation of the text summary belonging to the second language is output through the third prior network. Finally, a summary text of the text to be processed can be generated based on the third hidden vector and the target embedding vector. Through the above method, it is realized that through the first prior network and the second prior network, the target embedding vector belonging to the first language can be synchronously converted into a first hidden vector belonging to the second language and a second hidden vector representing the text summary. By using the first hidden vector, the second hidden vector, and the target embedding vector as the input of the third prior network, a third hidden vector capable of representing the text summary belonging to the second language can be obtained. Then, by decoding the third hidden vector and the target embedding vector, synchronous translation and summary generalization of the target embedding vector can be realized. Without going through a series of cumbersome step-by-step processing procedures, the summary generalization and translation of the text to be processed can be completed synchronously, which can reduce the complexity of text processing and thus improve the efficiency of text processing. Description of the Drawings

[0061] Figure 1 is a schematic diagram of the architecture of the text processing control system in an embodiment of the present application;

[0062] Figure 2 is a schematic diagram of an embodiment of the method for text processing in an embodiment of the present application;

[0063] Figure 3 is a schematic diagram of another embodiment of the method for text processing in an embodiment of the present application;

[0064] Figure 4 is a schematic diagram of another embodiment of the method for text processing in an embodiment of the present application;

[0065] Figure 5 is a schematic diagram of another embodiment of the method for text processing in an embodiment of the present application;

[0066] Figure 6 is a schematic diagram of another embodiment of the method for text processing in an embodiment of the present application;

[0067] Figure 7 is a schematic diagram of another embodiment of the method for text processing in an embodiment of the present application;

[0068] Figure 8 It is another schematic diagram of an embodiment of the text processing method in the embodiments of the present application;

[0069] Figure 9 It is a schematic diagram of the principle process of an embodiment of the text processing method in the embodiments of the present application;

[0070] Figure 10 It is another schematic diagram of the principle process of an embodiment of the text processing method in the embodiments of the present application;

[0071] Figure 11 It is a schematic diagram of an embodiment of the text processing device in the embodiments of the present application;

[0072] Figure 12 It is a schematic diagram of an embodiment of a computer device in the embodiments of the present application. Detailed implementation manners

[0073] The embodiments of the present application provide a text processing method, related device and equipment, which can realize the synchronous conversion of the first hidden vector and the second hidden vector of the target embedding vector through the first prior network and the second prior network, and use the first hidden vector, the second hidden vector and the target embedding vector as the input of the third prior network to obtain a third hidden vector representing that the text summary belongs to the second language. Then, by decoding the third hidden vector and the target embedding vector, it is possible to synchronously implement the summary and translation of the text to be processed corresponding to the target embedding vector without going through a series of cumbersome distributed processing processes, which can reduce the complexity of text processing and thus improve the efficiency of text processing.

[0074] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0075] It should be understood that the text processing method provided in this application can be applied to scenarios where text summarization and translation are completed through modeling. As an example, for instance, through modeling, an English paper is summarized and translated into a Chinese abstract. As another example, for instance, through modeling, a French news item is summarized and translated into an English abstract. As yet another example, for instance, through modeling, an English story is summarized and translated into a French abstract. In all of the above scenarios, in order to complete text summarization and translation, the solution provided in the prior art is to process the summary through an end-to-end training method. However, the end-to-end summary processing method generally translates the text to be processed first and then summarizes the translated text, or summarizes the text to be processed first and then translates the summarized text, and finally translates the summary of the text to be processed. This processing method can only perform summary after translation is completed, or translation after summary is completed, and requires a series of cumbersome step-by-step processing procedures to complete the summary translation of the text to be processed, increasing the complexity of text processing and thus resulting in a reduction in text processing efficiency.

[0076] To solve the above problems, this application proposes a text processing method, which is applied to Figure 1 the text processing control system shown in Figure 1 , Figure 1 which is a schematic architecture diagram of the text processing control system in an embodiment of this application. As shown in Figure 1As shown, the server obtains N word embedding vectors corresponding to the text to be processed in the first language sent by the terminal device, generates a target embedding vector based on the N word embedding vectors, and then inputs the target embedding vector into the first prior network. The first prior network outputs a first hidden vector that can be used to represent that the target embedding vector belongs to the second language. At the same time, the target embedding vector can be input into the second prior network, and the second prior network outputs a second hidden vector that can be used to represent the vector representation of the text summary. Then, the first hidden vector, the second hidden vector, and the target embedding vector can be input into the third prior network, and the third prior network outputs a third hidden vector that can be used to represent the vector representation of the text summary belonging to the second language. Finally, the summary text of the text to be processed can be generated based on the third hidden vector and the target embedding vector. Through the above method, by using the first prior network and the second prior network, the target embedding vector in the first language can be synchronously converted into the first hidden vector in the second language and the second hidden vector representing the text summary. By using the first hidden vector, the second hidden vector, and the target embedding vector as the input of the third prior network, the third hidden vector that can be used to represent the text summary belonging to the second language can be obtained. Then, by decoding the third hidden vector and the target embedding vector, the synchronous translation and summary of the target embedding vector can be realized. Without going through a series of cumbersome step-by-step processing procedures, the summary and translation of the text to be processed can be completed synchronously, which can reduce the complexity of text processing and thus improve the efficiency of text processing.

[0077] It can be understood that Figure 1 only one type of terminal device is shown. In actual scenarios, more types of terminal devices can participate in the data processing process, such as a personal computer (PC). The specific quantity and types depend on the actual scenario and are not specifically limited here. Additionally, Figure 1 one server is shown, but in actual scenarios, multiple servers can also participate. Especially in the scenario of multi-model training interaction, the number of servers depends on the actual scenario and is not specifically limited here.

[0078] It should be noted that in this embodiment, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The terminal device can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods. The terminal device and the server can be connected to form a blockchain network, which is not limited in this application.

[0079] To solve the above problems, the present application proposes a text processing method, which is generally executed by a server or a terminal device. Correspondingly, the device applied to text processing is generally set in the server or the terminal device.

[0080] It can be understood that, as the text processing method, related devices and devices disclosed in the present application, multiple servers / terminal devices can form a blockchain, and the server / terminal device is a node on the blockchain. In practical applications, data sharing between nodes may be required in the blockchain, and text data may be stored on each node.

[0081] The text processing method in the present application will be introduced below. Please refer to Figure 2 and Figure 9 , an embodiment of the text processing method in the embodiment of the present application includes:

[0082] In step S101, N word embedding vectors corresponding to the text to be processed are obtained. The text to be processed belongs to the text corresponding to the first language, and N is an integer greater than 1;

[0083] In this embodiment, when there is text to be processed that needs to be translated and summarized, the text to be processed can be obtained. It can be understood that in order to enable machine learning to better learn the features in the text, the text to be processed can be converted into a form that is easy for machine learning algorithms to utilize, that is, the obtained text to be processed can be converted into N word embedding vectors, so that the text to be processed belonging to the first language can be translated and summarized into an abstract by performing corresponding processing on the obtained N word embedding vectors.

[0084] Among them, the first language can specifically be Chinese, English, French, etc., or it can be other languages, without specific limitations here. The text to be processed belonging to the first language can specifically be a sentence, a paragraph, an article, etc., or it can be other forms of text, without specific limitations here. For example, an English biography, a Chinese note, a French town news, or an English paper, etc., without specific limitations here.

[0085] Specifically, as Figure 9 shown, after obtaining the text to be processed, this embodiment can obtain N word embedding vectors corresponding to the text to be processed. Specifically, it can first perform word segmentation on the text to be processed, and then, based on the one-hot encoding method, map each processed word or character to an array or list composed of the total number of words or characters. Among them, each character or word can be uniquely corresponding to an array or list, and the uniqueness of the array or list is represented by 1. Each text can be integrated into a sparse matrix, that is, the sparse matrix corresponding to the text to be processed.

[0086] Furthermore, to avoid situations such as an increase in the amount of text operation processing and an increase in operation complexity due to excessive text length, this embodiment can pass the sparse matrix corresponding to the text to be processed through an Embedding layer, such as a token embeddings layer and a positional embeddings layer, etc., and map the sparse matrix into a dense matrix. Specifically, through some linear transformations (such as look-up table operations), the sparse matrix can be transformed into a dense matrix. The dense matrix can be characterized by N features to represent all the words. It can be understood that in the dense matrix, although it superficially represents a one-to-one correspondence between the dense matrix and a single word, in fact, it also contains a large number of internal relationships between words, between words and even between sentences. These relationships can be characterized by the parameters learned by the Embedding layer. By performing dimensionality reduction processing on the sparse matrix, the operation complexity can be effectively reduced, thereby improving the efficiency of text processing to a certain extent.

[0087] For example, assume that there is a text to be processed "The princess is very beautiful", then the text to be processed is tokenized to obtain three words "princess / very / beautiful". Furthermore, "princess / very / beautiful" can be encoded into a sparse matrix, and then the sparse matrix can be mapped into a dense matrix through an embedding layer. For example, "gong" = [0 0 0 0 1 0 0 0 00] and "zhu" = [0 0 0 1 0 0 0 0 0 0] can be mapped into a word vector of "princess" = [1.0 0.25 1.0]. Similarly, other words or characters can be mapped to generate the dense matrix corresponding to the text "The princess is very beautiful".

[0088] Furthermore, after obtaining the dense matrix corresponding to the text to be processed, a word encoder (Encoder) can encode the dense matrix. It can be understood as accurately obtaining the vector representation corresponding to each word or character in the text to be processed based on the mapping relationship between the vocabulary and the matrix, and obtaining N word embedding vectors corresponding to the text to be processed.

[0089] In step S102, a target embedding vector is generated according to the N word embedding vectors;

[0090] In this embodiment, after obtaining the N word embedding vectors, the N word embedding vectors can be converted into a target embedding vector that can be used to represent or summarize the core semantics of the text to be processed, aiming to prepare for the subsequent acquisition of the hidden vector, i.e., the collection of hidden variables, corresponding to the text to be processed, so that the subsequent model can better perform text summarization or translation based on the learning of the hidden vector.

[0091] Specifically, as Figure 9 shown, after obtaining the N word embedding vectors, this embodiment can generate a target embedding vector based on the N word embedding vectors. Specifically, it can be to take the average of the vector dimensions of the N word embedding vectors, or take the maximum value of the vector dimensions, or other methods can also be used to obtain the target embedding vector, which is not specifically limited here.

[0092] For example, assume that a text to be processed corresponds to 10 word embedding vectors of 512 dimensions. The 10 word embedding vectors of 512 dimensions can be regarded as a matrix representation of 10*512 dimensions. Furthermore, by taking the maximum value (a total of 10 values) on each dimension of a 512-dimensional vector and then adding and fusing them, a 1*512-dimensional vector, i.e., the target embedding vector, can be obtained.

[0093] In step S103, the target embedding vector is input into the first prior network, and the first hidden vector is output through the first prior network. The first hidden vector represents the vector representation of the target embedding vector belonging to the second language.

[0094] In this embodiment, after obtaining the target embedding vector, the target embedding vector can be input into a first prior network capable of aligning the semantics in the text for learning, so as to output, through the first prior network, a vector representation capable of indicating that the target embedding vector belongs to a second language, that is, the first hidden vector.

[0095] Among them, the first prior network can specifically be manifested as a language translation model with a prior distribution added, which can be used to improve the accuracy and integrity of predicting specific information about the semantic relationships between words, between words, and even between sentences in the text to be processed. The first hidden vector refers to a latent variable that can be used to indicate that the target embedding vector belongs to a second language. It can be understood that there may be features in the target embedding vector that can directly indicate the belonging of the target embedding vector to the second language but cannot be directly observed. Therefore, a latent variable can be introduced to describe the features of the belonging of the target embedding vector to the second language, that is, the first latent variable. Among them, the second language is another language category different from the first language. For example, if the first language is Chinese, the second language can be English, Japanese, Korean, etc., or it can be other languages, which is not specifically limited here.

[0096] Specifically, as Figure 9 shown, after obtaining the target embedding vector, since the target embedding vector can be understood as a parameter representation of the internal relationships between a large number of words, between words, and even between sentences in the text to be processed, but there may still be some semantic expressions, semantic relationships, or association relationships that are difficult to observe hidden in it, and these hidden information have an obvious impact on the translation of the text to be processed, but these hidden information cannot be directly observed. Therefore, in this embodiment, the target embedding vector can be input into a first prior network capable of aligning the semantics in the text for learning, so as to obtain a vector representation capable of describing these hidden information and capable of indicating that the target embedding vector belongs to a second language, that is, the first hidden vector such as Z mt , so that the subsequent accurate translation of the text to be processed can be realized according to the first hidden vector, thereby improving the accuracy of text translation to a certain extent.

[0097] In step S104, the target embedding vector is input into a second prior network, and a second hidden vector is output through the second prior network. The second hidden vector is a vector representation of the text summary, and the text summary is a summary expression of the text to be processed;

[0098] In this embodiment, after obtaining the target embedding vector, the target embedding vector can be input into a second prior network capable of generalizing the semantics of the text for learning, so as to output, through the second prior network, a vector representation capable of indicating the text summary, that is, the second hidden vector.

[0099] Among them, the second prior network can specifically be manifested as a monolingual abstract generation model with a prior distribution added, which can be used to improve the accuracy and integrity of predicting specific information with general or summary meaning in the text. The second hidden vector can be understood as a feature that may exist in the target embedding vector and can directly represent the attribution with general or summary meaning, but cannot be directly observed. Therefore, a latent variable can be introduced to describe the attribution feature of the general or summary meaning, that is, the second latent variable.

[0100] Specifically, as Figure 9 shown, after obtaining the target embedding vector, since the target embedding vector can contain a large number of parameter representations of the internal relationships between words, between words, and even between sentences, and there may also be some hidden information that is more difficult to observe and can better reflect the semantic expression, semantic relationship, or association relationship of the text to be processed, and these hidden information can better help the model to perform generalization learning on the text to be processed, but these hidden information cannot be directly observed. Therefore, in this embodiment, the target embedding vector can be input into the second prior network that can be used to semantically generalize the text for learning, so as to obtain a vector representation that can be used to describe these hidden information and can be used to represent the text summary, that is, the second hidden vector such as Z ms , so that the subsequent text summary of the text to be processed can be accurately extracted or generalized according to the second hidden vector, thereby improving the accuracy of obtaining the text summary to a certain extent.

[0101] In step S105, the first hidden vector, the second hidden vector, and the target embedding vector are input into the third prior network, and the third hidden vector, the vector representation of the text summary belonging to the second language, is output through the third prior network;

[0102] In this embodiment, after obtaining the target embedding vector, the target embedding vector can be input into the third prior network that can be used to align and generalize the semantics in the text for learning, so as to output, through the third prior network, a vector representation that can be used to represent the text summary belonging to the second language, that is, the third hidden vector.

[0103] Among them, the third prior network can specifically be manifested as a cross-language abstract generation model with a prior distribution added, which not only has the ability to further extract the semantic relationship and semantic expression in the target embedding vector, but also has the ability to obtain the attribution feature representing the text summary. The third hidden vector is a feature used to describe the attribution with general or summary meaning, and the attribution features with general or summary meaning belong to the feature representation of the second language.

[0104] Specifically, asFigure 9 As shown in Figure 9 , after obtaining the target embedding vector, the target embedding vector can be concatenated with the first hidden vector and the second hidden vector to obtain a vector representation rich in features. Then, in this embodiment, the concatenated vector representation can be input into a third prior network capable of aligning and summarizing the semantics of the text for learning, so as to more accurately and fully obtain the features used to describe the attribution with general or summary meaning, and the vector representation of these attribution features with general or summary meaning belonging to the features of the second language, that is, the third hidden vector of the text summary belonging to the second language, such as Z. cls , so that in the subsequent process, the accurate extraction or summary of the text summary of the text to be processed can be realized according to the third hidden vector, as well as accurate translation, thereby improving the accuracy of obtaining the text summary belonging to the second language to a certain extent.

[0105] In step S106, according to the third hidden vector and the target embedding vector, a summary text of the text to be processed is generated, where the summary text belongs to the text corresponding to the second language.

[0106] In this embodiment, after obtaining the third hidden vector and the target embedding vector, the third hidden vector and the target embedding vector can be concatenated, and then the concatenated vector representation is decoded through a classifier (Softmax) to obtain the summary text belonging to the second language corresponding to the text to be processed, which can accurately summarize and translate the text to be processed belonging to the first language into the summary text belonging to the second language, thereby improving the accuracy of obtaining the text summary to a certain extent, and enabling the subsequent user to quickly understand the main content and core expression in the text to be processed according to the obtained summary text.

[0107] Specifically, as Figure 9 and Figure 10 shown, after obtaining the third hidden vector and the target embedding vector, the third hidden vector, the target embedding vector, and the translation result of the (t - 1)-th word can be used as the input of the decoder. First, the target embedding vector and the translation result of the (t - 1)-th word are compiled. Specifically, the decoder can use the self-attention mechanism to encode and learn the translation result of the (t - 1)-th word to obtain a compiled vector representation capable of representing the t-th word, as shown in the following formula (1):

[0108] H y =MultiHead(y, y, y) (1)

[0109] where y is the vector representation corresponding to the translation result of the (t - 1)-th word, and MultiHead is the multi-head self-attention mechanism.

[0110] Further, the target embedding vector is encoded and learned through another self-attention mechanism in the decoder to obtain the corresponding compiled vector representation as H x , then, the compiled vector representation as H x interacts with H y to obtain the interaction representation in the following formula (2):

[0111]

[0112] where FFN is the feed-forward neural network.

[0113] Further, the compiled vector representations corresponding to the target embedding vector and the translation result of the (t - 1)-th word can be concatenated with the third hidden vector, and the vector representation in the following formula (3) can be obtained through concatenation to enrich the reference factors for learning text translation and summary. Then, the concatenated vector representation is decoded through the classifier to accurately obtain the translation result of the t-th word in the following formula (4). When t is 1, the vector input to the decoder is a preset vector with a fixed dimension:

[0114]

[0115] p t = Softmax(W o o t + b o ) (4)

[0116] where z cls is the third hidden vector, W p , W o are learning matrix parameters, and b o is the learning parameter.

[0117] Further, the process of repeatedly compiling the target embedding vector and the translation result of the (t - 1)-th word, concatenating the obtained vector representation with the third hidden vector, and then decoding the concatenated vector representation through the classifier to obtain the translation result of the t-th word is repeated until a special termination symbol is reached, completing the translation and summary of the text to be processed and obtaining the summary text of the text to be processed.

[0118] In an embodiment of the present application, a text processing method is provided. Through the above method, it is realized that through the first prior network and the second prior network, the target embedding vector belonging to the first language can be synchronously converted into the first hidden vector belonging to the second language and the second hidden vector representing the text summary. By using the first hidden vector, the second hidden vector, and the target embedding vector as the input of the third prior network, the third hidden vector that can be used to represent the text summary belonging to the second language can be obtained. Then, by decoding the third hidden vector and the target embedding vector, the synchronous translation and summary of the target embedding vector can be realized. Without going through a series of cumbersome step-by-step processing procedures, the summary and translation of the text to be processed can be completed synchronously, which can reduce the complexity of text processing and thus improve the efficiency of text processing.

[0119] Optionally, on the basis of the above Figure 2 corresponding embodiment, in another optional embodiment of the text processing method provided by the embodiments of the present application, as Figure 3 shown, generating a target embedding vector according to N word embedding vectors includes:

[0120] In step S301, convert N word embedding vectors into an N*d-dimensional vector matrix, where d is an integer greater than 1;

[0121] In step S302, perform dimensionality reduction processing on the N*d-dimensional vector matrix to obtain a target embedding vector, and the target embedding vector is a 1*d-dimensional vector.

[0122] In this embodiment, after obtaining N word embedding vectors, the obtained N word embedding vectors can be converted into an N*d-dimensional vector matrix, and then the N*d-dimensional vector matrix can be subjected to dimensionality reduction processing by taking the average to obtain a target embedding vector, so that the hidden vector corresponding to the text to be processed can be obtained based on the target embedding vector in the subsequent process, preparing for the acquisition of the hidden variable, so that the subsequent model can better summarize or translate the text based on the learning of the hidden vector.

[0123] Specifically, as Figure 9 shown, when N word embedding vectors are obtained, in this embodiment, the target embedding vector can be obtained by taking the average of the vector dimensions of the N word embedding vectors. Specifically, the obtained N word embedding vectors can be stacked to obtain an N*d-dimensional vector matrix, and then, the average can be taken for the N*d-dimensional vector matrix to obtain a 1*d-dimensional vector that can be used to represent the core content of the text to be processed, that is, the target embedding vector.

[0124] For example, assume that there are 10 word embedding vectors of 512 dimensions corresponding to a text to be processed. The 10 word embedding vectors can be encoded to obtain a vector matrix of 10 * 512 dimensions. Then, the average value of the vector matrix is calculated to obtain a 1 * 512-dimensional vector, which is the target embedding vector.

[0125] Optionally, based on the above Figure 2 corresponding embodiment, in another optional embodiment of the text processing method provided by the embodiments of the present application, as Figure 4 shown, the method further includes:

[0126] In step S401, a sample training set is obtained. The sample training set includes a first original sample, a first target sample, a second original sample, a second target sample, a third original sample, and a third target sample. The first target sample is the translated text of the first original sample, the second target sample is the abstract text of the second original sample, and the third target sample is the abstract text after translation of the third original sample;

[0127] In step S402, a first predicted probability distribution is output through a first prior network according to the first original sample, a second predicted probability distribution is output through a second prior network according to the second original sample, and a third predicted probability distribution is output through a third prior network according to the third original sample;

[0128] In step S403, a first true probability distribution is output through a first recognition network according to the first target sample, a second true probability distribution is output through a second recognition network according to the second target sample, and a third true probability distribution is output through a third recognition network according to the third target sample, the first true probability distribution, and the second true probability distribution;

[0129] In step S404, the model parameters of the first prior network are updated according to the divergence between the first predicted probability distribution and the first true probability distribution;

[0130] In step S405, the model parameters of the second prior network are updated according to the divergence between the second predicted probability distribution and the second true probability distribution;

[0131] In step S406, the model parameters of the third prior network are updated according to the divergence between the third predicted probability distribution and the third true probability distribution.

[0132] In this embodiment, the first original sample and the first target text belong to the translation corpus training set. Among them, the first original sample can specifically be a text belonging to the first language, and the first target sample is the translated text obtained by translating the first original sample, and the first target sample belongs to the second language. The second original sample and the second target text belong to the monolingual corpus training set. Among them, the second original sample can specifically be a text belonging to the first language, and the second target sample is the abstract text of the second original sample, and the second target sample belongs to the first language. The third original sample and the third target text belong to the cross-lingual corpus training set. Among them, the third original sample can specifically be a text belonging to the first language, and the third target sample is the abstract text after translating the third original sample, and the third original sample belongs to the second language.

[0133] Specifically, when the first original sample, the first target sample, the second original sample, the second target sample, the third original sample, and the third target sample are obtained, first, the first original sample, the first target sample, the second original sample, the second target sample, the third original sample, and the third target sample are respectively subjected to vector conversion through the embedding layer, and the word embedding vectors corresponding to the first original sample, the first target sample, the second original sample, the second target sample, the third original sample, and the third target sample can be obtained.

[0134] Further, as Figure 9 shown, the word embedding vectors corresponding to the first original sample such as X mt , the second original sample such as X ms , and the third original sample such as X cls After being encoded by the encoder respectively, the corresponding vector representations can be obtained, and the average processing is respectively performed on the vector representations. Among them, the embedding vector obtained by the average processing corresponding to the first original sample can be expressed as The embedding vector obtained by the average processing corresponding to the second original sample can be expressed as The embedding vector obtained by the average processing corresponding to the third original sample can be expressed as

[0135] Further, the embedding vector representation obtained by the average processing corresponding to the first original sample can be output through the first prior network to obtain the first predicted probability distribution, and the embedding vector representation obtained by the average processing corresponding to the second original sample can be output through the second prior network to obtain the second predicted probability distribution, and the embedding vector representation obtained by the average processing corresponding to the third original sample can be output through the third prior network to obtain the third predicted probability distribution.

[0136] Similarly, as Figure 9 shown, the first target sample such as Y mt , the second target sample such as Yms and the third target sample such as Y cls The word embedding vectors corresponding to them respectively, after being encoded by the encoder, can obtain the corresponding vector representations respectively, and average processing is performed on the vector representations respectively. The embedding vector obtained by the average processing corresponding to the first target sample can be expressed as The embedding vector obtained by the average processing corresponding to the second target sample can be expressed as The embedding vector obtained by the average processing corresponding to the third target sample can be expressed as

[0137] Furthermore, the embedding vector representation obtained by the average processing corresponding to the first target sample can output the first true probability distribution through the first recognition network, and the embedding vector representation obtained by the average processing corresponding to the second target sample can output the second true probability distribution through the second recognition network, and the embedding vector representation obtained by the average processing corresponding to the third target sample can output the third predicted true distribution through the third recognition network. Among them, the first recognition network, the first recognition network, and the first recognition network can specifically be expressed as a Gaussian distribution model, or other recognition networks, such as Naive Bayes, etc., which are not specifically limited here.

[0138] Furthermore, as Figure 9 shown, after obtaining the first predicted probability distribution and the first true probability distribution, the second predicted probability distribution and the second true probability distribution, and the third predicted probability distribution and the third true probability distribution, the model parameters of the first prior network can be updated according to the divergence between the first predicted probability distribution and the first true probability distribution. Similarly, the model parameters of the second prior network can be updated according to the divergence between the second predicted probability distribution and the second true probability distribution, and the model parameters of the third prior network can be updated according to the divergence between the third predicted probability distribution and the third true probability distribution. Specifically, the model parameters can be updated by using the gradient descent method, and there can also be parameter update methods, which are not specifically limited here, and can stably update and converge in the direction of gradient update to better update the model parameters, thereby improving the learning ability of the model, improving the training accuracy of the model, and thus being able to improve the accuracy of the model for obtaining text summaries to a certain extent.

[0139] Among them, the calculation of the divergence between the first predicted probability distribution and the first true probability distribution, the divergence between the second predicted probability distribution and the second true probability distribution, and the divergence between the third predicted probability distribution and the third true probability distribution can specifically use the relative entropy (Kullback-Leibler, KL) to calculate the KL divergence, and other divergence representations can also be used, such as the JS divergence, which is not specifically limited here.

[0140] Optionally, based on the above Figure 4 corresponding embodiment, in another optional embodiment of the text processing method provided by the embodiments of the present application, as Figure 5 shown, outputting a first predicted probability distribution through a prior network according to a first original sample includes:

[0141] In step S501, obtaining at least two word embedding vectors corresponding to the first original sample, where the first original sample belongs to the text corresponding to the first language;

[0142] In step S502, generating an embedding vector of the first original sample according to at least two word embedding vectors corresponding to the first original sample;

[0143] In step S503, inputting the embedding vector of the first original sample into the first prior network, and outputting a first predicted probability distribution through the first prior network.

[0144] Specifically, after obtaining the first original sample, the first original sample can be first converted into at least two word embedding vectors. Among them, the method of converting the first original sample into at least two word embedding vectors is similar to the method of obtaining N word embedding vectors corresponding to the text to be processed in step S101, which will not be elaborated here.

[0145] Furthermore, the at least two word embedding vectors corresponding to the first original sample can be used to obtain an embedding vector of the first original sample. Among them, the method of generating an embedding vector of the first original sample according to at least two word embedding vectors corresponding to the first original sample is similar to the method of generating a target embedding vector according to N word embedding vectors in step S102, which will not be elaborated here.

[0146] Furthermore, inputting the embedding vector of the first original sample X mt into the first prior network, and outputting a first predicted probability distribution through the first prior network can specifically be to first obtain a latent variable Z corresponding to the embedding vector of the first original sample through the first prior network mt , and then by defining the latent variable Z mt to follow a multivariate Gaussian distribution, that is where I represents the first original sample. Then, according to the definition, it can be inferred that the first predicted probability distribution composed of the following formulas (5) and (6) can be obtained, so that the model parameters of the first prior network can be updated according to the obtained first predicted probability distribution in the subsequent process, enabling the first prior network to stably update and converge in the gradient update direction, thereby improving the learning ability of the first prior network and the training accuracy of the first prior network, and thus improving the accuracy of the first prior network in obtaining vector representations to a certain extent:

[0147]

[0148]

[0149] Similarly, the first target sample Y mt The embedding vector of can be input into the first recognition network, and the first true probability distribution is output through the first recognition network. Specifically, the latent variable Z corresponding to the embedding vector of the first target sample can be obtained through the first recognition network mt , and then by defining the latent variable Z mt obeys a multivariate Gaussian distribution, that is where I represents the first target sample. Then, according to the definition, the first true probability distribution composed of the following formulas (7) and (8) can be inferred:

[0150]

[0151]

[0152] Optionally, on the basis of the above Figure 2 corresponding embodiment, in another optional embodiment of the text processing method provided by the embodiment of the present application, as Figure 6 shown, the second predicted probability distribution is output through the prior network according to the second original sample, including:

[0153] In step S601, at least two word embedding vectors corresponding to the second original sample are obtained. The second original sample belongs to the abstract text corresponding to the first language;

[0154] In step S602, an embedding vector of the second original sample is generated according to at least two word embedding vectors corresponding to the second original sample;

[0155] In step S603, the embedding vector of the second original sample is input into the second prior network, and the second predicted probability distribution is output through the second prior network.

[0156] Specifically, after the second original sample is obtained, the second original sample can be first converted into at least two word embedding vectors. Among them, the method of converting the second original sample into at least two word embedding vectors is similar to the method of obtaining N word embedding vectors corresponding to the text to be processed in step S101, and will not be elaborated here.

[0157] Further, at least two word embedding vectors corresponding to the second original sample may be used to obtain the embedding vector of the second original sample. Among them, the manner of generating the embedding vector of the second original sample based on at least two word embedding vectors corresponding to the second original sample is similar to the manner of generating the target embedding vector based on N word embedding vectors in step S102, which will not be elaborated here.

[0158] Further, the embedding vector of the second original sample is input into the second prior network, and the second predicted probability distribution is output through the second prior network. Specifically, the latent variable Z corresponding to the embedding vector of the second original sample may be obtained through the second prior network first. ms , and then by defining the latent variable Z ms to follow a multivariate Gaussian distribution, that is where I represents the second original sample. Then, according to the definition, the second predicted probability distribution composed of the following equations (9) and (10) can be inferred, so that the model parameters of the second prior network can be updated based on the obtained second predicted probability distribution in the subsequent process, enabling the second prior network to stably update and converge in the gradient update direction. Furthermore, the learning ability of the second prior network can be improved, the training accuracy of the second prior network can be enhanced, and thus the accuracy of the second prior network in obtaining vector representations can be improved to a certain extent:

[0159]

[0160]

[0161] Similarly, the embedding vector of the second target sample Y ms can be input into the second recognition network, and the second true probability distribution is output through the second recognition network. Specifically, the latent variable Z corresponding to the embedding vector of the second target sample may be obtained through the second recognition network first. ms , and then by defining the latent variable Z ms to follow a multivariate Gaussian distribution, that is where I represents the second target sample. Then, according to the definition, the second true probability distribution composed of the following equations (11) and (12) can be inferred:

[0162]

[0163]

[0164] Optionally, based on the corresponding embodiments described above Figure 2 in another optional embodiment of the text processing method provided by the embodiments of the present application, as Figure 7As shown, the third predicted probability distribution is output by the prior network based on the third original sample, including:

[0165] In step S701, at least two word embedding vectors corresponding to the third original sample are obtained. The third original sample belongs to the text corresponding to the first language.

[0166] In step S702, an embedding vector of the third original sample is generated according to at least two word embedding vectors corresponding to the third original sample.

[0167] In step S703, the embedding vector of the third original sample is input into the third prior network, and the third predicted probability distribution is output through the third prior network.

[0168] Specifically, after obtaining the third original sample, the third original sample can be first converted into at least two word embedding vectors. Among them, the method of converting the third original sample into at least two word embedding vectors is similar to the method of obtaining N word embedding vectors corresponding to the text to be processed in step S101, which will not be elaborated here.

[0169] Furthermore, at least two word embedding vectors corresponding to the third original sample can be used to obtain the embedding vector of the third original sample. Among them, the method of generating the embedding vector of the third original sample according to at least two word embedding vectors corresponding to the third original sample is similar to the method of generating the target embedding vector according to N word embedding vectors in step S102, which will not be elaborated here.

[0170] Furthermore, inputting the embedding vector of the third original sample into the third prior network and outputting the third predicted probability distribution through the third prior network can specifically be to first obtain the latent variable Z corresponding to the embedding vector of the third original sample through the third prior network clc , and then by defining the latent variable Z cls to follow a multivariate Gaussian distribution, that is where I represents the third original sample. Then, according to the definition, it can be inferred that the third predicted probability distribution composed of the following equations (13) and (14) is obtained, so that the model parameters of the third prior network can be updated according to the obtained third predicted probability distribution in the subsequent process, enabling the third prior network to stably update and converge in the gradient update direction, thereby improving the learning ability of the third prior network, improving the training accuracy of the third prior network, and thus being able to improve the accuracy of the third prior network in obtaining vector representations to a certain extent:

[0171]

[0172]

[0173] Similarly, the third target sample Y can bems The embedding vector, the latent variable Z corresponding to the embedding vector of the first original sample mt and the latent variable Z corresponding to the embedding vector of the second original sample ms are first concatenated, and then the vector representation obtained by concatenation is input into the third recognition network, and the third true probability distribution is output through the third recognition network. Specifically, the latent variable Z corresponding to the embedding vector of the third target sample can be obtained through the third recognition network cls and then, by defining the latent variable Z cls to follow a multivariate Gaussian distribution, that is where I represents the third target sample. Then, according to the definition, the third true probability distribution composed of the following equations (15) and (16) can be inferred

[0174]

[0175]

[0176] Optionally, based on the corresponding embodiment above Figure 4 In another optional embodiment of the text processing method provided by the embodiments of the present application, as Figure 8 shown, updating the model parameters of the first prior network according to the divergence between the first predicted probability distribution and the first true probability distribution includes

[0177] In step S801, obtaining the divergence between the first predicted probability distribution and the first true probability distribution

[0178] In step S802, taking the divergence between the first predicted probability distribution and the first true probability distribution as the first loss value

[0179] In step S803, updating the model parameters of the first prior network according to the first loss value

[0180] Updating the model parameters of the second prior network according to the divergence between the second predicted probability distribution and the second true probability distribution includes

[0181] In step S804, obtaining the divergence between the second predicted probability distribution and the second true probability distribution

[0182] In step S805, taking the divergence between the second predicted probability distribution and the second true probability distribution as the second loss value

[0183] In step S806, updating the model parameters of the second prior network according to the second loss value

[0184] Updating the model parameters of the third prior network according to the divergence between the third predicted probability distribution and the third true probability distribution, including:

[0185] In step S807, obtaining the divergence between the third predicted probability distribution and the third true probability distribution;

[0186] In step S808, taking the divergence between the third predicted probability distribution and the third true probability distribution as the third loss value;

[0187] In step S809, updating the model parameters of the third prior network according to the third loss value.

[0188] Specifically, as Figure 9 shown, after obtaining the first predicted probability distribution and the first true probability distribution, the divergence between the first predicted probability distribution and the first true probability distribution can be obtained according to the following formula (17), and the divergence between the first predicted probability distribution and the first true probability distribution can be taken as the first loss value. Then, the model parameters of the first prior network are updated according to the first loss value. That is, taking the KL divergence between the first predicted probability distribution and the first true probability distribution as the first loss value can generate a more stable gradient update direction to better update the model parameters, thereby improving the training accuracy of the first prior network and making the effect better:

[0189]

[0190] where q(Z mt ) is the first predicted probability distribution, and p(Z mt ) is the first true probability distribution.

[0191] Furthermore, after obtaining the second predicted probability distribution and the second true probability distribution, the divergence between the second predicted probability distribution and the second true probability distribution can be obtained according to the following formula (18), and the divergence between the second predicted probability distribution and the second true probability distribution can be taken as the second loss value. Then, the model parameters of the second prior network are updated according to the second loss value. That is, taking the KL divergence between the second predicted probability distribution and the second true probability distribution as the second loss value can generate a more stable gradient update direction to better update the model parameters, thereby improving the training accuracy of the second prior network and making the effect better:

[0192]

[0193] where q(Z ms ) is the second predicted probability distribution, and p(Z ms ) is the second true probability distribution.

[0194] Further, after obtaining the third predicted probability distribution and the third true probability distribution, the divergence between the third predicted probability distribution and the third true probability distribution can be calculated according to the following formula (19), and the divergence between the third predicted probability distribution and the third true probability distribution can be used as the third loss value. Then, the model parameters of the third prior network are updated according to the third loss value. That is, using the KL divergence between the third predicted probability distribution and the third true probability distribution as the third loss value can generate a more stable gradient update direction to better update the model parameters, thereby improving the training accuracy of the third prior network and achieving better results:

[0195]

[0196] where q(Z cls ) is the third predicted probability distribution, and p(Z cls ) is the third true probability distribution.

[0197] The following is a detailed description of the text processing device in this application. Please refer to Figure 11 , Figure 11 which is a schematic diagram of an embodiment of the text processing device in the embodiment of this application. The text processing device 20 includes:

[0198] An acquisition unit 201, configured to acquire N word embedding vectors corresponding to the text to be processed. The text to be processed belongs to the text corresponding to the first language, and N is an integer greater than 1;

[0199] A generation unit 202, configured to generate a target embedding vector according to the N word embedding vectors;

[0200] A processing unit 203, configured to input the target embedding vector into the first prior network, and output a first hidden vector through the first prior network. The first hidden vector represents the vector representation of the target embedding vector belonging to the second language;

[0201] The processing unit 203 is further configured to input the target embedding vector into the second prior network, and output a second hidden vector through the second prior network. The second hidden vector is the vector representation of the text summary, and the text summary is the summary expression of the text to be processed;

[0202] The processing unit 203 is further configured to input the first hidden vector, the second hidden vector, and the target embedding vector into the third prior network, and output a third hidden vector through the third prior network. The third hidden vector is the vector representation of the text summary belonging to the second language;

[0203] The generation unit 202 is further configured to generate a summary text of the text to be processed according to the third hidden vector and the target embedding vector, where the summary text belongs to the text corresponding to the second language.

[0204] Optionally, in the aboveFigure 11 Based on the corresponding embodiment, in another embodiment of the text processing device provided by the embodiment of the present application, the generating unit 202 may specifically be used for:

[0205] Convert N word embedding vectors into an N*d dimensional vector matrix, where d is an integer greater than 1;

[0206] Perform dimensionality reduction processing on the N*d dimensional vector matrix to obtain a target embedding vector, where the target embedding vector is a 1*d dimensional vector.

[0207] Optionally, based on the above Figure 11 Based on the corresponding embodiment, in another embodiment of the text processing device provided by the embodiment of the present application,

[0208] The obtaining unit 201 is further configured to obtain a sample training set, where the sample training set includes a first original sample, a first target sample, a second original sample, a second target sample, a third original sample, and a third target sample. The first target sample is the translated text of the first original sample, the second target sample is the abstract text of the second original sample, and the third target sample is the abstract text after translation of the third original sample;

[0209] The processing unit 203 is further configured to output a first predicted probability distribution through a first prior network according to the first original sample, and output a second predicted probability distribution through a second prior network according to the second original sample, and output a third predicted probability distribution through a third prior network according to the third original sample;

[0210] The processing unit 203 is further configured to output a first true probability distribution through a first recognition network according to the first target sample, and output a second true probability distribution through a second recognition network according to the second target sample, and output a third true probability distribution through a third recognition network according to the third target sample, the first true probability distribution, and the second true probability distribution;

[0211] The processing unit 203 is further configured to update the model parameters of the first prior network according to the divergence between the first predicted probability distribution and the first true probability distribution;

[0212] The processing unit 203 is further configured to update the model parameters of the second prior network according to the divergence between the second predicted probability distribution and the second true probability distribution;

[0213] The processing unit 203 is further configured to update the model parameters of the third prior network according to the divergence between the third predicted probability distribution and the third true probability distribution.

[0214] Optionally, based on the above Figure 11Based on the corresponding embodiment, in another embodiment of the text processing device provided by the embodiments of the present application, the processing unit 203 may specifically be used for:

[0215] Obtain at least two word embedding vectors corresponding to a first original sample, where the first original sample belongs to a text corresponding to a first language;

[0216] Generate an embedding vector of the first original sample according to the at least two word embedding vectors corresponding to the first original sample;

[0217] Input the embedding vector of the first original sample into a first prior network, and output a first predicted probability distribution through the first prior network.

[0218] Optionally, based on the above Figure 11 Based on the corresponding embodiment, in another embodiment of the text processing device provided by the embodiments of the present application, the processing unit 203 may specifically be used for:

[0219] Obtain at least two word embedding vectors corresponding to a second original sample, where the second original sample belongs to an abstract text corresponding to the first language;

[0220] Generate an embedding vector of the second original sample according to the at least two word embedding vectors corresponding to the second original sample;

[0221] Input the embedding vector of the second original sample into a second prior network, and output a second predicted probability distribution through the second prior network.

[0222] Optionally, based on the above Figure 11 Based on the corresponding embodiment, in another embodiment of the text processing device provided by the embodiments of the present application, the processing unit 203 may specifically be used for:

[0223] Obtain at least two word embedding vectors corresponding to a third original sample, where the third original sample belongs to a text corresponding to the first language;

[0224] Generate an embedding vector of the third original sample according to the at least two word embedding vectors corresponding to the third original sample;

[0225] Input the embedding vector of the third original sample into a third prior network, and output a third predicted probability distribution through the third prior network.

[0226] Optionally, based on the above Figure 11 Based on the corresponding embodiment, in another embodiment of the text processing device provided by the embodiments of the present application, the processing unit 203 may specifically be used for:

[0227] Obtain the divergence between the first predicted probability distribution and the first true probability distribution;

[0228] Take the divergence between the first predicted probability distribution and the first true probability distribution as the first loss value;

[0229] Update the model parameters of the first prior network according to the first loss value;

[0230] The processing unit 203 can specifically be used for:

[0231] Obtain the divergence between the second predicted probability distribution and the second true probability distribution;

[0232] Take the divergence between the second predicted probability distribution and the second true probability distribution as the second loss value;

[0233] Update the model parameters of the second prior network according to the second loss value;

[0234] The processing unit 203 can specifically be used for:

[0235] Obtain the divergence between the third predicted probability distribution and the third true probability distribution;

[0236] Take the divergence between the third predicted probability distribution and the third true probability distribution as the third loss value;

[0237] Update the model parameters of the third prior network according to the third loss value.

[0238] Another aspect of this application provides a schematic diagram of another computer device, as Figure 12 shown, Figure 12 is a schematic diagram of the structure of a computer device provided by an embodiment of this application. The computer device 300 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 331 or data 332 (for example, one or more mass storage devices). Among them, the memory 320 and the storage media 330 can be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device 300. Further, the central processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the computer device 300.

[0239] The computer device 300 may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 333, such as Windows Server TM, Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0240] The above computer device 300 is also used to execute the steps in the corresponding embodiments as Figures 2 to 8 shown.

[0241] On the other hand, the present application provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the steps in the method described in the embodiments as Figures 2 to 8 shown.

[0242] On the other hand, the present application provides a computer program product containing instructions. When it runs on a computer or a processor, it causes the computer or the processor to execute the steps in the method described in the embodiments as Figures 2 to 8 shown.

[0243] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0244] In the several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0245] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0246] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0247] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

Claims

1. A method for text processing, characterized in that, Including: Obtain N word embedding vectors corresponding to the text to be processed, where the text to be processed belongs to the text corresponding to the first language, and N is an integer greater than 1; Generate a target embedding vector according to the N word embedding vectors; Input the target embedding vector into a first prior network, and output a first hidden vector through the first prior network, where the first hidden vector represents the vector representation of the target embedding vector belonging to the second language; Input the target embedding vector into a second prior network, and output a second hidden vector through the second prior network, where the second hidden vector is the vector representation of the text summary, and the text summary is the summary expression of the text to be processed; Input the first hidden vector, the second hidden vector, and the target embedding vector into a third prior network, and output a third hidden vector through the third prior network, where the third hidden vector is the vector representation of the text summary belonging to the second language; Generate a summary text of the text to be processed according to the third hidden vector and the target embedding vector, where the summary text belongs to the text corresponding to the second language.

2. The method according to claim 1, characterized in that, The generating a target embedding vector according to the N word embedding vectors includes: Convert the N word embedding vectors into an N*d dimensional vector matrix, where d is an integer greater than 1; Perform dimensionality reduction processing on the N*d dimensional vector matrix to obtain the target embedding vector, and the target embedding vector is a 1*d dimensional vector.

3. The method according to claim 1, wherein The method further includes: Obtain a sample training set, where the sample training set includes a first original sample, a first target sample, a second original sample, a second target sample, a third original sample, and a third target sample, the first target sample is the translated text of the first original sample, the second target sample is the summary text of the second original sample, and the third target sample is the summary text after translation of the third original sample; Output a first predicted probability distribution through the first prior network according to the first original sample, output a second predicted probability distribution through the second prior network according to the second original sample, and output a third predicted probability distribution through the third prior network according to the third original sample; Output a first true probability distribution through a first recognition network according to the first target sample, output a second true probability distribution through a second recognition network according to the second target sample, and output a third true probability distribution through the third recognition network according to the third target sample, the first true probability distribution, and the second true probability distribution; Update the model parameters of the first prior network according to the divergence between the first predicted probability distribution and the first true probability distribution; Update the model parameters of the second prior network according to the divergence between the second predicted probability distribution and the second true probability distribution; Update the model parameters of the third prior network according to the divergence between the third predicted probability distribution and the third true probability distribution.

4. The method according to claim 3, wherein The outputting a first predicted probability distribution through a prior network according to the first original sample includes: Obtain at least two word embedding vectors corresponding to the first original sample, where the first original sample belongs to the text corresponding to the first language; Generate an embedding vector of the first original sample according to the at least two word embedding vectors corresponding to the first original sample; Input the embedding vector of the first original sample into the first prior network, and output the first predicted probability distribution through the first prior network.

5. The method according to claim 3, wherein The outputting the second predicted probability distribution through the prior network according to the second original sample includes: Obtain at least two word embedding vectors corresponding to the second original sample, where the second original sample belongs to the abstract text corresponding to the first language; Generate an embedding vector of the second original sample according to the at least two word embedding vectors corresponding to the second original sample; Input the embedding vector of the second original sample into the second prior network, and output the second predicted probability distribution through the second prior network.

6. The method according to claim 3, characterized in that The outputting the third predicted probability distribution through the prior network according to the third original sample includes: Obtain at least two word embedding vectors corresponding to the third original sample, where the third original sample belongs to the text corresponding to the first language; Generate an embedding vector of the third original sample according to the at least two word embedding vectors corresponding to the third original sample; Input the embedding vector of the third original sample into the third prior network, and output the third predicted probability distribution through the third prior network.

7. The method according to claim 3, wherein The updating the model parameters of the first prior network according to the divergence between the first predicted probability distribution and the first true probability distribution includes: Obtain the divergence between the first predicted probability distribution and the first true probability distribution; Take the divergence between the first predicted probability distribution and the first true probability distribution as the first loss value; Update the model parameters of the first prior network according to the first loss value; The updating the model parameters of the second prior network according to the divergence between the second predicted probability distribution and the second true probability distribution includes: Obtain the divergence between the second predicted probability distribution and the second true probability distribution; Take the divergence between the second predicted probability distribution and the second true probability distribution as the second loss value; Update the model parameters of the second prior network according to the second loss value; The updating the model parameters of the third prior network according to the divergence between the third predicted probability distribution and the third true probability distribution includes: Obtain the divergence between the third predicted probability distribution and the third true probability distribution; Take the divergence between the third predicted probability distribution and the third true probability distribution as the third loss value; Update the model parameters of the third prior network according to the third loss value.

8. An apparatus for text processing, characterized in that, Includes: An obtaining unit, configured to obtain N word embedding vectors corresponding to the text to be processed, where the text to be processed belongs to the text corresponding to the first language, and N is an integer greater than 1; A generating unit, configured to generate a target embedding vector according to the N word embedding vectors; A processing unit, configured to input the target embedding vector into a first prior network, and output a first hidden vector through the first prior network, where the first hidden vector represents a vector representation of the target embedding vector belonging to a second language; The processing unit is further configured to input the target embedding vector into a second prior network, and output a second hidden vector through the second prior network, where the second hidden vector is a vector representation of a text summary, and the text summary is a summary expression of the text to be processed; The processing unit is further configured to input the first hidden vector, the second hidden vector, and the target embedding vector into a third prior network, and output a third hidden vector through the third prior network, where the third hidden vector is a vector representation of the text summary belonging to the second language; The generating unit is further configured to generate a summary text of the text to be processed according to the third hidden vector and the target embedding vector, where the summary text belongs to the text corresponding to the second language.

9. A computer device, characterized in that, Comprising: A memory, a transceiver, a processor, and a bus system; Wherein, the memory is used for storing programs; The processor is configured to implement the method according to any one of claims 1 to 7 when executing the programs in the memory; The bus system is used for connecting the memory and the processor, so that the memory and the processor can communicate with each other.

10. A computer-readable storage medium, comprising instructions, which when running on a computer, cause the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text feature extraction method and device and storage medium

    CN110321929A

  • Unsupervised semantic representation automatic identification method and device

    CN112364659A