Multilingual knowledge fusion method, text processing method, and model training method
Patent Information
- Application Number
- CN202211089347.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2042-09-07
AI Technical Summary
而传统的匹配方法由于基于文本之间的相似度计算;难以实现复杂语种情况下的语义识别分析,进而导致基于对话的语义分析结果准确率低
[0074] This invention provides a cross-lingual knowledge fusion method. In this method, after acquiring first corpus information, cross-lingual corpus with linguistic differences is further acquired. Then, the first corpus information and the cross-lingual corpus are fused and concatenated using a fully connected method, and then a semilinear transformation is performed to obtain a second corpus information after multilingual knowledge fusion. By fusing corpus materials from different languages, this method can obtain a large number of semantically identical but differently expressed semantic entities. Using these semantic entities as knowledge reserves or corpus support in tasks such as human-computer interaction and text processing can improve the accuracy of text processing and enhance the readability and usability of text during human-computer interaction. Furthermore, this invention employs two fully connected methods and two activation processes, allowing for more thorough fusion of semantic entities from two different corpus sets, thereby improving the accuracy of multilingual knowledge fusion processing.
Smart Images

Figure CN117010401B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, and in particular to multilingual knowledge fusion methods, text processing methods, and model training methods. Background Technology
[0002] In human-computer interaction (HCI) processes, to enhance the user experience of computers and other terminal devices, it is typically necessary to analyze the content of human-computer dialogue using appropriate language models to clarify the intent of the target audience. Currently, semantic recognition and natural language processing technologies have made significant progress and can be widely applied in fields such as autonomous driving, complex scene recognition, intelligent search, and intelligent authentication. For example, in the field of autonomous driving, human-computer dialogue can be used to obtain the driver's operating instructions, which can then be processed through semantic analysis to execute corresponding actions. Similarly, in human-computer dialogue scenarios for intelligent decision-making, it is often necessary to obtain the target audience's needs through multi-turn question-and-answer dialogue and make recommendations based on those needs.
[0003] However, current text processing methods based on semantic recognition utilize linguistic knowledge, multi-grammar, and lexical features, employing traditional matching algorithms for classification and prediction. Traditional matching methods, based on similarity calculations between texts, struggle to achieve semantic recognition and analysis in complex language scenarios, leading to low accuracy in dialogue-based semantic analysis results. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a multilingual knowledge fusion method, a text processing method, and a model training method with high prediction accuracy, so as to improve the prediction accuracy of target text in multilingual language scenarios.
[0005] One aspect of this invention provides a multilingual knowledge fusion method:
[0006] Obtain information from the first corpus;
[0007] Obtain reference text corpus that differs in language from the first corpus information, and determine cross-language corpus based on the reference text corpus;
[0008] The first corpus information and the cross-language corpus are subjected to a first fusion and splicing process to obtain the first fused corpus;
[0009] The first fused corpus is subjected to a nonlinear transformation to obtain the second fused corpus;
[0010] After performing a second fusion splicing process on the second fused corpus, a nonlinear transformation is applied to obtain the second corpus information after multilingual knowledge fusion.
[0011] Another aspect of the present invention provides a text processing method, including:
[0012] Once the target text is obtained, the first word vector of all words in the target text is generated;
[0013] Feature relationships are extracted from the first word vector to obtain the text feature information of the target text;
[0014] The text feature information is classified and identified to obtain the target result corresponding to the target text;
[0015] Wherein, at least one of the target text, the first word vector, and the text feature information is obtained after multilingual knowledge fusion processing using the multilingual knowledge fusion method described in the preceding aspect.
[0016] Another aspect of this invention provides a method for training a multilingual knowledge fusion model, the method comprising:
[0017] Obtain training corpus information and determine the target word vector in the training corpus information;
[0018] The training corpus information of different languages is subjected to a first fusion and splicing process to obtain the first fused corpus;
[0019] The first fused corpus is subjected to a nonlinear transformation to obtain the second fused corpus;
[0020] The second fused corpus is subjected to a second fusion splicing process and a nonlinear transformation to obtain the predicted word vectors after multilingual knowledge fusion;
[0021] The parameters of the multilingual knowledge fusion model are adjusted based on the loss value between the predicted word vector and the target word vector.
[0022] Another aspect of the present invention provides a text processing model training method, including:
[0023] Obtain the training text and determine the target result in the training text;
[0024] Generate the first word vectors for all words in the training text;
[0025] Feature relationships are extracted from the first word vector to obtain the text feature information of the target text;
[0026] The text feature information is classified and identified to obtain the prediction result corresponding to the target text;
[0027] The parameters of the text processing model are adjusted based on the loss value between the prediction result and the target result; wherein, at least one of the training text, the first word vector, and the text feature information is obtained after multilingual knowledge fusion processing using the multilingual knowledge fusion method described in the preceding aspect.
[0028] Another aspect of the present invention provides a multilingual knowledge fusion device, comprising:
[0029] The first module is used to obtain information from the first corpus.
[0030] The second module is used to acquire reference text corpus that differs in language from the first corpus information, and to determine cross-language corpus based on the reference text corpus;
[0031] The third module is used to perform a first fusion and splicing process on the first corpus information and the cross-language corpus to obtain the first fused corpus.
[0032] The fourth module is used to perform a nonlinear transformation on the first fused corpus to obtain the second fused corpus;
[0033] The fifth module is used to perform a second fusion splicing process on the second fused corpus and then a nonlinear transformation to obtain the second corpus information after multilingual knowledge fusion.
[0034] In one possible implementation, the second module includes:
[0035] The first unit is used to perform language conversion on the reference text corpus to obtain cross-language text;
[0036] The second unit is used to determine reference words from the reference text corpus, and to perform word vector processing on the reference words to obtain cross-language word vectors.
[0037] The third unit is used to extract the feature relationships of word vectors in the cross-language word vectors to obtain cross-language text feature corpus;
[0038] The fourth unit is used to construct the cross-language corpus based on the cross-language text, the cross-language word vectors, and the cross-language text feature corpus.
[0039] In one possible implementation, the fourth module includes:
[0040] The fifth unit is used to determine the first weight matrix based on the corpus vector dimension of the first fused corpus and the corpus vector dimension of the second fused corpus.
[0041] The sixth unit is used to determine the first bias vector based on the corpus vector dimension of the second fused corpus;
[0042] The seventh unit is used to perform activation processing on the first fused corpus according to the first activation function, the first weight matrix and the first bias vector to obtain the first function value;
[0043] The eighth unit is used to determine the second fused corpus based on the first function value and the first corpus information.
[0044] In one possible implementation, the fifth module includes:
[0045] The ninth unit is used to determine the second weight matrix based on the corpus vector dimension of the second fused corpus and the corpus vector dimension of the first corpus information;
[0046] The tenth unit is used to determine the second bias vector based on the vector dimension of the first corpus information;
[0047] The eleventh unit is used to activate the second fused corpus according to the second activation function, the second weight matrix and the second bias vector to obtain the second function value;
[0048] The twelfth unit is used to determine the second corpus information based on the second function value and the second fused corpus.
[0049] Another aspect of the present invention provides a text processing apparatus, comprising:
[0050] The sixth module is used to obtain the target text and generate the first word vector of all words in the target text;
[0051] The seventh module is used to extract feature relationships from the first word vector to obtain the text feature information of the target text;
[0052] The eighth module is used to classify and identify the text feature information to obtain the target result corresponding to the target text;
[0053] Wherein, at least one of the target text, the first word vector, and the text feature information is obtained after multilingual knowledge fusion processing using the aforementioned multilingual knowledge fusion method.
[0054] In one possible implementation, the sixth module includes:
[0055] Unit 13 is used to extract text content from the text to be processed, resulting in the first text, the second text, and the third text.
[0056] The fourteenth unit is used to perform a first word segmentation process on the first text according to the first text length of the text paragraphs in the first text to obtain a first word sequence;
[0057] The fifteenth unit is used to perform second word segmentation on the second text based on the second text length of the question statement in the second text to obtain a second word sequence;
[0058] The sixteenth unit is used to perform third word segmentation on the third text based on the length of the third text of the candidate answers in the third text to obtain a third word sequence;
[0059] The seventeenth unit is used to determine candidate words based on the first word sequence, the second word sequence, and the third word sequence;
[0060] The eighteenth unit is used to perform multilingual corpus fusion on the candidate words to obtain the first word vector.
[0061] In one possible implementation, the eighteenth unit includes:
[0062] The first subunit is used to perform vectorization processing on the candidate words to obtain candidate word vectors;
[0063] The second subunit is used to perform multilingual corpus fusion on the candidate word vectors to obtain the first word vector.
[0064] In one possible implementation, the seventh module includes:
[0065] The eighteenth unit is used to extract feature relationships from the first word vector and generate feature vectors;
[0066] The nineteenth unit is used to perform multilingual corpus fusion on the feature vector to obtain the second word vector;
[0067] The twentieth unit is used to determine the text feature information based on the second word vector.
[0068] In one possible implementation, the eighteenth unit includes:
[0069] The third subunit is used to determine the context relationship based on the information carried by the adjacent position vectors of the first word vector; the adjacent position vectors include at least one of the vectors of the first word vector at the first few positions or the vectors of the first word vector at the last few positions.
[0070] The fourth subunit is used to encode the context relationship and determine several encoded vectors corresponding to different adjacent position vectors;
[0071] The fifth subunit is used to integrate the various encoded vectors to obtain the feature vector.
[0072] Another aspect of the present invention provides a computer-readable storage medium storing a program that is executed by a processor to implement the aforementioned multilingual knowledge fusion method, text processing method, multilingual knowledge fusion model training method, and text processing model training method.
[0073] This application also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned multilingual knowledge fusion method, text processing method, multilingual knowledge fusion model training method, and text processing model training method.
[0074] This invention provides a cross-lingual knowledge fusion method. In this method, after acquiring first corpus information, cross-lingual corpus with linguistic differences is further acquired. Then, the first corpus information and the cross-lingual corpus are fused and concatenated using a fully connected method, and then a semilinear transformation is performed to obtain a second corpus information after multilingual knowledge fusion. By fusing corpus materials from different languages, this method can obtain a large number of semantically identical but differently expressed semantic entities. Using these semantic entities as knowledge reserves or corpus support in tasks such as human-computer interaction and text processing can improve the accuracy of text processing and enhance the readability and usability of text during human-computer interaction. Furthermore, this invention employs two fully connected methods and two activation processes, allowing for more thorough fusion of semantic entities from two different corpus sets, thereby improving the accuracy of multilingual knowledge fusion processing.
[0075] This invention proposes a text processing method based on multilingual knowledge fusion. This method can perform multilingual corpus fusion during the text preprocessing stage to obtain the target text; or it can perform multilingual corpus fusion during the generation of a candidate word set for the target text to obtain a first word vector set; or it can perform multilingual corpus fusion during feature information extraction to obtain text feature information incorporating multilingual knowledge. In each of the aforementioned processing steps, multilingual corpus fusion can be selectively performed, resulting in a large number of multilingual semantic entities. Finally, the target result is classified and identified based on the multilingual corpus fusion results, improving the accuracy of semantic analysis and inference prediction in complex linguistic environments. Furthermore, this invention performs multilingual corpus fusion, outputting multilingual semantic entities at different stages. The method in this invention performs multilingual corpus fusion at least once and can be flexibly applied to any stage of text processing, effectively reducing the number of parameters in model building and making the actual operation simpler and more efficient. Attached Figure Description
[0076] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0077] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of the present invention;
[0078] Figure 2 A flowchart illustrating the steps of a text processing method for predicting target results in relevant technical solutions;
[0079] Figure 3 This is a flowchart illustrating the steps of the multilingual knowledge fusion method provided in this embodiment of the invention.
[0080] Figure 4 This is a schematic diagram of the structure of a multilingual knowledge fusion model provided in an embodiment of the present invention;
[0081] Figure 5 This is a schematic diagram of the structure of a text processing model provided in an embodiment of the present invention;
[0082] Figure 6 This is a flowchart of the steps of a text processing method provided in an embodiment of the present invention;
[0083] Figure 7 This is a flowchart illustrating the steps of word segmentation for the text to be processed in an embodiment of the present invention.
[0084] Figure 8 This is a schematic diagram of the cross-language knowledge representation acquisition model in an embodiment of the present invention;
[0085] Figure 9 This is a flowchart illustrating the steps of a multilingual knowledge fusion model training method provided in an embodiment of the present invention.
[0086] Figure 10 This is a schematic diagram of the structure of the multilingual knowledge fusion device provided in the embodiment of the present invention;
[0087] Figure 11 This is a schematic diagram of the structure of the text processing device provided in the embodiment of the present invention;
[0088] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0089] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0090] It is understood that the terms "first," "second," etc., used in this application may be used to describe various concepts herein, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of this application, the first set of word vectors may be referred to as the second set of word vectors, and the second set of word vectors may be referred to as the first set of word vectors.
[0091] In addition, the terms “at least one,” “multiple,” “each,” “any,” etc., used in this application, “at least one” includes one, two, or more than two, “multiple” includes two or more than two, “each” refers to each of the corresponding multiple, and “any” refers to any one of the multiple.
[0092] The specification will describe exemplary embodiments in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0093] Before providing a detailed description of the embodiments of this application, necessary explanations will be given for the technical terms that may be involved in the embodiments of the application's technical solutions:
[0094] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0095] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning. The display device with image acquisition components shown in this application mainly relates to computer vision, machine learning / deep learning, autonomous driving, and intelligent transportation.
[0096] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learn-by-doing.
[0097] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, intent recognition, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.
[0098] Multi-layer Perceptron (MLP): Based on biological neuron models, researchers have constructed the basic structure of multi-layer feedforward neural networks. The most typical multi-layer feedforward neural network mainly includes three layers: input layer, hidden layer, and output layer. Furthermore, the different layers of the multi-layer feedforward neural network are connected in a fully connected manner. Fully connected means that any neuron in the previous layer is connected to all neurons in the next layer.
[0099] Long Short-Term Memory (LSTM) is a special type of Recurrent Neural Network (RNN) primarily designed to address the vanishing and exploding gradient problems during training of long sequences. In short, LSTM performs better with longer sequences than regular RNNs.
[0100] Based on the aforementioned theoretical foundations, and with the research and advancements in artificial intelligence technology, AI technology has been studied and applied in multiple fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. As technology develops, AI technology will be applied in more fields and play an increasingly important role.
[0101] It should be understood that the multilingual knowledge fusion method, text processing method, multilingual knowledge fusion model training method, and text processing model training method provided in the embodiments of the present invention can all be applied to any computer device with data processing and computing capabilities, and this computer device can be various types of terminals or servers. When the computer device in the embodiments is a server, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. Optionally, the terminal can be a smartphone, tablet computer, laptop computer, or desktop computer, but it is not limited to the aforementioned options.
[0102] It should be further noted that the terminals involved in the embodiments of this application include, but are not limited to, smartphones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, and aircraft. The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.
[0103] In some possible implementations, the computer programs for the multilingual knowledge fusion method, text processing method, multilingual knowledge fusion model training method, and text processing model training method provided in the embodiments of the present invention can be deployed and executed on a computer device, or executed on multiple computer devices located in one location; or, executed on multiple computer devices distributed in multiple locations and interconnected through a communication network, wherein the multiple computer devices distributed in multiple locations and interconnected through a communication network can form a blockchain system.
[0104] Based on the fact that multiple computer devices can form a blockchain system, the multilingual knowledge fusion method, text processing method, multilingual knowledge fusion model training method, and text processing model training method implemented in this embodiment of the invention can be a node in the blockchain. This node can pre-store the corresponding text processing model. During model training, training text data can be broadcast to the blockchain nodes. Based on the synchronized training text data, the local text processing model to be trained on the node is trained. Through training, the parameters and structure of the model are optimized and adjusted to obtain the trained text processing model. Similarly, in the process of text content prediction, the text content to be predicted can be broadcast to the blockchain nodes. The trained text processing model on the node predicts the target text based on the text content. The obtained target text can also be broadcast to other nodes in the blockchain for distributed storage.
[0105] like Figure 1 The diagram shown is a schematic representation of an implementation environment provided by an embodiment of the present invention. (Refer to...) Figure 1In this implementation environment, it mainly includes at least one interactive terminal 110 and a text processing backend server 120. Based on the terminal 110 and the server 120, the text processing method or text processing model training method in the embodiments of the present invention can be implemented collaboratively. Furthermore, those skilled in the art will understand that the text processing method or text processing model training method in the embodiments of the present invention can also be implemented independently by the terminal device or independently by the server. The terminal 110 is connected to the text processing backend server 120 via a wireless network or a wired network. The terminal 110 can be at least one of a smartphone, camera, desktop computer, tablet computer, MP4 player, and laptop computer. The terminal 110 has an application that supports human-computer interaction or human-computer dialogue installed and running. For example, the terminal 110 can be a terminal equipped with an interactive application client, and the application client running on the terminal contains the account of the account owner. The terminal 110 can refer to one of multiple terminals; those skilled in the art will understand that the number of terminals can be more or less. For example, there may be only one remote terminal, or there may be dozens or hundreds of remote terminals, or even more. This application embodiment does not limit the number and type of terminals 110. The text processing backend server 120 can be configured as a single server, multiple servers, or a combination of cloud servers. The text processing backend server 120 is used to provide text processing capabilities, and based on these capabilities, it can perform semantic analysis, reading comprehension, and intent analysis on natural language text, and output and feedback the corresponding results.
[0106] In related technical solutions, during the process of text processing backend server 120 performing tasks such as semantic analysis, reading comprehension, and intent analysis, it typically utilizes linguistic knowledge, multi-grammar, and lexical features, and predicts the target text content based on matching algorithms. For example... Figure 2 As shown, some other related technical solutions also provide a text processing method for predicting target results based on a deep learning model. This model mainly includes an encoding layer and a prediction layer, performing feature encoding on the target text before classification and prediction. While the aforementioned solutions can improve text processing efficiency and prediction accuracy, in more complex linguistic environments, such as when the text content includes sentences in multiple different languages, the efficiency improvement effect of the deep learning models provided in these solutions is quite limited. Furthermore, the accuracy and readability of the model outputs also decrease.
[0107] based on Figure 1To address the potential low prediction accuracy during the matching prediction process in the implementation environment, this invention proposes a multilingual knowledge fusion method. The method first acquires corpus information from the speech analysis process. This corpus information can include material content at different text granularities, including but not limited to text passages, word vectors, or relational feature information between word vectors. After determining the corpus information, the embodiment acquires reference text corpus in a different language than the corpus information. Based on the text granularity of the corpus information, this reference text corpus undergoes necessary splitting processing to obtain cross-lingual corpus with the same text granularity as the corpus information. Further, the embodiment fuses and concatenates semantic entities with the same expressive meaning in the corpus information and the cross-lingual corpus. Through two fully connected fusion concatenations and two nonlinear transformations, the final fused corpus information is obtained. The obtained corpus information can serve as a knowledge base and data support for text processing tasks such as reading comprehension and semantic analysis.
[0108] After obtaining the fused corpus information, embodiments of the present invention can also obtain conversational text content as target text through human-computer interaction via terminal 110, or obtain preliminary text to be processed as text material through human-computer interaction via terminal 110, and perform multilingual corpus fusion processing. During the fusion process, the fused text material is obtained through the aforementioned multilingual knowledge fusion processing. After obtaining the text material, the embodiments perform word-level segmentation of the text material based on the architecture of deep learning models such as semantic recognition models, and vectorize each segmented word text to obtain the corresponding word vector representation. All word vectors in the target text are integrated to obtain a word vector set; or, after the text is segmented, a candidate word set is obtained. In this stage, multilingual corpus fusion is performed, and word texts in other languages are translated and merged to obtain a knowledge-fused word vector set. After obtaining a set of word vectors through vectorization, semantic analysis and target result prediction are performed on the vectorized word vector set following the architectural principles of deep learning models such as semantic recognition models. In this embodiment, feature relationships can be extracted from each word vector in the word vector set. Extraction methods include, but are not limited to, based on the positional relationships or semantic context relationships of words in the text, or calculating the similarity between word vectors to determine this feature relationship. In this embodiment, multilingual corpora can also be fused during the relationship extraction stage. By fusing the vector representations of feature relationships between word vectors in other language environments with the currently extracted feature relationships, fused text feature information is obtained. After obtaining the text feature information, this embodiment uses prediction processing such as classification and recognition to obtain the final target result. The target result obtained in this embodiment has different expressions in different implementation environments; in human-computer interaction, it can be the target answer to the corresponding question in the input dialogue text; or it can be the emotional expression obtained from semantic analysis based on the language content of the target object in the behavior analysis. It should be noted that the multilingual corpus fusion process in this embodiment can be integrated into any of the processes such as input preprocessing, word vector processing, and feature extraction to fuse multilingual knowledge; however, to ensure the successful fusion of multilingual knowledge and thus improve the accuracy and usability of the target results, the knowledge fusion model in this embodiment should perform multilingual corpus fusion processing at least once. In summary, in Figure 1 In the implementation environment shown, the text processing method provided by the embodiments of the present invention can effectively reduce the number of parameters and save computational costs on the basis of existing deep learning models, while also improving the accuracy of reading comprehension in interactive scenarios.
[0109] Furthermore, based on Figure 1In the implementation environment shown, this embodiment of the invention also provides a text processing model training method. The method primarily uses existing text data as training text to train the corresponding text processing model. The training process of the text processing model is similar to the aforementioned text processing method, except that the text content input to the embodiment is replaced; therefore, it will not be described in detail here.
[0110] It should be further noted that in all specific embodiments of this application, when processing is required based on the target object's information, behavioral data, historical data, and interaction data generated during human-computer interaction, the target object's permission or consent must be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards of the relevant countries and regions. In addition, when embodiments of this application require obtaining sensitive information about the target object, separate permission or consent from the target object will be obtained through pop-ups or redirection to a confirmation page. Only after obtaining the target object's separate permission or consent will the necessary target object-related data for the normal operation of the embodiments of this application be obtained.
[0111] like Figure 3 The diagram illustrates a multilingual knowledge fusion method provided by an embodiment of the present invention. The method can be derived from... Figure 1 The text processing backend server 120 executes the process, or it executes on some terminals 110 with certain data processing capabilities, or it executes the process interactively between terminal 110 and text processing backend server 120. (See reference) Figure 3 The method mainly includes steps S310-S350:
[0112] S310, Obtain information from the first corpus;
[0113] The first corpus information can be a collection of textual material at any textual granularity, including but not limited to text chapters, text paragraphs, sentences, phrases, word vectors, and the relationship features between word vectors, etc.
[0114] For example, taking an implementation scenario where semantic analysis is performed by a text processing model, the embodiment can perform vectorization processing on all word texts in the candidate word set input into the model through the word vector layer in the text processing model to obtain corresponding word vectors, and integrate the word vectors to obtain a candidate word vector set. It should be noted that, in the embodiment, when the text processing model adopts a natural language processing model such as Bert as the basic architecture, the word segmentation and word vector processing can be performed through the word vector layer in the architecture. For example, after the Tokenizer built in the word vector layer performs word segmentation according to a word list or dictionary to obtain a candidate word set, the word vector layer further performs matching according to the word list to obtain the word vector corresponding to each candidate word. When the text processing model adopts the basic architecture of other machine learning models, other tools, such as word2vec, can also be used to measure parts of speech, emotional tendency, degree and other aspects, and a set of values is used to represent a word, so as to determine the word vector representation of candidate words.
[0115] S320: Acquire reference text corpus that has language difference from the first corpus information, and determine cross-language corpus according to the reference text corpus;
[0116] Wherein, the reference text corpus may be a text corpus acquired from stored historical text data, or a text corpus acquired from other open source databases; and the text content of the acquired reference text corpus shall be different from the text content in the first corpus information in the embodiment in terms of expression language of at least a single word.
[0117] Specifically in the embodiment, the reference text corpus may be acquired from stored historical text data, or acquired from other open source databases; for example, if the first corpus information is selected as text content in Chinese, then the corresponding reference text corpus shall be selected from text in English, Japanese or other languages as reference. For another example, there may be expressions in both Chinese and English in the first corpus information, then when selecting the reference text, there is a difference in the expression of at least a single word in the reference text; when the words or phrases "乐意" and "be glad to" exist in the target text, the corresponding expression in French in the reference text is "volontiers". After acquiring the reference text corpus, the embodiment can translate the reference text corpus through a machine translation tool, so that the language of the reference text corpus is consistent with that of the target text; then perform word segmentation and vectorization processing on the translated reference text corpus to obtain a cross-language corpus with the same text granularity as the first corpus information.
[0118] S330: Perform a first fusion and concatenation process on the first corpus information and the cross-language corpus to obtain a first fused corpus;
[0119] The first fused corpus is intermediate data in this embodiment, obtained by fusing and concatenating the information from the first corpus with cross-linguistic corpora. Specifically, in this embodiment, as... Figure 4 As shown, this embodiment provides a Cross-lingual Knowledge Fusion Network (CKFN) model to implement the multilingual knowledge fusion method achieved through fully connected layers and activation functions, as described in this embodiment. Furthermore, the CKFN model can be integrated into different reading comprehension models, such as BERT-based or LSTM-based reading comprehension models. In this embodiment, the CKFN model may include two fully connected layers and two activation function layers. The input text content is first fully connected through a first fully connected layer, then activated by a first activation function. The result is then input to a second fully connected layer for a second fully connected layer, followed by another activation function. After nonlinear transformation through the two activation functions, the corrected text content is output. It should be noted that the correction process based on the CKFN model is the process of fusing reference text (knowledge) from other languages into the target text.
[0120] Based on the two-layer architecture of this model, in the embodiment, after obtaining the first corpus information (denoted as D) and the cross-lingual corpus (denoted as E) with the same text granularity, D and E are input into the CKFN model in the embodiment. First, the word vectors in the two sets D and E are structurally concatenated and combined. Then, the concatenated vector is passed through the first fully connected layer of the CKFN model to fuse the cross-lingual vector representation in set E with the candidate word vectors in set D, and output to the first activation function layer of the CKFN model.
[0121] S340. Perform a nonlinear transformation on the first fused corpus to obtain the second fused corpus;
[0122] The second fused corpus can be corpus content obtained after nonlinear transformation through a similar activation function, and is also intermediate data in the multilingual knowledge fusion process. Specifically, in the embodiment, the nonlinear transformation can refer to the activation operation performed by the two activation functions in the CKFN model of the embodiment. In the embodiment, after fully connecting the vectors in set D and set E in the first fully connected layer, the fully connected vectors are input one by one into the first activation function layer for activation operation, that is, the hidden layer vectors in the CKFN model obtained by nonlinear transformation are used as the second fused corpus in the embodiment. It should be noted that the significance of the hidden layer of the model is to abstract the features of the input data to another dimension space to show its more abstract features, which can be better linearly divided; correspondingly, in the CKFN model of this embodiment, the two fully connected layers and the activation function layer together constitute the hidden layer of the model.
[0123] For example, in the CKFN model of this embodiment, the first layer activation function is selected as the ReLU function. The embodiment obtains the activation function by performing the ReLU function on the vector output by the first fully connected layer. Vector r is the vector representation of the cross-linguistic vector representation E and the candidate word vector representation D after passing through the first linear layer; d r This is the hidden layer dimension of the CKFN model, which refers to the size of the vector representation r.
[0124] S350. After performing second fusion splicing processing on the second fused corpus, a nonlinear transformation is performed to obtain the second corpus information after multilingual knowledge fusion.
[0125] Specifically, in this embodiment, the second fully connected layer and the second activation function layer in the CKFN model operate in the same way as the first fully connected layer and the first activation function layer. This two-layer architecture of the CKFN model results in better text knowledge fusion compared to a single-layer model. For example, in this embodiment, after obtaining the hidden layer vector r through a nonlinear transformation of the first activation layer, the corrected corpus information is finally obtained through a second fully connected layer and a ReLU nonlinear transformation layer. Among them, corpus information It can be used to represent cross-language knowledge (text) as vectors, and the output will be several vectors. The integration process yields the complete content of the second corpus information after the fusion of language knowledge.
[0126] In order to achieve multilingual knowledge fusion at any stage of the text processing process, or to achieve knowledge fusion model integration at any level in the text processing model, the step S320 of obtaining reference text corpus that differs in language from the first corpus information and determining cross-lingual corpus based on the reference text corpus may include steps S321-S324:
[0127] S321. Perform language conversion on the reference text corpus to obtain cross-language text;
[0128] Specifically, in the embodiments, for text materials with large text granularity such as chapters or paragraphs, the embodiments can directly translate the text chapters or paragraphs to achieve language conversion and obtain the corresponding cross-language text. For example, in the embodiments, English text paragraphs can be gradually translated into Chinese to obtain Chinese-expressed text content as cross-language text.
[0129] S322. Determine reference words from the reference text corpus, perform word vector processing on the reference words, and obtain cross-language word vectors;
[0130] Specifically, in the embodiments, for reference text corpora at the granularity of phrases or word vectors, the embodiments can use word segmentation and word vectorization to obtain the vector representation of words, which is more convenient for subsequent concatenation processing and activation operations. For example, the embodiments perform simple preprocessing on the acquired reference text corpora to remove noise content from the text. For instance, for English text, firstly, misspelled words and meaningless symbols are removed. After noise removal, the text is divided into segments at different granularities based on its content. For example, longer articles are first broken down into sentences, and paragraphs are divided according to punctuation marks to obtain sentence-level text material. Based on the sentence-level granularity, the text material is further segmented into words, and sentences are split according to a preset English thesaurus or dictionary to obtain word-level text material. After obtaining a set of words or phrases at the word granularity, each English word in the set is translated one by one to obtain a set of Chinese phrases. Then, each word or phrase in the set is vectorized to obtain a corresponding reference word vector. Finally, the reference word vectors are integrated to obtain a set of cross-language word vectors.
[0131] S323. Extract the feature relationships of word vectors in cross-linguistic word vectors to obtain cross-linguistic text feature corpus;
[0132] In this context, the feature relationship mainly refers to the association and connection between various word vectors that can be used to describe the relationship between them in the embodiment. Specifically, in the embodiment, based on the cross-language word vectors (set) generated in step S322, the embodiment can further encode the cross-language word vectors in the form of a sequence. During the feature encoding process, all cross-language word vectors in the text sequence are read at once, so that the embodiment can extract features based on both sides of the cross-language word vector sequence. For the context features of the cross-language word vectors, they are added to the corresponding word embedding token of the word vector through encoding, and a set of cross-language text feature corpus is output. In this set of cross-language text feature corpus, each vector corresponds to a token with the same index.
[0133] S324. A cross-language corpus is constructed based on cross-language text, cross-language word vectors, and cross-language text feature corpus.
[0134] Specifically, in the embodiments, since the multilingual knowledge fusion method can be applied in various processing flows of the text processing method, or the knowledge fusion model corresponding to the multilingual knowledge fusion method can be integrated into any level of the text processing model, the content or format of the corpus used as input for knowledge fusion is not the same in different processing flows of the method or in different levels of the model. Therefore, the technical solution of this application can select the cross-language text, cross-language word vectors and cross-language text feature corpus obtained in steps S321-S323 to jointly constitute the cross-language corpus in the embodiments, and can select the cross-language corpus with the corresponding text granularity as the corpus material to be fused and spliced with the first corpus information according to the specific application stage or the level of the model.
[0135] For example, in the process of constructing a cross-language corpus based on a reference text corpus, the embodiment can split the reference text corpus according to its length structure to obtain the text S1, the question text Q1, and the option text O1. The three text sets obtained from the splitting are then concatenated to obtain [S1, Q1, O1]. By vectorizing the text content in the concatenated set one by one, [E] is obtained. i1 ,…,E ij ,…,E il ], where E ij This refers to the vector representation of the corresponding word. d e Here, l represents the word vector dimension, and l is the length of the input text content. After vectorizing all the text content in the reference text corpus, the set of vector representations [S1, Q1, O1] is translated by a machine translation model to obtain the converted text S. t Question Q tOption O t The text vector representation. Furthermore, in this embodiment of the invention, during the language conversion process, multiple reference texts in different languages can be selected for multilingual corpus fusion and splicing. That is, in this embodiment, the translated text vector representations in multiple different languages can be spliced together to obtain the input to the natural language processing model for the next step [S]. t Q t O t In this embodiment, the natural language processing model can be either the BERT model or the Transformer for semantic analysis. Taking the BERT model as an example, in this embodiment, [S] t Q t O t Inputting BERT yields vector representations of cross-linguistic knowledge. Integrating these vector representations results in a set of cross-linguistic corpora.
[0136] In some feasible embodiments, step S400 of performing a nonlinear transformation on the first fused corpus to obtain the second fused corpus in the embodiment method may include steps S410-S440:
[0137] S410. Determine the first weight matrix based on the corpus vector dimension of the first fused corpus and the corpus vector dimension of the second fused corpus.
[0138] The first weight matrix is the weight matrix corresponding to the concatenation matrix of the first corpus information and the cross-linguistic corpus. Specifically, in the first activation function layer of the CKFN model provided in the embodiment, the weight matrix W... (0) As a necessary parameter in the ReLU activation function, in the example, Where, d r It is the hidden layer dimension of the CKFN model, where d c =d D +d e d D For the vector dimension of cross-linguistic corpora, d e The first fused corpus is defined by its vector dimension; the hidden layer vector in the model serves as the vector representation of the second fused corpus in this embodiment. Furthermore, in the first activation function of this embodiment, the value of the weight matrix is determined based on the vector dimension of the first fused corpus, the vector dimension of the cross-linguistic corpus, and the vector dimension of the second fused corpus.
[0139] S420. Determine the first bias vector based on the corpus vector dimension of the second fused corpus;
[0140] Here, the first bias vector refers to the vector used in the nonlinear transformation process of the concatenated vector matrix in the first activation function layer. Its main function is to add translation capability to the nonlinear transformation. Specifically, in the first activation function layer of the CKFN model in the embodiment, the bias vector...
[0141] S430. Activate the first fused corpus according to the first activation function, the first weight matrix, and the first bias vector to obtain the first function value;
[0142] S440. Determine the second fused corpus based on the first function value and the information of the first corpus;
[0143] Specifically in the CKFN model, the first activation function can be the ReLU activation function, and its calculation process is as follows:
[0144] r = ReLU(W (0) ·[D;e]+b (0) )+D
[0145] Where r is the vector representation of the second fused corpus or the hidden layer vector of the CKFN model in the aforementioned embodiment, W (0) Let b be the weight matrix. (0) W is the bias vector. (0) and b (0) These are hyperparameters that have been determined during the training of the CKFN model. [;] represents vector concatenation operation. d c =d D +d e In other words, in the CKFN model, the process of nonlinear transformation in the first activation function layer firstly involves multiplying the vector representation of the concatenated first fused corpus with the weight matrix, then assigning a certain bias vector to the result of the multiplication, and finally performing a nonlinear transformation through the ReLU function. Finally, the result of the nonlinear transformation is summed with the candidate word vectors of the initial input to obtain the vector representation of the second fused corpus.
[0146] In some feasible embodiments, the step S600 of performing a second fusion splicing process on the second fused corpus and then a nonlinear transformation to obtain the second corpus information after multilingual knowledge fusion in the embodiment method may include steps S610-S640:
[0147] S610. Determine the second weight matrix based on the corpus vector dimension of the second fused corpus and the corpus vector dimension of the first corpus information;
[0148] Here, the second weight matrix is the weight matrix corresponding to the corpus vectors of the second fused corpus. Taking the CKFN model as an example, in the second activation function layer of the model, the weight matrix W...(1) As a necessary parameter in the ReLU activation function, in the embodiment, d h d represents the dimension of the final output vector of the CKFN model; as described in the preceding embodiments, d r This refers to the hidden layer dimension of the CKFN model. In the aforementioned embodiments, it was already indicated that the hidden layer vectors in the CKFN model serve as the vector representation of the second fused corpus in the embodiments, i.e., d. r It can also refer to the corpus vector dimension of the second fused corpus. Furthermore, in the second activation function of the embodiment, the value of the weight matrix is determined based on the corpus vector dimension of the first fused corpus, the vector dimension of the cross-language corpus, and the corpus vector dimension of the second fused corpus.
[0149] S620. Determine the second bias vector based on the vector dimension of the first corpus information;
[0150] Similar to the first bias vector, the second bias vector refers to the vector used in the nonlinear transformation process of the concatenated vector matrix in the second activation function layer. Its main function is to add translation capabilities to the nonlinear transformation. Specifically, in the second activation function layer of the CKFN model, the bias vector...
[0151] S630. Activate the second fused corpus according to the second activation function, the second weight matrix, and the second bias vector to obtain the second function value.
[0152] S640. The second corpus information is determined based on the second function value and the second fused corpus.
[0153] Specifically in the CKFN model, the second activation function can also be the ReLU activation function, and its calculation process is as follows:
[0154]
[0155] in, W refers to the vector representation of the information in the second corpus. (1) Let b be the weight matrix. (1) W is the bias vector. (1) and b (1) These are hyperparameters that have been determined during the training of the CKFN model. [;] represents vector concatenation operation. In other words, in the CKFN model, the process of nonlinear transformation in the second activation function layer first involves multiplying the corpus vector dimension of the concatenated second fused corpus with the weight matrix, then assigning a certain bias vector to the result of the multiplication, and finally performing a nonlinear transformation through the ReLU function. Finally, the result of the nonlinear transformation is summed with the input second fused corpus to obtain the second corpus information.
[0156] Refer to the instruction manual. Figure 5 The complete implementation process of the multilingual knowledge fusion method of this invention is described below:
[0157] In implementation scenarios where semantic analysis of the target object is required, such as Figure 5 As shown, since the text processing model is constructed by integrating a knowledge fusion model on top of the machine learning model architecture, the text processing model in this embodiment retains the basic hierarchical architecture of an input layer, a hidden layer (intermediate layer), and an output layer. In this embodiment, the knowledge fusion model, such as the CKFN model, can be integrated into various layers of the text processing model to perform fusion processing of corpus content from different languages. Furthermore, in this embodiment, the operation of the knowledge fusion model in fusing multilingual corpus data must be performed at least once during the text processing process; that is, in this embodiment, the text processing model after training integrates at least one knowledge fusion model. Exemplarily, the knowledge fusion model integrated in the text processing model of this embodiment can perform preliminary translation of text corpus data from different languages, and then fuse and concatenate the translated text content with the initial input text corpus to obtain cross-lingual knowledge content, which is then input into the next layer of the text processing model. It should be noted that, since cross-linguistic knowledge content is obtained by splicing the output of any level in the text processing model through the knowledge fusion model, the text granularity of the corpus content obtained from the output of any level in the text processing model in this embodiment of the invention should be consistent with the text content obtained from the translation in the knowledge fusion model.
[0158] Taking cross-linguistic knowledge content at the word granularity level as an example, in one embodiment, the word vector layer of the text processing model denoises the text to be processed, resulting in target text. Furthermore, the word sequence of the target text contains expressions such as "willing" and "pleasant." Before the word vector layer of the text processing model inputs the denoised text content to the next encoding layer, the embodiment can use a knowledge fusion model to translate and fuse corpora from other languages that have the same semantics and text granularity as expressions such as "willing" and "pleasant" into the output text content of the word vector layer. For example, the knowledge fusion model can translate the English expression "be willing to" into its word vector form and fuse it with the Chinese corpus content output by the word vector layer to obtain cross-linguistic knowledge content, which serves as the input to the next level of the text processing model.
[0159] It should be further noted that, since the knowledge fusion model in this invention can be integrated into different machine learning models, the text processing model infrastructure in this embodiment can adopt any model architecture capable of achieving text reading comprehension or semantic analysis, including but not limited to BERT model, LSTM model, and other neural network models. This embodiment does not limit this.
[0160] like Figure 6 The illustration shows a text processing method provided by an embodiment of the present invention, which can be implemented by... Figure 1 The text processing backend server 120 executes the process, or it executes on some terminals 110 with certain data processing capabilities, or it executes the process interactively between terminal 110 and text processing backend server 120. (See reference) Figure 6 The text processing method mainly includes steps S610-S630:
[0161] S610. Obtain the target text and generate the first word vector of all words in the target text;
[0162] In this embodiment, the target text can refer to the text content obtained by simple preprocessing of the text to be processed. The text to be processed refers to the initial text before processing, which can be pre-stored text data or text data acquired in real-time through, for example, dialogue. For instance, in the real-time acquisition method of this embodiment, the voice information of the target object is first acquired, and after necessary spectral amplification and speech-to-text processing, the corresponding text to be processed can be obtained. The text to be processed in this embodiment can include corpus content in different languages; for example, the text to be processed can contain corpus text in both Chinese and English. The preprocessing process of the text to be processed can include removing noise data from the text to be processed, such as removing obviously missing text content and incorrect punctuation. Alternatively, the target text can also refer to the text content obtained by preprocessing the text to be processed and then fusing the corpus content in different languages through a knowledge fusion model.
[0163] The first word vector in this embodiment can be obtained by segmenting the target text into words and then vectorizing the segmented word sequence. In an implementation that requires inserting a knowledge fusion model at the word vector stage, the target text is first segmented and vectorized to obtain candidate word vectors, which are then integrated to form a candidate word set. Then, multilingual corpus fusion is performed to construct the first word vector set. In the latter implementation, the word vectors in the first word vector set are all multilingual expression fusion word vectors.
[0164] Specifically in an embodiment, for target text input to a trained text processing model, it is first necessary to perform preliminary processing on the text corpus. Large chunks of text are gradually split to gradually obtain text content with smaller granularities such as text sentences and words; for example, word segmentation is performed on text sentences to obtain individual word texts. Illustratively, in an embodiment of the present invention, text paragraphs are decomposed into text sentences first according to tone marks or separation symbols, and then word segmentation is performed on the text sentences by means of traversal and matching based on a preset dictionary or vocabulary in the embodiment to obtain individual words. For example, a certain sentence obtained after splitting a text paragraph is "I am very willing to help you", according to the preset dictionary in the embodiment, part-of-speech tagging and splitting are performed on possible words in the sentence, and the obtained word set is {"I", "very", "willing", "help", "you"}; then vectorization encoding is performed on the words in the word set to obtain a vectorized representation of the word set.
[0165] In some other embodiments, multi-language corpus fusion processing can be performed in the word vector processing stage to add semantic subjects in the word vector processing stage, so as to improve the accuracy of the final prediction result. Furthermore, after obtaining the candidate word set through the aforementioned word segmentation and vectorization encoding in the embodiment, expressions in different languages can be subjected to word segmentation processing in the current language, and after translation and vectorization encoding, they are spliced and fused with the word vectors in the candidate word set. For example, for the sentence "I am glad to help you", the embodiment first performs word segmentation to obtain "I", "am", "glad", "to", "help", "you"; then the word segmentation results are translated one by one, for example, "help" is translated into "帮助", the knowledge fusion model in the embodiment can fuse and splice the word vector obtained after translating "help" and performing vectorization encoding with the vector representation of the aforementioned word "帮助" to obtain a vector representation after multi-language corpus fusion. The foregoing processes of translation, vectorization and fusion splicing are repeated to obtain a plurality of vector representations fused with multi-language corpora, and after integration, a first word vector set is obtained.
[0166] It should be further noted that, in the embodiments of the present invention, when a natural language processing model such as BERT is selected as the basic architecture of the text processing model in the embodiments, word segmentation and word vectorization can be performed through the word vector layer built into the model; and the first set of word vectors output by the word vector layer can be output in the form of a matrix of word vectors. However, when other neural network models are selected as the basic architecture of the text processing model in the embodiments, word segmentation and word vectorization can be performed by third-party word segmentation tools and part-of-speech tagging tools. For example, in the embodiments, the jieba word segmentation tool can be used to segment words in the target text.
[0167] S620. Extract feature relationships from the first word vector to obtain the text feature information of the target text;
[0168] In this embodiment, text feature information mainly refers to the feature relationships that can be used to describe the relationships between word vectors. Based on these relationship relationships, the classification prediction layer in the text processing model can predict the target text content. In this embodiment, text feature information includes, but is not limited to, similarity features between word vectors, contextual features between word vectors, etc. Alternatively, in other feasible embodiments, the feature information of cross-language word vectors can be fused and spliced with the feature information of the current language to obtain the feature information after multilingual corpus fusion as text feature information.
[0169] Specifically, in this embodiment, the word vectors in the first word vector set obtained in step S610 are input one by one into the (feature) encoding layer in the text processing model to extract feature relationships. Taking the implementation scenario of extracting contextual features between word vectors as an example, in this embodiment, the first word vector set is input as a text sequence into the feature encoding layer. The feature encoding layer reads all the word vectors in the text sequence at once, so that the text processing model of this embodiment can extract features based on both sides of the word vectors. For the contextual features of the word vectors, they are added to the corresponding word embedding tokens of the word vectors through encoding, and the target word vector sequence is output. In the target word vector sequence, each vector corresponds to a token with the same index.
[0170] In some other feasible implementations, after extracting the target word vector sequence, the embodiments can perform word segmentation, vectorization encoding, and feature encoding on reference text in other languages to output a reference word vector sequence with the same text granularity as the target word vector sequence; and the method of extracting feature information and the word embedding token encoding method are the same as the feature information extraction process in the aforementioned embodiments, and will not be repeated here. The cross-language reference word vector sequence is then fused and concatenated with the target word vector sequence. The resulting vector sequence has each vector corresponding to a token with the same index, and the corresponding word vector or target word can be determined based on the token.
[0171] S630. Classify and identify the text feature information to obtain the target result corresponding to the target text;
[0172] The target result refers to the answer output by the text processing model after inputting the target text, which corresponds to the semantics of the target text. The content of this answer can be determined according to the specific implementation scenario. For example, in the process of human-computer interaction question and answer, the target result may refer to the keywords that are omitted in the dialogue. The computer needs to determine the statement to be triggered in the next round of dialogue based on these keywords. Therefore, the dialogue content is output as the target text to the text processing model in the embodiment, and its output is the target result, which is the keywords that may be omitted in the dialogue content.
[0173] Specifically, in this embodiment, after the feature encoding layer (encoder), the text processing model uses a classifier consisting of a fully connected layer and a softmax function layer, also known as a classification prediction layer, to calculate the probability distribution of the word vectors encoded based on text feature information in step S330. This outputs the predicted answer distribution, which describes the probability of each target word vector being the final target result. It can be understood that the higher the probability value, the greater the likelihood that the word vector will be the final target word. In this embodiment, the word vector corresponding to the maximum probability value is selected as the target result. The target word is determined and output as the final target result based on the token with the same index corresponding to each vector.
[0174] For example, in the embodiment, after fusing and concatenating the cross-language reference word vector sequence with the target word vector sequence to obtain a cross-language vector sequence, the embodiment uses an attention mechanism to calculate the similarity value between each word embedding token, calculates the similarity score corresponding to the input word vector based on the weight matrix and performs linear mapping on the token, and after standardizing the similarity score, determines the target result by sorting the standardized scores from high to low.
[0175] It should be noted that the fusion process at any stage of the embodiments is achieved through the multilingual knowledge fusion method or the CKFN knowledge fusion model provided in the foregoing embodiments. Therefore, in the embodiments of the present invention, at least one of the target text, the first word vector, and the text feature information is obtained through multilingual knowledge fusion processing.
[0176] like Figure 7 As shown, in some feasible implementations, in order to improve the processing efficiency of the text processing model corresponding to the text processing method in the embodiment, the embodiment can perform necessary preprocessing on the target text before performing word vector processing. Therefore, in the step S610 of the present invention, after generating a candidate word set of all words in the target text, multilingual corpus fusion is performed on the candidate word set, the step S611-S616 may be included:
[0177] S611. Extract the text content of the text to be processed to obtain the first text, the second text, and the third text;
[0178] The first text, the second text, and the third text are three sets of different text materials obtained by decomposing the text materials in the text to be processed according to different partitioning rules. The partitioning rules in the embodiments may include, but are not limited to, partitioning the text according to the length of the text materials, or partitioning the text according to the number of subjects contained in the text materials. The text content refers to all the corpus materials contained in the text to be processed.
[0179] For example, in this embodiment, the text content of the text to be processed is extracted and segmented according to the length of the text. Based on the length of the corpus material in the text to be processed, for example, simple extraction and segmentation of text content is performed based on the punctuation marks contained in the material text; for example, text material containing many commas, periods, or ellipses is extracted to obtain the chapter text S, which is the first text; text material containing question marks is extracted to obtain the question text Q, which is the second text; and content is extracted based on fixed sequence numbers and commas to obtain the option O. Furthermore, in this embodiment, there can be a subordinate relationship between the chapter S, question Q, and option O; for example, a chapter may contain 5 questions, and each question may have 4 options.
[0180] S612. Based on the length of the first text paragraph in the first text, perform the first word segmentation on the first text to obtain the first word sequence;
[0181] S613. Based on the length of the second text of the question statement in the second text, perform second word segmentation on the second text to obtain the second word sequence;
[0182] S614, performing third word segmentation processing on the third text according to the third text length of the candidate answer in the third text to obtain a third word sequence;
[0183] wherein, the first text length, the second text length and the third text length are all lengths determined by text materials under various granularities in corresponding texts, that is, division granularities; for illustration, the first text length can be jointly determined by the number of text sentences contained in the passage S and the number of words contained in the text sentences; the second text length can be determined by the number of words contained in the question Q; similarly, the third text length can be determined by the number of words contained in the option O.
[0184] Specifically in the embodiment, after determining the corresponding division granularity in each text, word segmentation processing is performed on each text respectively to obtain a word sequence under word granularity.
[0185] For example, passage S is a passage containing a series of sentences [s1,…,s i ,…s n , wherein s i is the i-th sentence in the passage, and n represents the number of sentences in the passage; firstly, simple clause processing is performed according to punctuation marks in the passage to obtain w ij represents the j-th word in the i-th sentence in the passage, l i is the length corresponding to this sentence. After obtaining words, more detailed decomposition can be performed through the input layer of a text processing model or a word vector layer. For example, for Chinese, it will be split into characters, and for English, it will be split into word units; for example, "hello你好" will be split into "he#llo你好".
[0186] For another example, the question Q in the embodiment can be processed by word segmentation to obtain wherein, q k represents the k-th word in the question, l q is the length of the question.
[0187] For another example, the option O in the embodiment can also be processed by word segmentation to obtain o m represents the m-th word in the option, l o is the length of the answer.
[0188] S615, determining a candidate word set according to the first word sequence, the second word sequence and the third word sequence;
[0189] The candidate word set mainly refers to the collection of all word material texts obtained through word segmentation before multilingual corpus fusion in the word vector processing stage. Specifically, in this embodiment, the three word sequences obtained from word segmentation in steps S322-S324 are integrated to obtain this candidate word set. It should be noted that during the integration process in this embodiment, necessary processing steps can be used to remove noise data present in the three word sequences. For example, punctuation marks and obvious errors left over from word segmentation can be removed from each word sequence.
[0190] S616. The candidate word set is fused with multilingual corpora using a knowledge fusion model to obtain the first word vector set.
[0191] Specifically, in this embodiment, at the text granularity of words, multilingual corpus fusion is performed on the candidate word set and the cross-lingual word set. Choosing to fuse and concatenate using vectors can reduce the computational resource consumption of the text processing model to a certain extent, making the fusion process more efficient. In this embodiment, the text of each word in the candidate word set is input to the word vector layer of the text processing model. Taking the BERT model as the basic architecture of the text processing model as an example, in this embodiment, the passage sequence S, the question sequence Q, and the option sequence O are first concatenated to obtain the input candidate word set denoted as [S,Q,O]. [S,Q,O] is input to the word vector layer. The text processing model can look up the vocabulary for each word in the candidate word set according to the pre-stored vocabulary in the word vector layer to obtain the vector representation corresponding to each word, thus obtaining [E...]. i1 ,…,E ij ,…,E il ], where E ij This refers to the vector representation of the corresponding word. d e 'l' represents the word vector dimension, and 'l' represents the length of the input candidate word set. Furthermore, the knowledge fusion model translates reference text that differs from the candidate word set in language, and then performs word segmentation and vectorization encoding on the translated reference text to obtain its word vector set. Finally, a full connection is established between the vectorized candidate word set and the vectorized reference text set to obtain the cross-language fused knowledge text, i.e., the first word vector set.
[0192] In some feasible implementations, the step S616 of fusing multilingual corpora to obtain the first word vector set in this embodiment of the invention may include steps S6161-S6162:
[0193] S6161. Vectorize the candidate words to obtain candidate word vectors;
[0194] Specifically in an embodiment, all word texts in the candidate word set input into the model may be subjected to vectorization processing through the word vector layer in a text processing model, to obtain corresponding word vectors, and the word vectors are integrated to obtain a candidate word vector set. In an embodiment, when the text processing model adopts a natural language processing model such as Bert as a basic architecture, the word segmentation and word vector processing processes can be executed through the word vector layer in the architecture. For example, after the word tokenizer built in the word vector layer performs word segmentation according to a word list or dictionary to obtain a candidate word set, the word vector layer further performs matching according to the word list to obtain the word vector corresponding to each candidate word. When the text processing model adopts the basic architecture of other machine learning models, other tools, such as word2vec, can also be used to measure the words from aspects such as part of speech, emotional color, degree, etc., and use a set of values to represent a word, so as to determine the word vector representation of the candidate words.
[0195] For example, the embodiment may obtain a reference text from stored historical text data, or obtain a reference text from other open source databases; the text content of the obtained reference text shall be different from the target text in the embodiment in terms of the expression language of at least a single word. For example, if the target text is selected as text content in Chinese, the corresponding reference text shall be selected as text in English, Japanese or other languages as the reference text. For another example, there may be expressions in both Chinese and English in the target text, so when selecting the reference text, there is a difference in the expression of at least a single word in the reference text; if there are words or phrases "乐意" and "be glad to" in the target text, the corresponding reference text has the French expression "volontiers". After the reference text corresponding to the target text is constructed, the reference text is first translated by a machine translation tool, so that the language of the reference text is consistent with the language of the target text. Then, the translated reference text is also subjected to word segmentation and vectorization processing to obtain a cross-language vector set with the same text granularity as the candidate word vector set. In the embodiment, the reference text can be processed in the same processing manner as that for the target text, which will not be repeated herein.
[0196] S6162, performing multi-lingual corpus fusion on candidate word vectors to obtain first word vectors;
[0197] The multilingual corpus fusion process refers to the process of fusing and splicing reference text corpora that have been translated and have the same text granularity as the target text corpus. For example, the aforementioned candidate word vector set and the reference text vector set are spliced together using a fully connected method, and the resulting vector set is the first word vector set. In some other feasible implementations, after fully connecting the two sets, a modified first word vector can be obtained through nonlinear transformation or activation function operations.
[0198] In embodiments of the present invention, based on the hot-swappable nature of the knowledge fusion model, which can be integrated into any layer of the text processing model, in some other implementations, after extracting feature relationships from the text materials in the target text, multilingual corpus fusion can be performed to obtain cross-lingual text feature information. Furthermore, in embodiments of the present invention, the step S620 of extracting feature relationships from the first word vector to obtain the text feature information of the target text may include steps S621-S623:
[0199] S621. Extract feature relationships from the first word vector to generate a feature vector;
[0200] Similar to text feature information, feature relationships can be used to describe the semantic information existing between word vectors in the first word vector set. This semantic information may include, but is not limited to, the positional information of word vectors in the text sequence, the contextual information of word vectors in the text sequence, and the similarity between word vectors, etc. Specifically, in the embodiment, the feature extraction layer or feature encoding layer in the text processing model structure determines which feature to use for semantic analysis of word vectors and to express the feature information during the feature relationship extraction process, outputting the corresponding feature vector.
[0201] For example, the embodiment can use a convolutional neural network as the basic architecture layer of the text processing model. Specifically, the embodiment extracts features from word vectors in the first word vector set through a convolutional layer in the feature extraction layer, where the convolutional kernel represents the weight of the corresponding word vector; the convolutional kernel is then multiplied by the word vector to output a feature vector. Alternatively, in some other implementations, a natural language processing model such as Transformer can be used. The encoder layer in the Transformer encodes the position of the word vector in the text sequence and encodes it based on an attention mechanism to obtain an encoding matrix. The feature relationship described by this encoding matrix is the same as that of the feature vector set.
[0202] S622. Perform multilingual corpus fusion on the feature vectors to obtain the second word vector;
[0203] Specifically, in this embodiment, a set of word vectors for the language-converted reference text is obtained. During the feature relationship extraction process, the reference text is segmented, vectorized, and translated to obtain a second set of word vectors. In this embodiment, feature relationships between the second word vectors can be extracted based on the obtained set of second word vectors. It should be noted that, for ease of subsequent fusion processing, the method for extracting feature relationships from the second word vectors can be the same as the method for extracting feature relationships in general.
[0204] For example, when a convolutional neural network is used as the basic architecture layer of the text processing model in the embodiment, the word vectors in the second cross-linguistic vector are also input one by one into the feature extraction layer, and the feature vector of the second cross-linguistic vector is obtained by performing convolution operation through convolution kernels. In the embodiment, in the feature information extraction stage, the knowledge fusion model can also adopt the CKFN model in the aforementioned embodiment; specifically, the feature vectors corresponding to the first set of word vectors and the feature vectors corresponding to the second set of word vectors are input into the CKFN model, and fully connected processing is performed through two layers of fully connected layers, and nonlinear transformation is performed through two layers of activation functions to output the second word vector; this second word vector is the vector after cross-linguistic knowledge fusion; and due to the aforementioned feature extraction and encoding process, the output second word vector also contains the feature relationship of the corresponding word vectors.
[0205] S623. Determine text feature information based on the second word vector;
[0206] Specifically, in the embodiment, all the second word vectors output in step S333 are integrated to obtain a second word vector set. Since the second word vectors can be used to describe the feature relationship of the corresponding word vectors, the integrated second word vector set contains the text feature information corresponding to each word vector under different languages.
[0207] In some feasible embodiments, to improve the processing efficiency of the text processing model and reduce the consumption of computing resources, the text processing model provided in the embodiments can be built on the basic architecture of the BERT model. That is, when extracting feature relationships from the word vectors of the target text and the reference text in the embodiments, the focus is mainly on calculating and extracting the contextual relationships of the word vectors. Furthermore, the step S621 of extracting feature relationships from the first word vector and generating feature vectors in the embodiments of the present invention can include S6211-S6213:
[0208] S6211. Determine the context relationship based on the information carried by the adjacent position vectors of the first word vector;
[0209] The adjacent position vectors can include at least one of the vectors of the first word vector at the first few positions or the vectors of the first word vector at the last few positions. In the embodiments, the contextual features of the word vectors are dynamic, that is, different word vectors express different semantics in different text sequences, depending on the context around the word vector; for example, in the text sequence "one bird was flying below another bird", the semantics expressed by the word vector corresponding to the word "bird" will be different, and "one", "flying", "below", and "another" all belong to the context of "bird".
[0210] S6212. Encode the context relationship using features to determine several encoded vectors corresponding to different adjacent position vectors;
[0211] Specifically, in this embodiment, the first set of word vectors is input as a complete text sequence into the trained text processing model. The encoder in the model learns the context of each word vector through a multi-head attention mechanism, and outputs the context content as a vector encoding with the word vector set to obtain the embedding vector corresponding to that word vector, which is the encoded vector. It should be noted that the process of context learning by the encoder based on the multi-head attention mechanism in this embodiment can adopt the mature implementation method of multi-head attention mechanism in related technical solutions, and therefore will not be elaborated here.
[0212] For example, the target text is in English, and the text content corresponding to the first input word vector is "May I help you". The word vectors obtained after word segmentation and vectorization are input into the text processing model in this embodiment. The model outputs the context embedding vector representation of each word in the text content. The text processing model can stack several encoders. Therefore, in this embodiment, the embedding vector R... I The word vector representing the word "I" can represent the size of each encoder layer; for example, the size of each word vector should be the size of each encoder layer.
[0213] S6213. Integrate the various encoding vectors to obtain a set of feature vectors;
[0214] Specifically, in this embodiment, the text sequence containing the first word vector, i.e. the first word vector set, is input into the text processing model in this embodiment, and the encoded vectors output by the model are integrated to obtain a feature vector set.
[0215] It should be noted that, in this embodiment, before the final answer distribution is calculated through the softmax function layer, the text processing model also includes at least one Decoder layer, which is mainly used to decode and translate the encoded information matrix output by the encoding layer during the feature extraction process.
[0216] The following is in conjunction with the instruction manual appendix. Figure 4 The following is a general description of the complete implementation process of the embodiments of the present invention, along with specific application scenarios:
[0217] Taking a conversational human-computer interaction scenario as an example, in this implementation scenario, the machine or computer needs to interact with the target object through dialogue. During the dialogue, the machine or computer needs to perform semantic analysis on the target object's real-time expressions and provide corresponding feedback based on the analysis results. For example, in a certain dialogue scenario, target object A: "What is the relationship between Tom and Jerry?" Computer B: "Tom and Jerry are good friends." Target object A: "Why are they good friends?" In this scenario, the computer needs to perform semantic analysis on the target object's latest question. The analysis process includes, but is not limited to, determining that "they" in the statement refers to "Tom and Jerry," and learning the expressions of "Tom and Jerry" in different languages, such as "Tom and Jerry," and providing the answer to the corresponding question through cross-language knowledge fusion analysis.
[0218] In this implementation scenario, the text processing model provided in this embodiment of the invention can be used to perform cross-language knowledge fusion analysis. The text processing model in the embodiment is a model obtained by integrating a knowledge fusion model on the basis of the Bert model architecture. The knowledge fusion model can perform fusion processing of corpus content of different languages at various levels of the text processing model. The knowledge fusion model can be integrated into any layer of the text processing model architecture.
[0219] In this embodiment, all dialogue content in the scenario, or pre-set dialogue content in the computer, is input as target text into the text processing model in this embodiment. In the text processing model, the target text is first decomposed into sections such as chapters S, questions Q, and options O. For example, the pre-set dialogue content in the computer consists of chapters, questions, and options. The target text contains a chapter, which may contain 5 pre-set questions, each with 4 pre-set answer options. Chapter S comprises a series of sentences [s1,…,s…]. i ,…s n ], where s i It is the i-th sentence in the passage, where n represents the number of sentences in the passage. w ijThe j-th word in the i-th sentence of the passage; l i This refers to the length of the sentence. (Question) q k L represents the k-th word in the question. q This refers to the length of the question. Options o m l represents the m-th word in the options. o This refers to the length of the answer. During the prediction of the target result, the passage S, question Q, and option O are concatenated to obtain the input [S,Q,O], which is then processed through a word vector layer to obtain [E]. i1 ,…,E ij ,…,E il ], where E ij This refers to the vector representation of the corresponding word. d e is the word vector dimension, and l is the length of the input. Then, the encoder in the BERT model framework outputs the context vector representation C of the entire input. d h This is the dimension of the BERT output vector. Finally, C passes through a classifier consisting of fully connected layers and softmax layers to obtain the predicted answer distribution.
[0220] In this embodiment, the knowledge fusion model can be integrated into different layers of the text processing model, such as the word vector layer, encoding layer, and classifier layer. At the word vector layer, the knowledge fusion model converts the original word vectors E... ij By incorporating cross-linguistic knowledge and information, a new word vector representation is obtained. At the classifier layer, the knowledge fusion model outputs to the encoder of the text processing model, that is, cross-linguistic knowledge information is incorporated into the context representation C of the entire input, resulting in a corrected context representation. Finally use The original C is used as the input to the classifier to obtain the predicted answer.
[0221] More specifically, in the knowledge fusion model of this embodiment, the model's input is divided into two parts: cross-lingual knowledge representation E and text representation D. The cross-lingual knowledge representation E is obtained through machine translation and initialization with another BERT model; while D can be text representations at various granularities, such as word representations or document representations. When the knowledge fusion model is integrated into the word vector layer of the text processing model, D = E. ij In the knowledge fusion model at the classifier layer, D = C. E and D have the same granularity; for example, both are word-level or document-level. In the knowledge fusion model, the model first concatenates the two input vectors, then passes them together through the first fully connected layer, followed by a ReLU nonlinear transformation layer, to obtain... d rThis is the hidden layer dimension of the knowledge fusion model. The formula for calculating vector r is as follows:
[0222] r = ReLU(W (0) ·[D;E]+b (0) )+D
[0223] Where [;] denotes vector concatenation operation. W represents the weight matrix, and b represents the bias vector. d c =d D +d e d e The dimension of a word vector. Then, vector r passes through a second fully connected layer and a ReLU nonlinear transformation layer to finally obtain the corrected text representation. Among them, text representation The calculation formula is as follows:
[0224]
[0225] in, d h This refers to the dimension of the output vector from the knowledge fusion model. Finally, the contextual features obtained after knowledge fusion post-processing are used as input to the classifier to obtain the predicted answer. The answer corresponds to the dialogue content in the implementation scenario, and its output can be "Tom and Jerry".
[0226] like Figure 8 As shown, the process of acquiring cross-linguistic knowledge representation E in the knowledge fusion model can be obtained through the output of machine translation and BERT models. First, a reference text with linguistic differences from the target text is obtained. Similarly, the passage S1, question Q1, and option O1 in the reference text are concatenated to obtain the input [S1, Q1, O1]. This input is then processed through word vectorization and vectorization to obtain [E]. i1 ,…,E ij ,…,E il ], where E ij This refers to the vector representation of the corresponding word. d e Here, l represents the word vector dimension, and l is the length of the input. Inputting [S,Q,O] into the machine translation model yields texts S in different languages. t Question Q t Option O t Text representation. Then, the text representations from different languages are concatenated to obtain the input [S]. t Q t O t ], will [S t Q t O tInputting BERT yields the cross-language knowledge representation E.
[0227] It should be noted that the application scenarios of the solutions provided in the embodiments of the present invention are not limited to the scenarios described above and below. In addition to the technical field of human-computer interaction, they can also be applied to other technical fields such as target object profiling, cloud technology, and speech recognition processing.
[0228] like Figure 9 As shown, in another aspect, this invention also provides a multilingual knowledge fusion model training method, the method including T100-T500:
[0229] T100: Obtain training corpus information and determine the target word vector in the training corpus information;
[0230] The training corpus can include corpus content in different languages. It can be text content obtained through simple preprocessing of historical texts. In this embodiment, the training corpus primarily serves as reference text in different languages for the corpus material in the first training text. Specifically, in this embodiment, the training corpus content is denoised and cleaned as necessary. Then, larger sections of text are broken down into sentence texts based on punctuation marks. After obtaining the sentence texts, further word segmentation is performed to obtain target words. The target word vector corresponding to the target word is determined through word list or dictionary matching.
[0231] T200. Perform a first fusion and splicing process on the different languages in the training corpus to obtain the first fused corpus;
[0232] Specifically, in the embodiment, through step T100, word vectors with the same meaning but different languages in the training corpus information are input into the first fully connected layer in the text fusion model. The fully connected layer performs fully connected processing on word vectors with the same meaning but different languages, and outputs the first fused corpus.
[0233] T300: The second fused corpus is obtained by performing a nonlinear transformation on the first fused corpus;
[0234] Specifically, in the embodiment, the fusion vector output by the first fully connected layer in the knowledge fusion model, i.e. the first fusion corpus, is input into the first activation function layer in the model. The first fusion corpus is then transformed nonlinearly by the activation function to output the second fusion corpus.
[0235] T400: Perform second fusion splicing processing on the second fused corpus and perform nonlinear transformation to obtain the predicted word vectors after multilingual knowledge fusion;
[0236] Specifically, in this embodiment, to improve the fusion processing effect of the knowledge fusion model and enhance the prediction accuracy of the overall model, i.e., the text processing model, this embodiment of the invention adopts a two-layer fully connected layer and activation function layer model architecture in the knowledge fusion model. Therefore, after the knowledge fusion model obtains the second fused corpus through the first nonlinear transformation, the fused vector is further input into the second fully connected layer, where the second fused corpus undergoes self-connection to output the third fused corpus. After obtaining the third fused corpus, the fused vector is input into the second activation function layer in the knowledge model, where a nonlinear transformation is performed to output the final candidate word vectors for knowledge fusion, which are then integrated to obtain the predicted word vectors.
[0237] T500: Adjust the parameters of the multilingual knowledge fusion model based on the loss value between the predicted word vector and the target word vector;
[0238] Specifically, in the embodiment, during the training phase of the text processing model, the knowledge fusion model is also trained simultaneously. The candidate word vector set obtained from the output of steps T100-T400 is input into the classifier in the text processing model to obtain the prediction results during the training process. The predicted word vectors corresponding to the prediction results are compared with the original word vectors in the second training text, i.e., the target word vectors, to determine the second loss value between the two vectors.
[0239] It should be noted that during the training phase of the text processing model and knowledge fusion model in this embodiment of the invention, residual connections can be added to the model according to specific needs. By incorporating residual connections, the gradient vanishing or gradient explosion problems that accompany the increase of network depth can be reduced, thereby improving the final prediction performance of the model. Furthermore, during training, layers or units in the model can be removed with a certain probability to obtain a simpler model structure, thereby reducing the computational resources occupied by the model and improving both processing efficiency and robustness.
[0240] The following is a general description of the complete training process of the text processing model in this embodiment of the invention, using specific application scenarios as examples:
[0241] In the field of vehicle networking technology, in-vehicle terminals typically need to have a built-in natural language processing model to recognize the voice commands of the target object and execute corresponding actions. To cope with more complex language usage environments, the in-vehicle terminal can incorporate the text processing model provided in this embodiment. The training process of this text processing model is as follows: First, the historical command content or text content stored in the terminal server is preprocessed, such as by denoising, to form training text, which is then input to the text processing model to be trained. In the text processing model, the text structure of the training text is broken down into sections such as passages, questions, and options. These sections are then concatenated to obtain the model input [S,Q,O]. This input is then vectorized through the word vector layer in the text processing model to obtain [E]. i1 ,…,E ij ,…,E il ], where E ij This refers to the vector representation of the corresponding word. d e is the word vector dimension, and l is the length of the input. Then, the encoder in the text processing model outputs a context vector representation of the entire input. This context vector representation is then processed by a classifier consisting of fully connected layers and softmax layers to obtain the predicted answer distribution.
[0242] During training, the knowledge fusion model in this embodiment can be integrated into different layers of the text processing model, such as the word vector layer, encoding layer, and classifier layer. At the word vector layer, the knowledge fusion model modifies the original word vectors E... ij By incorporating cross-linguistic knowledge and information, a new word vector representation is obtained. At the classifier layer, the knowledge fusion model outputs to the encoder of the text processing model, that is, cross-linguistic knowledge information is incorporated into the context representation C of the entire input, resulting in a corrected context representation. Finally use The original C is used as the input to the classifier to obtain the predicted answer.
[0243] After obtaining the predicted answer, the loss is calculated by comparing it with the corresponding answer content in the training text, such as the options in the training text or the real answer in the training text. The parameters of the text processing model are then continuously tuned based on the calculated loss value until the loss value converges, and finally the trained text processing model is obtained.
[0244] In another aspect, this invention also provides a text processing model training method, the method comprising steps R100-R500:
[0245] R100: Obtain the training text and determine the target result in the training text;
[0246] The training text refers to stored training text data, which may include text content from several dialogue interactions. The text statements that respond to questions or other text requiring feedback during these interactions can serve as the standard results in the embodiments. The training text may include corpus content in different languages. The first training text can be text content obtained by performing simple preprocessing on the training text. The preprocessing mainly involves removing noisy data from the text to be trained.
[0247] Specifically, in this embodiment, the text processing model architecture to be trained integrates a knowledge fusion model. The knowledge fusion model is primarily used to fuse different language corpora at various levels of the text processing model. The knowledge fusion model can be integrated into any layer of the text processing model architecture. In this embodiment, the knowledge fusion model can perform preliminary translation of text corpora in different languages, and then fuse and concatenate the translated text content with the initial input text corpora to obtain cross-lingual knowledge content. This cross-lingual knowledge content is then input into the next level of the model.
[0248] R200: Generate the first word vectors for all words in the training text;
[0249] Specifically, in this embodiment, for the target text input to the text processing model to be trained, the training text first needs to undergo preliminary text corpus processing. This involves progressively breaking down large sections of text into smaller, more granular text content such as sentences and words. For example, sentences are segmented into individual words. Exemplarily, in this embodiment, text paragraphs are first decomposed into sentences based on tone marks or delimiters. Then, using a pre-defined dictionary or vocabulary, the sentences are segmented into individual words, which are then vectorized to obtain a first word vector.
[0250] In some other implementations, multilingual corpus fusion processing can be performed during the word vector processing stage to incorporate semantic subjects and improve the accuracy of the final prediction results.
[0251] R300: Extract feature relationships from the first word vector to obtain the text feature information of the target text;
[0252] The text feature information mainly refers to the feature relationships that can be used to describe the various word vectors in the training text. Based on these feature relationships, the target text content can be predicted through the classification prediction layer in the text processing model. Specifically, in this embodiment, the word vectors obtained in the previous step are input one by one into the encoding layer of the text processing model. The feature encoding layer in the text processing model reads all the word vectors in the text sequence at once, so that the text processing model in this embodiment can extract features based on both sides of the word vectors. For the contextual features of the word vectors, they are added to the corresponding word embedding tokens of the word vectors through encoding, and the target word vector sequence is output. In this target word vector sequence, each vector corresponds to a token with the same index.
[0253] In some other feasible implementations, after extracting the target word vector sequence, the embodiment can use a knowledge fusion model to perform word segmentation, vectorization encoding, and feature encoding on other training texts in other languages to output a reference word vector sequence with the same text granularity as the target word vector sequence. Furthermore, the method for extracting feature information and the word embedding token encoding method are the same as the feature information extraction process in the aforementioned embodiment, and will not be repeated here. The cross-language reference word vector sequence is then fused and concatenated with the target word vector sequence. The resulting vector sequence can determine its corresponding word vector or target word.
[0254] R400 classifies and identifies text feature information to obtain the prediction result corresponding to the target text.
[0255] The prediction result refers to the answer that the text processing model predicts during the training phase after inputting the first training text, which satisfies the current semantics. Specifically, in the embodiment, after the feature encoding layer (encoder), the text processing model uses a classifier composed of a fully connected layer and a softmax function layer to calculate the probability distribution of word vectors encoded based on text feature information, and outputs the predicted answer distribution. This answer distribution can be used to describe the probability of each target word vector as the predicted result output by the model.
[0256] R500: Adjust the parameters of the text processing model based on the loss value between the predicted result and the target result;
[0257] Specifically, in this embodiment, the predicted result obtained during training is compared with the true answer (i.e., the standard result) in the current semantic environment of the first training text. The loss value between the predicted answer and the true answer is calculated. Before the loss value fully converges, the parameters in the text processing model are adjusted until the loss value converges. It should be noted that the loss function used in the loss value calculation process in this embodiment includes, but is not limited to, the mean squared error loss function, the Euclidean distance loss function, and the Manhattan distance loss function, etc. Furthermore, the calculation process of the loss function in this embodiment is not limited.
[0258] During the training of the text processing model, at least one of the training text, the first word vector, and the text feature information is obtained by the aforementioned multilingual knowledge fusion method. The fusion process will not be elaborated here.
[0259] like Figure 10 As shown, in another aspect, the present invention also provides a multilingual knowledge fusion device, which includes:
[0260] The first module 1001 is used to obtain information from the first corpus.
[0261] The second module 1002 is used to obtain reference text corpus that differs in language from the information in the first corpus, and to determine cross-language corpus based on the reference text corpus;
[0262] The third module 1003 is used to perform a first fusion and splicing process on the first corpus information and the cross-language corpus to obtain the first fused corpus.
[0263] The fourth module 1004 is used to perform nonlinear transformation on the second fused corpus after second fusion splicing processing to obtain the second corpus information after multilingual knowledge fusion;
[0264] The fifth module 1105 is used to perform a second fusion splicing process on the second fused corpus and then perform a nonlinear transformation to obtain the second corpus information after multilingual knowledge fusion.
[0265] It should be noted that the device provided in the above embodiments is only illustrated by the division of the above functional modules when running the application. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the multilingual knowledge fusion model training method embodiment, which will not be repeated here.
[0266] like Figure 11As shown, in another aspect, the present invention also provides a text processing apparatus, the apparatus comprising:
[0267] The sixth module 1101 is used to obtain the target text and generate the first word vector of all words in the target text;
[0268] The seventh module 1102 is used to extract feature relationships from the first word vector to obtain the text feature information of the target text;
[0269] Module 8, 1103, is used to classify and identify text feature information to obtain the target result corresponding to the target text.
[0270] Specifically, in the embodiments, the sixth module 1101 performs word segmentation on the text to be processed input to the embodiment device to obtain the target text at the word granularity; or, after denoising the text to be processed, a knowledge fusion model is used to translate and fuse corpora in other languages that have the same semantics and text granularity as the target text to obtain the target text. The sixth module 1101 further decomposes the obtained target text step by step to obtain the word set with the smallest text granularity, and vectorizes each word in the word set to obtain the first word vector set; or, during the vectorization process, a knowledge fusion model is used to obtain the word vectors of other languages after translation, and then the word vectors of the two different languages are fused to construct the first word vector set. The seventh module 1102 in the embodiments inputs each word vector in the first word vector set output by the sixth module 1101 into the text processing model to extract feature relationships to obtain text feature information; or, after extracting feature relationships, a knowledge fusion model is used to obtain the feature relationships between the word vectors of other languages after translation, and then the feature relationship expressions of the two different languages are fused to obtain text feature information. Finally, in the embodiment, the eighth module 1103 predicts the answer distribution of the prediction result based on the input text feature information, and uses the word with the highest probability in the answer distribution as the answer content corresponding to the target text.
[0271] like Figure 12 As shown, the technical solution of this application provides an electronic device that can implement the aforementioned multilingual knowledge fusion method, text processing method, multilingual knowledge fusion model training method, and text processing model training method. This device includes, but is not limited to, smartphones, tablets, MP4 (Moving Picture Experts Group AudioLayer IV) players, laptops, or desktop computers. The electronic device may also be referred to as a mobile device, portable terminal, laptop terminal, desktop terminal, or other names.
[0272] Typically, electronic devices include a processor 1201 and a memory 1202. The processor 1201 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 1201 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1201 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state.
[0273] In some embodiments, processor 1201 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1201 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0274] The memory 1202 may include one or more computer-readable storage media, which may be non-transitory. The memory 1202 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1202 are used to store at least one instruction, which is executed by the processor 1201 to implement at least one of the text processing method and text processing model training method provided in the method embodiments of this application.
[0275] In some embodiments, the electronic device may also optionally include: a peripheral device interface 1203 and at least one peripheral device. The processor 1201, memory 1202, and peripheral device interface 1203 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1203 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1204, a display screen 1205, a camera assembly 1206, an audio circuit 1207, a positioning assembly 1208, and a power supply 1209.
[0276] This invention also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform a multilingual knowledge fusion method, a text processing method, a multilingual knowledge fusion model training method, and a text processing model training method.
[0277] In summary, the text processing method for cross-linguistic knowledge fusion for reading comprehension provided by this invention can explicitly incorporate cross-linguistic knowledge, thereby effectively improving the accuracy of reading comprehension. The knowledge fusion model provided in this solution, due to its simple structure and flexible pluggable nature, can be applied to any natural language processing task and model that requires the fusion of cross-linguistic knowledge.
[0278] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this invention are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and sub-operations described as part of a larger operation are executed independently.
[0279] Furthermore, although the invention has been described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the described functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the invention. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional skill of an engineer. Therefore, those skilled in the art can implement the invention as set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of the invention, which is determined by the full scope of the appended claims and their equivalents.
[0280] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0281] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0282] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0283] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0284] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0285] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
[0286] The above is a detailed description of the preferred embodiments of the present invention, but the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A multilingual knowledge fusion method, characterized in that, include: Obtain information from the first corpus; Obtain reference text corpus that differs in language from the first corpus information, determine cross-language corpus based on the reference text corpus, the cross-language corpus being obtained by segmenting and vectorizing the translated reference text corpus, and having the same text granularity as the first corpus information, the language of the translated reference text corpus being consistent with the language of the first corpus information; The first corpus information and the cross-language corpus are subjected to a first fusion and splicing process to obtain the first fused corpus; The first fused corpus is subjected to a nonlinear transformation to obtain the second fused corpus; After performing a second fusion splicing process on the second fused corpus, a nonlinear transformation is performed to obtain the second corpus information after multilingual knowledge fusion. Wherein, the first fusion splicing process is used to represent the processing of the first corpus information and the cross-language corpus after structural splicing and passing through the first fully connected layer, and the second fusion splicing process is used to represent the processing of the second fused corpus through the second fully connected layer.
2. The multilingual knowledge fusion method according to claim 1, characterized in that, The step of acquiring reference text corpus that differs in language from the first corpus information, and determining cross-language corpus based on the reference text corpus, includes: The reference text corpus is converted to a different language to obtain cross-language text; Reference words are determined from the reference text corpus, and word vector processing is performed on the reference words to obtain cross-language word vectors; The feature relationships of word vectors in the cross-language word vectors are extracted to obtain cross-language text feature corpus; The cross-language corpus is constructed based on the cross-language text, the cross-language word vectors, and the cross-language text feature corpus.
3. The multilingual knowledge fusion method according to claim 1, characterized in that, The step of performing a nonlinear transformation on the first fused corpus to obtain the second fused corpus includes: The first weight matrix is determined based on the vector dimension of the first fused corpus and the vector dimension of the second fused corpus. The first bias vector is determined based on the corpus vector dimension of the second fused corpus; The first fused corpus is activated based on the first activation function, the first weight matrix, and the first bias vector to obtain the first function value; The second fused corpus is determined based on the first function value and the first corpus information.
4. The multilingual knowledge fusion method according to claim 1, characterized in that, The second fused corpus is then subjected to a second fusion and splicing process followed by a nonlinear transformation to obtain the second corpus information after multilingual knowledge fusion, including: The second weight matrix is determined based on the corpus vector dimension of the second fused corpus and the corpus vector dimension of the first corpus information; The second bias vector is determined based on the vector dimension of the first corpus information; The second fused corpus is activated based on the second activation function, the second weight matrix, and the second bias vector to obtain the second function value; The second corpus information is determined based on the second function value and the second fused corpus.
5. A text processing method, characterized in that, include: Once the target text is obtained, the first word vector of all words in the target text is generated; Feature relationships are extracted from the first word vector to obtain the text feature information of the target text; The text feature information is classified and identified to obtain the target result corresponding to the target text; Wherein, at least one of the target text, the first word vector, and the text feature information is obtained after multilingual knowledge fusion processing by the multilingual knowledge fusion method as described in any one of claims 1-4.
6. The text processing method according to claim 5, characterized in that, The first word vector is obtained by multilingual corpus fusion processing using a multilingual knowledge fusion method. Generating the first word vector for all words in the target text includes: The text to be processed is extracted to obtain the first text, the second text, and the third text. Based on the first text length of the text paragraphs in the first text, the first text is processed by first word segmentation to obtain the first word sequence; Based on the second text length of the question statement in the second text, the second text is processed by second word segmentation to obtain a second word sequence; Based on the length of the third text of the candidate answers in the third text, the third text is processed by third word segmentation to obtain a third word sequence; Candidate words are determined based on the first word sequence, the second word sequence, and the third word sequence; The candidate words are fused using multilingual corpora to obtain the first word vector.
7. The text processing method according to claim 6, characterized in that, The step of fusing multilingual corpora to obtain the first word vector includes: The candidate words are vectorized to obtain candidate word vectors; The candidate word vectors are fused using multilingual corpora to obtain the first word vector.
8. The text processing method according to claim 5, characterized in that, The text feature information is obtained through multilingual corpus fusion processing using a multilingual knowledge fusion method. The step of extracting feature relationships from the first word vector to obtain the text feature information of the target text includes: Extract feature relationships from the first word vector to generate a feature vector; The feature vectors are fused with multilingual corpora to obtain the second word vectors; The text feature information is determined based on the second word vector.
9. The text processing method according to claim 8, characterized in that, The feature relationships include contextual relationships; the step of extracting feature relationships from the first word vector to generate feature vectors includes: The contextual relationship is determined based on the information carried by the adjacent position vectors of the first word vector; the adjacent position vectors include at least one of the vectors of the first word vector at the first few positions or the vectors of the first word vector at the last few positions. The contextual relationships are feature-encoded to determine several encoded vectors corresponding to different adjacent position vectors; The individual encoded vectors are integrated to obtain the feature vector.
10. A training method for a multilingual knowledge fusion model, characterized in that, include: Obtain training corpus information and determine the target word vector in the training corpus information; The training corpus information of different languages is subjected to a first fusion and splicing process to obtain the first fused corpus; The first fused corpus is subjected to a nonlinear transformation to obtain the second fused corpus; The second fused corpus is subjected to a second fusion splicing process and a nonlinear transformation to obtain the predicted word vectors after multilingual knowledge fusion; The parameters of the multilingual knowledge fusion model are adjusted based on the loss value between the predicted word vector and the target word vector.
11. A text processing model training method, characterized in that, include: Obtain the training text and determine the target result in the training text; Generate the first word vectors for all words in the training text; Feature relationships are extracted from the first word vectors to obtain the text feature information of the training text; The text feature information is classified and identified to obtain the prediction result corresponding to the training text; The parameters of the text processing model are adjusted based on the loss value between the prediction result and the target result; wherein at least one of the training text, the first word vector, and the text feature information is obtained after being fused by the multilingual knowledge fusion method as described in any one of claims 1-4.
12. A multilingual knowledge fusion device, characterized in that, include: The first module is used to obtain information from the first corpus. The second module is used to acquire reference text corpus that differs in language from the first corpus information, and to determine cross-language corpus based on the reference text corpus. The cross-language corpus is obtained by segmenting and vectorizing the translated reference text corpus, and is cross-language corpus with the same text granularity as the first corpus information. The language of the translated reference text corpus is consistent with the language of the first corpus information. The third module is used to perform a first fusion and splicing process on the first corpus information and the cross-language corpus to obtain the first fused corpus. The fourth module is used to perform a nonlinear transformation on the first fused corpus to obtain the second fused corpus; The fifth module is used to perform a nonlinear transformation on the second fused corpus after performing a second fusion splicing process to obtain the second corpus information after multilingual knowledge fusion; Wherein, the first fusion splicing process is used to represent the processing of the first corpus information and the cross-language corpus after structural splicing and passing through the first fully connected layer, and the second fusion splicing process is used to represent the processing of the second fused corpus through the second fully connected layer.
13. A text processing device, characterized in that, include: The sixth module is used to obtain the target text and generate the first word vector of all words in the target text; The seventh module is used to extract feature relationships from the first word vector to obtain the text feature information of the target text; The eighth module is used to classify and identify the text feature information to obtain the target result corresponding to the target text; Wherein, at least one of the target text, the first word vector, and the text feature information is obtained after being fused by the multilingual knowledge fusion method according to any one of claims 1-4.
14. An electronic device, characterized in that, Including the processor and memory; The memory is used to store programs; The processor executes the program to implement the multilingual knowledge fusion method as described in any one of claims 1-5, or the text processing method as described in any one of claims 6-9, or the multilingual knowledge fusion model training method as described in claim 10, or the text processing model training method as described in claim 11.
15. A computer-readable storage medium, characterized in that, The storage medium stores a program that is executed by a processor to implement the multilingual knowledge fusion method as described in any one of claims 1-5, or the text processing method as described in any one of claims 6-9, or the multilingual knowledge fusion model training method as described in claim 10, or the text processing model training method as described in claim 11.
Citation Information
Patent Citations
A multilingual text classification method fusing theme information and BiLSTM-CNN
CN109885686A
Data processing method and device, computer equipment and storage medium
CN114328809A