Text dialogue method and device suitable for multiple languages, terminal equipment and storage medium
By identifying the language in which the user enters the text and combining the context information of the dialogue content, we generate reply text that meets the user's expectations, solving the problem of replies that do not conform to the language and grammar errors in traditional technology, realizing multilingual text dialogue and a more natural and accurate dialogue experience.
Patent Information
- Application Number
- CN202510176854.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional text dialogue technology cannot recognize the language in which the text is entered, resulting in the reply content that may not meet the language that users expect, and the reply text may have grammatical errors or incoherence, resulting in low accuracy and low naturalness during the conversation, making multilingual text dialogue impossible.
By obtaining the source text and dialogue content input by the user, the preset text dialogue model generates the local features of the phrase structure of the source text and the global features of the sentence structure, identify the target language corresponding to the source text, and generate dialogue features based on the context information of the dialogue content, and combine local features, global features, dialogue features and grammatical rules of the target language to generate reply text.
It improves the accuracy and naturalness of reply text during the conversation, realizes multilingual text dialogue, optimizes the user's dialogue experience, and ensures the language and grammatical accuracy of the reply.
Smart Images

Figure CN120104742A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text dialogue processing, and in particular to a text dialogue method, device, terminal equipment and storage medium applicable to multiple languages. Background Art
[0002] Text dialogue refers to the process of information exchange between humans, machines or users in the form of text. When users have conversations in a multilingual environment, the dialogue system needs to be able to understand and respond to texts in different languages input by users. To achieve this goal, the dialogue system usually needs to be able to process multiple languages and generate appropriate responses based on the context information of the conversation.
[0003] When processing text input by users, traditional text dialogue technology directly performs context analysis on the source text, tries to understand the user's intention and context information, and then generates a reply text based on the results of the context analysis. In traditional dialogue technology, it is impossible to identify the language of the input text, which may result in the reply content not being in the language expected by the user. Moreover, since the language of the current input text is not considered, it is impossible to reply according to the actual grammatical rules of the language, so that the reply text may contain grammatical errors, or appear stiff or incoherent, resulting in low accuracy and low naturalness of the reply during the dialogue process, and it is impossible to achieve multi-language text dialogue. Summary of the invention
[0004] The embodiments of the present invention provide a text conversation method, apparatus, terminal device and storage medium applicable to multiple languages, which not only improve the accuracy and naturalness of reply texts during a conversation, but also realize text conversations in multiple languages and optimize the user's conversation experience.
[0005] An embodiment of the present invention provides a text conversation method applicable to multiple languages, comprising:
[0006] Obtaining the source text input by the user in the current conversation and the conversation content of the current conversation; wherein the conversation content is used to represent the user input content that has been replied to and the output content of the replied user in the current conversation;
[0007] The source text and the dialogue content are input into a preset text dialogue model, so that the text dialogue model generates local features for representing the phrase structure of the source text and global features for representing the sentence structure of the source text according to the source text; generates the target language corresponding to the source text according to the local features and the global features; generates dialogue features according to the context information of the dialogue content; generates a reply text corresponding to the source text based on the local features, the global features, the dialogue features and the grammatical rules corresponding to the target language; wherein the preset text dialogue model integrates the grammatical rules of several different languages;
[0008] The reply text is used as the output content of the current conversation.
[0009] Preferably, the text dialogue model generates, based on the source text, local features for characterizing the phrase structure of the source text and global features for characterizing the sentence structure of the source text, including:
[0010] The text dialogue model converts the source text into a corresponding character vector sequence;
[0011] Performing a convolution operation on the character vector sequence to generate local features for representing the phrase structure of the source text;
[0012] A global feature for characterizing a sentence structure of a source text is generated according to the context dependency in the character vector sequence.
[0013] Preferably, the text dialogue model comprises: a multi-scale convolutional layer;
[0014] The performing a convolution operation on the character vector sequence to generate local features for representing the phrase structure of the source text includes:
[0015] By sliding different convolution kernels in a multi-scale convolution layer on the character vector sequence, phrase structures of different lengths are extracted; wherein the phrase structures include: subject-predicate phrases, verb-object phrases, attributive phrases, prepositional phrases, quantity phrases, and directional phrases;
[0016] Each phrase structure is downsampled through the pooling layer in the multi-scale convolutional layer to generate local features for characterizing the phrase structure of the source text.
[0017] Preferably, the text dialogue model comprises: a forward recurrent neural network and a backward recurrent neural network;
[0018] The step of generating a global feature for characterizing the sentence structure of the source text according to the context dependency in the character vector sequence includes:
[0019] Extracting forward context information of the character vector sequence according to the forward order of the character vector sequence through a forward recurrent neural network to generate a forward dependency relationship corresponding to each character vector; wherein the forward dependency relationship is used to represent the dependency relationship from the beginning of the sentence to the character vector;
[0020] Extracting reverse context information of the character vector sequence in reverse order of the character vector sequence through a backward recurrent neural network to generate a backward dependency relationship of each character vector; wherein the backward dependency relationship is used to represent a dependency relationship from the beginning of the character vector to the end of the sentence;
[0021] According to each forward dependency and each backward dependency, a global feature for characterizing the sentence structure of the source text is generated.
[0022] Preferably, generating the target language corresponding to the source text according to the local features and the global features includes:
[0023] Assign different weights to local features and global features based on the attention mechanism;
[0024] Generate a target feature vector including not only the local features of each character in the sentence of the source text but also the dependency relationship between the characters in the sentence of the source text according to the weights corresponding to the local features and the weights corresponding to the global features;
[0025] The target feature vector is matched with each language, and the target language corresponding to the source text is generated according to the matching result.
[0026] Preferably, the dialogue features include: a character feature vector and a dialogue context feature vector;
[0027] The generating of the dialogue feature according to the context information of the dialogue content includes:
[0028] Convert the replied user input content in the conversation content into a corresponding first word vector sequence; and convert the conversation content into a corresponding second word vector sequence;
[0029] Sliding different convolution kernels in a multi-scale convolution layer on the first word vector sequence to extract a character feature vector for characterizing the characteristics of the user's question;
[0030] Different convolution kernels in the multi-scale convolution layer slide on the second word vector sequence to extract a conversation context feature vector for representing the context information of the conversation content.
[0031] Preferably, the generation process of the preset text dialogue model includes:
[0032] Using source text samples input by the user in the conversation and conversation content samples as inputs of a generative adversarial network, using reply text samples corresponding to the source text samples as outputs of the generative adversarial network, and performing alternating iterative training on the generator and the discriminator in the generative adversarial network;
[0033] When the convergence of the generative adversarial network is detected, the generator after training is used as the preset text dialogue model.
[0034] Based on the above method embodiments, the present invention provides corresponding device embodiments.
[0035] An embodiment of the present invention provides a text dialogue device applicable to multiple languages, comprising: a dialogue data acquisition module, a reply text generation module and a text output module;
[0036] The dialogue data acquisition module is used to acquire the source text input by the user in the current dialogue and the dialogue content of the current dialogue; wherein the dialogue content is used to represent the user input content that has been replied to and the output content of the replied user in the current dialogue;
[0037] The reply text generation module is used to input the source text and the dialogue content into a preset text dialogue model, so that the text dialogue model generates local features for representing the phrase structure of the source text and global features for representing the sentence structure of the source text according to the source text; generates the target language corresponding to the source text according to the local features and the global features; generates dialogue features according to the context information of the dialogue content; generates the reply text corresponding to the source text based on the local features, the global features, the dialogue features and the grammatical rules corresponding to the target language; wherein the preset text dialogue model integrates the grammatical rules of several different languages;
[0038] The text output module is used to use the reply text as the output content of the current conversation.
[0039] Based on the above method embodiments, the present invention provides corresponding terminal device embodiments.
[0040] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for text conversation applicable to multiple languages described in the above-mentioned embodiment of the invention is implemented.
[0041] Based on the above method embodiments, the present invention provides corresponding storage medium item embodiments.
[0042] Another embodiment of the present invention provides a storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute a text conversation method applicable to multiple languages as described in the above-mentioned embodiment of the invention.
[0043] The following beneficial effects are achieved by implementing the present invention:
[0044] The embodiment of the present invention provides a text dialogue method, apparatus, terminal device and storage medium applicable to multiple languages. After obtaining the source text input by the user in the current dialogue, the present invention can first extract the local features used to characterize the phrase structure of the source text and the global features used to characterize the sentence structure of the source text, thereby accurately identifying the target language corresponding to the source text based on the phrase structure and sentence structure of the source text, and then selecting the corresponding grammatical rules based on the identified target language to generate the subsequent reply text. Furthermore, the present invention can also extract the corresponding dialogue features according to the context information of the dialogue content, so that the reply text corresponding to the source text can be generated based on the local features, global features, dialogue features and the grammatical rules corresponding to the target language. Compared with the prior art, the present invention can accurately determine the language used by the user based on the internal structure of the source text, thereby ensuring the language and grammatical correctness of the reply; and can also deeply understand the user's intention and context information based on the extracted local features, global features and dialogue features, thereby generating a reply text that conforms to the user's expected language, and based on the actual grammatical rules of the target language, the generated reply text is not only grammatically correct, but also naturally fluent in expression, avoiding stiffness or incoherence, thereby effectively solving the problems existing in traditional technologies, not only improving the accuracy and naturalness of the reply text during the dialogue process, but also by identifying the language of the current dialogue and making targeted replies, the present invention realizes multi-language text dialogue and optimizes the user's dialogue experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 The present invention is a flowchart of a text conversation method applicable to multiple languages provided by an embodiment of the present invention.
[0046] Figure 2 The diagram is a schematic diagram of a structure of a text dialogue device applicable to multiple languages provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] like Figure 1 FIG. 1 is a flow chart of a text conversation method applicable to multiple languages provided by an embodiment of the present invention. The text conversation method applicable to multiple languages includes:
[0049] Step S1: obtaining the source text input by the user in the current conversation and the conversation content of the current conversation; wherein the conversation content is used to represent the user input content that has been replied to and the output content of the replied user in the current conversation;
[0050] Step S2: inputting the source text and the dialogue content into a preset text dialogue model, so that the text dialogue model generates local features for characterizing the phrase structure of the source text and global features for characterizing the sentence structure of the source text according to the source text; generates the target language corresponding to the source text according to the local features and the global features; generates dialogue features according to the context information of the dialogue content; generates a reply text corresponding to the source text based on the local features, the global features, the dialogue features and the grammatical rules corresponding to the target language; wherein the preset text dialogue model integrates the grammatical rules of several different languages;
[0051] Step S3: Using the reply text as the output content of the current conversation.
[0052] For step S1, the present invention can automatically reply to the source text input by the user in the current dialogue during the text dialogue process, and comprehensively consider multiple levels of information input by the user, and use the text dialogue model to generate a reply, thereby solving the limitations of traditional multilingual text dialogue technology in processing user input and improving the accuracy and coherence of the dialogue.
[0053] In order to better obtain the language of the source text and generate the corresponding reply text, in the embodiment of the present invention, the source text input by the user in the current dialogue and the dialogue content of the current dialogue can be first obtained.
[0054] The replied user input content in the dialogue content refers to all the text content that the user has entered in the current dialogue session before entering the current source text. In principle, all the input text content may be one or more questions, statements or requests, which together constitute the interaction history between the user and the dialogue system.
[0055] The output content of the replied user in the dialogue content refers to all the replies or outputs given by the dialogue system or dialogue model to the content input by the user. These replies generally include direct answers to user questions, confirmation or refutation of statements, responses to requests, etc., reflecting the processing results of the dialogue system and the interaction status with the user.
[0056] The above contents are collectively referred to as dialogue content because they together constitute the context of the current dialogue, and this context provides key clues about user intentions, topic transitions, information supplements, etc. Inputting these dialogue contents and the source text entered by the user in the current dialogue into the preset text dialogue model can help the model understand the context and intention of the current source text more accurately, thereby generating more appropriate and accurate reply texts.
[0057] For step S2, the present invention is based on a preset text dialogue model. After receiving the source text and dialogue content as input, it can realize a series of functions such as language recognition, feature extraction, context understanding, grammar rule application and response generation, thereby improving the accuracy and coherence of the dialogue, and can support multi-language text dialogues.
[0058] In a preferred embodiment, the generation process of the preset text dialogue model includes:
[0059] Using source text samples input by the user in the conversation and conversation content samples as inputs of a generative adversarial network, using reply text samples corresponding to the source text samples as outputs of the generative adversarial network, and performing alternating iterative training on the generator and the discriminator in the generative adversarial network;
[0060] When the convergence of the generative adversarial network is detected, the generator after training is used as the preset text dialogue model.
[0061] It can be understood that the task of the generator is to generate pseudo reply text that is as close to the real reply text as possible, while the task of the discriminator is to distinguish whether the input reply text is real (from training data) or fake (generated by the generator).
[0062] The process of training the generator is as follows: in each iteration, the parameters of the discriminator are first fixed, and a batch of source text samples and dialogue content samples are sampled from the training data as the input of the generator. The generator generates a batch of pseudo reply texts based on these inputs and its own current parameters, and inputs this batch of pseudo reply texts and real reply text samples into the discriminator to obtain the real / fake classification results of these texts by the discriminator.
[0063] According to the feedback from the discriminator (usually the degree of classification error), the parameters of the generator are adjusted through the back-propagation algorithm so that the pseudo reply text it generates is closer to the real reply text.
[0064] The process of training the discriminator is as follows: in each iteration, the parameters of the generator are first fixed, a batch of source text samples and conversation content samples are sampled from the training data again, and the corresponding pseudo reply text is generated (this time using the current version of the generator). This batch of real reply text samples and pseudo reply texts are input into the discriminator together to enable the discriminator to distinguish whether these texts are real or fake. According to the error between the classification result of the discriminator and the real label, the parameters of the discriminator can be adjusted through the back propagation algorithm to improve its classification accuracy.
[0065] It is understandable that the generator can be trained first and then the discriminator, and then they can be trained alternately. In each iteration, the generator and the discriminator will adjust according to each other's current state, thus forming a dynamic game process. During the training process, the performance of the model can be evaluated by observing indicators such as the quality of the pseudo-reply text generated by the generator and the classification accuracy of the discriminator. When these indicators reach the preset threshold or no longer increase significantly, it can be considered that the generative adversarial network has converged. At this time, the generator after training is used as the preset text dialogue model for subsequent dialogue tasks.
[0066] In a preferred embodiment, when generating pseudo-reply text, the generator can deeply analyze the characteristics of the source text and the pseudo-reply text in terms of grammatical structure, word formation, lexical semantic distribution, etc., and construct a language feature map; and based on the self-supervised word alignment method of the attention mechanism, a high-precision word alignment matrix is obtained by constructing a multi-layer attention network to form an information-rich and robust multilingual representation, so that in the subsequent practical application process, the reply text corresponding to the source text can be accurately generated based on the learned word alignment matrix.
[0067] In a preferred embodiment, the generator can also use its internal convolutional neural network and bidirectional recurrent neural network structure, combined with the local features and global features of different languages in the language feature map, to extract and analyze the input text, and accurately determine the language to which the input text belongs by comparing it with the predefined language and referring to the mapping relationship between vocabulary and semantics in the word alignment matrix. For example, if the user inputs an English message, the generator will generate an English reply based on the English language features and conversation context, and ensure that the reply content is semantically consistent with the user's question and conforms to the English language habits.
[0068] Illustratively, in a preferred embodiment, the text dialogue model generates, based on the source text, local features for characterizing the phrase structure of the source text and global features for characterizing the sentence structure of the source text, including:
[0069] The text dialogue model converts the source text into a corresponding character vector sequence;
[0070] Performing a convolution operation on the character vector sequence to generate local features for representing the phrase structure of the source text;
[0071] A global feature for characterizing a sentence structure of a source text is generated according to the context dependency in the character vector sequence.
[0072] Specifically, when obtaining local features, local features can be extracted based on the multi-scale convolutional layer in the text dialogue model, and then:
[0073] By sliding different convolution kernels in a multi-scale convolution layer on the character vector sequence, phrase structures of different lengths are extracted; wherein the phrase structures include: subject-predicate phrases, verb-object phrases, attributive phrases, prepositional phrases, quantity phrases, and directional phrases;
[0074] Each phrase structure is downsampled through the pooling layer in the multi-scale convolutional layer to generate local features for characterizing the phrase structure of the source text.
[0075] Furthermore, when extracting global features, the text can be processed based on the forward recurrent neural network and the backward recurrent neural network in the text dialogue model to output accurate global features, and then:
[0076] The step of generating a global feature for characterizing the sentence structure of the source text according to the context dependency in the character vector sequence includes:
[0077] Extracting forward context information of the character vector sequence according to the forward order of the character vector sequence through a forward recurrent neural network to generate a forward dependency relationship corresponding to each character vector; wherein the forward dependency relationship is used to represent the dependency relationship from the beginning of the sentence to the character vector;
[0078] Extracting reverse context information of the character vector sequence in reverse order of the character vector sequence through a backward recurrent neural network to generate a backward dependency relationship of each character vector; wherein the backward dependency relationship is used to represent a dependency relationship from the beginning of the character vector to the end of the sentence;
[0079] According to each forward dependency and each backward dependency, a global feature for characterizing the sentence structure of the source text is generated.
[0080] It can be understood that the present invention can more accurately understand the intention and needs of the user input by analyzing the local and global features of the sentence and the context information of the conversation, thereby generating a more accurate response.
[0081] Schematically, in the process of local feature extraction, multi-scale convolution can capture phrase structures of different lengths and improve the diversity of feature extraction. Moreover, based on the pooling layer, the dimension of the feature vector can be reduced by downsampling operation to improve the computational efficiency. In a preferred embodiment, the text dialogue model of the present invention can use an internal algorithm (such as a convolutional neural network CNN, etc.) to extract local features for characterizing phrase structures from the source text, such as character combinations, vocabulary patterns, common phrase structures, etc. in a specific language.
[0082] In the process of global feature extraction, it can be completed through a bidirectional recurrent neural network (Bidirectional Recurrent Neural Network), which includes a forward recurrent neural network and a backward recurrent neural network, which can process forward and backward input information at the same time and can capture the contextual features of sequence data more comprehensively.
[0083] It can be understood that in a bidirectional recurrent neural network, there are two independent RNN layers: a forward RNN layer and a backward RNN layer. The forward RNN layer processes data in the forward order of the input sequence (e.g., from left to right) to capture the forward dependencies; while the backward RNN layer processes data in the reverse order of the input sequence (e.g., from right to left) to capture the backward dependencies.
[0084] Specifically, the forward RNN processes data in the forward order of the character vector sequence (for example, from left to right), which can gradually read the vector representation of each character (or word) and update its internal state based on the previous information (i.e., the forward context). In this way, the forward RNN can capture the forward dependency from the beginning of the sentence to the current character (or word). The forward dependency reflects the influence of the previous part of the sentence on the current part, such as the influence of the subject on the predicate, or the influence of the guide word of the clause on the content of the clause.
[0085] The backward RNN processes data in the reverse order of the character vector sequence (for example, from right to left). It also reads the vector representation of each character (or word) step by step, but this time updates its internal state based on later information (i.e., backward context). Therefore, the backward RNN is able to capture the backward dependency from the current character (or word) to the end of the sentence. The backward dependency reflects the influence of the later part of the sentence on the current part, such as the influence of the object on the verb, or the influence of the punctuation at the end of the sentence on the tone of the entire sentence.
[0086] Therefore, the forward dependencies and backward dependencies together constitute the complete contextual dependencies of the sentence. They not only reflect the grammatical relationship between the various parts of the sentence (such as subject-verb-object structure, clause structure, etc.), but also reflect the semantic connection (such as reference resolution, synonym replacement, etc.). Therefore, concatenating these dependencies can generate a global feature vector to more accurately characterize the sentence structure of the source text.
[0087] In a preferred embodiment, generating the target language corresponding to the source text according to the local features and the global features includes:
[0088] Assign different weights to local features and global features based on the attention mechanism;
[0089] Generate a target feature vector including not only the local features of each character in the sentence of the source text but also the dependency relationship between the characters in the sentence of the source text according to the weights corresponding to the local features and the weights corresponding to the global features;
[0090] The target feature vector is matched with each language, and the target language corresponding to the source text is generated according to the matching result.
[0091] It can be understood that in the embodiment of the present invention, the attention mechanism is used to assign weights to local features (mainly reflecting phrase structure) and global features (mainly reflecting sentence structure), so that the model can learn how to dynamically adjust the attention to each feature according to the current context information.
[0092] Due to the attention mechanism, each feature in the target feature vector is weighted according to its importance (i.e., weight), thereby ensuring that important features receive more attention in subsequent processing. The similarity or distance between the target feature vector and each language feature can be further calculated, and a language that best matches the source text feature can be selected as the target language.
[0093] Therefore, the embodiment of the present invention combines the attention mechanism, feature fusion and matching classification, so that the model can more accurately understand the semantic and structural information of the input text, thereby accurately identifying the target language.
[0094] Furthermore, in addition to obtaining the target language corresponding to the source text, the present invention can also obtain a character feature vector and a conversation context feature vector based on the context information of the conversation content, so as to deeply understand the user's intention and context information, thereby generating a more accurate reply and providing the user with a more accurate and personalized reply experience.
[0095] The generating of the dialogue feature according to the context information of the dialogue content includes:
[0096] Convert the replied user input content in the conversation content into a corresponding first word vector sequence; and convert the conversation content into a corresponding second word vector sequence;
[0097] Sliding different convolution kernels in a multi-scale convolution layer on the first word vector sequence to extract a character feature vector for characterizing the characteristics of the user's question;
[0098] Different convolution kernels in the multi-scale convolution layer slide on the second word vector sequence to extract a conversation context feature vector for representing the context information of the conversation content.
[0099] It can be understood that the character feature vector refers to a high-level summary of the characteristics of the user's questions, which mainly reflects the unique style, interests or concerns of the user in the conversation. The present invention can use different convolution kernels in the multi-scale convolution layer to slide on the first word vector sequence, thereby capturing phrase structures of different lengths and specific patterns of user questions, and then extract and integrate these features through convolution and pooling operations, and finally generate a character feature vector that can represent the user's characteristics.
[0100] The conversation context feature vector is a comprehensive reflection of the context information in the entire conversation process, which covers the historical information in the conversation, the background of the current problem, and the possible subsequent development direction. The embodiment of the present invention also uses different convolution kernels in the multi-scale convolution layer to slide on the second word vector sequence to extract the key information and contextual relationship in the conversation, and generates a feature vector that can fully reflect the conversation context by integrating these features.
[0101] Therefore, the character feature vector can enable the model to understand the user's personality and needs more deeply, and thus grasp the user's true intention more accurately. The conversation context feature vector provides comprehensive background information of the conversation, allowing the model to more accurately understand the context and background of the current conversation. By comprehensively considering the character feature vector and the conversation context feature vector, the model can more comprehensively understand the user's needs and conversation background, thereby generating responses that are more in line with the user's expectations and personality.
[0102] After generating the character feature vector and the conversation context feature vector, these two vectors will be input into the subsequent neural network layer for further processing and fusion.
[0103] Illustratively, the text dialogue model of the present invention is constructed on the basis of understanding and generating language. To achieve this goal, the text dialogue model can learn and integrate grammatical rules of various languages. These grammatical rules are the basic framework of the language, which define how words are combined into phrases, sentences, and the logical relationship between sentences. That is, the model can further generate a reply text corresponding to the grammatical rules based on the local features, global features, character feature vectors, dialogue context feature vectors, and the grammatical rules corresponding to the target language. By identifying the language of the current dialogue and making targeted replies, the present invention realizes multi-language text dialogue and optimizes the user's dialogue experience.
[0104] For step S3, in a preferred embodiment, after a series of processing and calculations by the text dialogue model of the present invention, text content for responding to the user input or request is generated according to the reply text, and then output to the user in the current dialogue.
[0105] like Figure 2 As shown, based on the above-mentioned various embodiments of the text conversation method applicable to multiple languages, the present invention provides a corresponding device embodiment;
[0106] An embodiment of the present invention provides a text dialogue device applicable to multiple languages, comprising: a dialogue data acquisition module, a reply text generation module and a text output module;
[0107] The dialogue data acquisition module is used to acquire the source text input by the user in the current dialogue and the dialogue content of the current dialogue; wherein the dialogue content is used to represent the user input content that has been replied to and the output content of the replied user in the current dialogue;
[0108] The reply text generation module is used to input the source text and the dialogue content into a preset text dialogue model, so that the text dialogue model generates local features for representing the phrase structure of the source text and global features for representing the sentence structure of the source text according to the source text; generates the target language corresponding to the source text according to the local features and the global features; generates dialogue features according to the context information of the dialogue content; generates the reply text corresponding to the source text based on the local features, the global features, the dialogue features and the grammatical rules corresponding to the target language; wherein the preset text dialogue model integrates the grammatical rules of several different languages;
[0109] The text output module is used to use the reply text as the output content of the current conversation.
[0110] It should be noted that the device embodiments described above are merely schematic, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, and may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without paying creative labor.
[0111] Those skilled in the art can clearly understand that, for the sake of convenience and simplicity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0112] Based on the above-mentioned various embodiments of the text conversation method applicable to multiple languages, the present invention provides corresponding embodiments of terminal equipment items.
[0113] An embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a text conversation method applicable to multiple languages as described in any method embodiment of the present invention is implemented.
[0114] The terminal device may be a computing terminal device such as a desktop computer, a notebook, a palm computer, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0115] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, and uses various interfaces and lines to connect various parts of the entire terminal device.
[0116] The memory can be used to store the computer program, and the processor realizes various functions of the terminal device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Med ia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device or other volatile solid-state storage device.
[0117] Based on the above-mentioned various embodiments of the text conversation method applicable to multiple languages, the present invention provides a corresponding storage medium item embodiment.
[0118] An embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a text conversation method applicable to multiple languages as described in any method embodiment of the present invention.
[0119] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0120] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A text conversation method applicable to multiple languages, characterized in that: include: Acquire the source text input by the user in the current conversation and the conversation content of the current conversation; wherein the conversation content is used to represent the user input content that has been replied to and the output content of the replied user in the current conversation; The source text and the dialogue content are input into a preset text dialogue model, so that the text dialogue model generates local features for representing the phrase structure of the source text and global features for representing the sentence structure of the source text according to the source text; generates the target language corresponding to the source text according to the local features and the global features; generates dialogue features according to the context information of the dialogue content; generates a reply text corresponding to the source text based on the local features, the global features, the dialogue features and the grammatical rules corresponding to the target language; wherein the preset text dialogue model integrates the grammatical rules of several different languages; The reply text is used as the output content of the current conversation.
2. A text conversation method applicable to multiple languages as claimed in claim 1, characterized in that: The text dialogue model generates, according to the source text, local features for representing the phrase structure of the source text and global features for representing the sentence structure of the source text, including: The text dialogue model converts the source text into a corresponding character vector sequence; Performing a convolution operation on the character vector sequence to generate local features for representing the phrase structure of the source text; A global feature for characterizing a sentence structure of a source text is generated according to the context dependency in the character vector sequence.
3. A text conversation method applicable to multiple languages as claimed in claim 2, characterized in that: The text dialogue model includes: a multi-scale convolutional layer; The performing a convolution operation on the character vector sequence to generate local features for representing the phrase structure of the source text includes: By sliding different convolution kernels in a multi-scale convolution layer on the character vector sequence, phrase structures of different lengths are extracted; wherein the phrase structures include: subject-predicate phrases, verb-object phrases, attributive phrases, prepositional phrases, quantity phrases, and directional phrases; Each phrase structure is downsampled through the pooling layer in the multi-scale convolutional layer to generate local features for characterizing the phrase structure of the source text.
4. A text conversation method applicable to multiple languages as claimed in claim 3, characterized in that: The text dialogue model includes: a forward recurrent neural network and a backward recurrent neural network; The step of generating a global feature for characterizing the sentence structure of the source text according to the context dependency in the character vector sequence includes: Extracting forward context information of the character vector sequence according to the forward order of the character vector sequence through a forward recurrent neural network to generate a forward dependency relationship corresponding to each character vector; wherein the forward dependency relationship is used to represent the dependency relationship from the beginning of the sentence to the character vector; Extracting reverse context information of the character vector sequence in reverse order of the character vector sequence through a backward recurrent neural network to generate a backward dependency relationship of each character vector; wherein the backward dependency relationship is used to represent a dependency relationship from the beginning of the character vector to the end of the sentence; According to each forward dependency and each backward dependency, a global feature for characterizing the sentence structure of the source text is generated.
5. A text conversation method applicable to multiple languages as claimed in claim 4, characterized in that: Generating the target language corresponding to the source text according to the local features and the global features includes: Assign different weights to local features and global features based on the attention mechanism; Generate a target feature vector including not only the local features of each character in the sentence of the source text but also the dependency relationship between the characters in the sentence of the source text according to the weights corresponding to the local features and the weights corresponding to the global features; The target feature vector is matched with each language, and the target language corresponding to the source text is generated according to the matching result.
6. A text conversation method applicable to multiple languages as claimed in claim 5, characterized in that: The dialogue features include: a character feature vector and a dialogue context feature vector; The generating of the dialogue feature according to the context information of the dialogue content includes: Convert the replied user input content in the conversation content into a corresponding first word vector sequence; and convert the conversation content into a corresponding second word vector sequence; Sliding different convolution kernels in a multi-scale convolution layer on the first word vector sequence to extract a character feature vector for characterizing the characteristics of the user's question; Different convolution kernels in the multi-scale convolution layer slide on the second word vector sequence to extract a conversation context feature vector for representing the context information of the conversation content.
7. The text conversation method applicable to multiple languages as claimed in claim 1, characterized in that: The generation process of the preset text dialogue model includes: Using source text samples input by the user in the conversation and conversation content samples as inputs of a generative adversarial network, using reply text samples corresponding to the source text samples as outputs of the generative adversarial network, and performing alternating iterative training on the generator and the discriminator in the generative adversarial network; When the convergence of the generative adversarial network is detected, the generator after training is used as the preset text dialogue model.
8. A text conversation device suitable for multiple languages, characterized in that: include: Dialogue data acquisition module, reply text generation module and text output module; The dialogue data acquisition module is used to acquire the source text input by the user in the current dialogue and the dialogue content of the current dialogue; wherein the dialogue content is used to represent the user input content that has been replied to and the output content of the replied user in the current dialogue; The reply text generation module is used to input the source text and the dialogue content into a preset text dialogue model, so that the text dialogue model generates local features for representing the phrase structure of the source text and global features for representing the sentence structure of the source text according to the source text; generates the target language corresponding to the source text according to the local features and the global features; generates dialogue features according to the context information of the dialogue content; generates the reply text corresponding to the source text based on the local features, the global features, the dialogue features and the grammatical rules corresponding to the target language; wherein the preset text dialogue model integrates the grammatical rules of several different languages; The text output module is used to use the reply text as the output content of the current conversation.
9. A terminal device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a text conversation method applicable to multiple languages as claimed in any one of claims 1 to 7 is implemented.
10. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute a text conversation method applicable to multiple languages as claimed in any one of claims 1 to 7.
Citation Information
Cited By
Multilingual intelligent interactive processing system
CN122366471A