Interaction Method, Device, Electronic Device, and Storage Medium

By determining the lattice position of the lattice entity in the Indo-European voice interaction system and reverting it to the original entity, the poor interaction accuracy caused by the lattice of the noun is solved, and the user experience is improved.

CN114925679BActive Publication Date: 2025-07-22IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210459426.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-27
Publication Date
2025-07-22
Estimated Expiration
2042-04-27

AI Technical Summary

Technical Problem

In the prior art, the Indo-European voice interaction system has poor interaction accuracy in the problem of noun change, resulting in poor user experience.

Method used

By performing entity extraction in the Indo-European voice interaction system, the lattice of the lattice entity in the command text is determined, and the lattice entity is restored to the original entity based on the lattice, and then responding to user commands.

Benefits of technology

It improves the accuracy and effectiveness of Indo-European language interactions and improves the user's interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114925679B_ABST
    Figure CN114925679B_ABST
Patent Text Reader

Abstract

The present invention provides an interaction method, device, electronic device and storage medium. The method includes: obtaining a command text of a user; when the language of the command text belongs to the Indo-European language family, performing entity extraction on the command text to obtain the declined entities in the command text, and determining the case positions to which the declined entities belong in the command text; based on the case positions to which the declined entities belong in the command text, determining the original entities corresponding to the declined entities; and responding to the command text based on the original entities. The method, device, electronic device and storage medium provided by the present invention, by extracting the declined entities in the command text and determining the case positions to which the declined entities belong in the command text when the language of the command text belongs to the Indo-European language family, and restoring the original entities corresponding to the declined entities based on this, thus realizing the solution to the noun declension problem existing in the Indo-European language interaction, improving the accuracy and effectiveness of the Indo-European language interaction, and greatly enhancing the user interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to an interaction method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of speech recognition technology and natural language understanding technology, speech interaction products have developed rapidly in all walks of life and are widely used in various fields such as industry, household appliances, communications, automotive electronics, medical care, home services, and electronic products. At the same time, single-language speech interaction alone cannot meet the market's demand for speech interaction, and multi-language speech interaction has more and more application scenarios. However, for speech interaction in the Indo-European language family, there is an urgent problem to be solved: noun declension is a common feature of Indo-European languages, which is different from Asian languages.

[0003] In the current human-computer interaction dialogue system, the common semantic understanding method based on the sequence labeling model can only obtain the entities in the actual sentence. However, for the Indo-European language family, the entities in the actual sentence are usually the declined entities, which results in the entities extracted by the existing technology not conforming to the entities actually required by the user, leading to the inability to proceed with the subsequent interaction process and affecting the user's interaction experience. Summary of the Invention

[0004] The present invention provides an interaction method, device, electronic device and storage medium to solve the problem of poor user experience in the speech interaction of the Indo-European language family in the existing technology.

[0005] The present invention provides an interaction method, including:

[0006] Obtain a command text of a user;

[0007] When the language of the command text belongs to the Indo-European language family, perform entity extraction on the command text to obtain the declined entity in the command text, and determine the case position to which the declined entity belongs in the command text;

[0008] Based on the case position to which the declined entity belongs in the command text, determine the original entity corresponding to the declined entity;

[0009] Respond to the command text based on the original entity.

[0010] According to an interaction method provided by the present invention, the determining the case position to which the declined entity belongs in the command text includes:

[0011] Based on the token representation of the pre-tokenization, select the case position of the inflected entity in the command text from each case position under the language of the command text, where the pre-tokenization is the token arranged before the inflected entity in the command text.

[0012] According to an interaction method provided by the present invention, determining the original entity corresponding to the inflected entity based on the case position of the inflected entity in the command text includes:

[0013] Select the part of speech to which the inflected entity belongs from each part of speech under the language of the command text;

[0014] Based on the part of speech to which the inflected entity belongs and the case position of the inflected entity in the command text, determine the original entity corresponding to the inflected entity.

[0015] According to an interaction method provided by the present invention, determining the original entity corresponding to the inflected entity based on the part of speech to which the inflected entity belongs and the case position of the inflected entity in the command text includes:

[0016] Based on the relationship between various pre-set combinations of parts of speech and case positions and declension rules, determine the declension rule corresponding to the part of speech and case position to which the inflected entity belongs;

[0017] Based on the declension rule corresponding to the part of speech and case position to which the inflected entity belongs, perform declension reduction on the inflected entity to obtain the original entity corresponding to the inflected entity.

[0018] According to an interaction method provided by the present invention, extracting the inflected entity from the command text includes:

[0019] Based on a general dictionary, encode the command text to obtain the text vector of the command text, where the general dictionary is constructed based on multilingual parallel corpora;

[0020] Based on the text vector of the command text, perform entity extraction on the command text to obtain the inflected entity in the command text.

[0021] According to an interaction method provided by the present invention, performing entity extraction on the command text based on the text vector of the command text to obtain the inflected entity in the command text includes:

[0022] Based on an entity extraction model, perform entity extraction on the text vector to obtain the inflected entity;

[0023] The entity extraction model is trained based on the sample text vectors and sample entities of the first sample text in the language of the command text on the basis of a general model, and the general model is trained based on the sample text vectors and sample entities of the second sample text in multiple languages.

[0024] According to an interaction method provided by the present invention, the responding to the command text based on the original entity includes:

[0025] Matching the original entity with a plurality of pre-stored candidate entities to obtain a matching result, and responding to the command text based on the matching result.

[0026] The present invention also provides an interaction device, including:

[0027] An acquisition unit for acquiring a command text of a user;

[0028] An extraction unit for performing entity extraction on the command text to obtain an inflected entity in the command text and determining a case position to which the inflected entity belongs in the command text when the language of the command text belongs to the Indo-European language family;

[0029] A restoration unit for determining an original entity corresponding to the inflected entity based on the case position to which the inflected entity belongs in the command text;

[0030] A response unit for responding to the command text based on the original entity.

[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements the interaction method as described in any one of the above when executing the program.

[0032] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and the computer program implements the interaction method as described in any one of the above when being executed by a processor.

[0033] The interaction method, device, electronic device, and storage medium provided by the present invention perform entity extraction on a command text to obtain an inflected entity in the command text and determine a case position to which the inflected entity belongs in the command text when the language of the command text belongs to the Indo-European language family, restore the original entity corresponding to the inflected entity based on this, and then respond to the command text based on the original entity, thereby solving the problem of noun inflection in Indo-European language interaction, improving the accuracy and effectiveness of Indo-European language interaction, and greatly enhancing the user's interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0035] Figure 1 is one of the schematic flowcharts of the interaction method provided by the present invention;

[0036] Figure 2 is one of the schematic flowcharts of the method for determining the original entity provided by the present invention;

[0037] Figure 3 is the second schematic flowchart of the method for determining the original entity provided by the present invention;

[0038] Figure 4 is an example diagram of the relationship between various combinations of word case forms and declension rules provided by the present invention;

[0039] Figure 5 is the schematic flowchart of the entity extraction method provided by the present invention;

[0040] Figure 6 is the second schematic flowchart of the interaction method provided by the present invention;

[0041] Figure 7 is the schematic structural diagram of the interaction device provided by the present invention;

[0042] Figure 8 is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0043] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0044] In practical applications, when analyzing users' usage behaviors, it is found that in vertical scenarios such as navigating to a specific POI (Point of Interest), calling a contact, or searching for a song, nouns in Indo-European language sentences often have the problem of declension. The existing slot extraction methods in NLP (Natural Language Processing) can only obtain the form of the noun after declension in the sentence. However, the declension of nouns involves gender, singular / plural, preceding prepositions or verbs. Using a way of listing rules is not only complex to process but also basically difficult to implement. Once the extracted entity does not match the original entity, it will lead to the failure of the entire interaction, greatly affecting the user's interaction experience.

[0045] For example, in the Chinese interaction scenario, assume that there is a contact named "Anton" in the address book. When the user issues a voice request "Call Anton", the system obtains the following semantic results through semantic understanding: (1) Domain: Phone; (2) Intent: Make a call; (3) Slot: Contact name (name)=Anton. Searching the address book through the semantic result "name=Anton", it can be found at this time.

[0046] In the Russian interaction scenario, assume that there is a contact named "Антон" in the Russian address book. When the user issues a voice request "Позвоните Антону", in this case, "Антону" should use the third case (i.e., the dative case) form, while the contacts stored in the address book generally use the first case (i.e., the nominative case, "Антон") form. Therefore, searching the address book through the semantic result "name=Антону" obtained by the system's semantic understanding, it cannot be found at this time, and the interaction process cannot continue.

[0047] In view of this, the present invention provides an interaction method. Figure 1 is one of the flow schematic diagrams of the interaction method provided by the present invention, as Figure 1 shown, the method includes:

[0048] Step 110, obtain the command text of the user.

[0049] Here, the command text is the text used to represent the user's interaction command, which can specifically be the text directly input by the user or the text obtained by voice transcription of the user's input voice data. The embodiments of the present invention do not make specific limitations on this.

[0050] Step 120, in the case that the language of the command text belongs to the Indo-European language family, perform entity extraction on the command text to obtain the declension entity in the command text, and determine the case position to which the declension entity belongs in the command text;

[0051] Step 130, determining the original entity corresponding to the declension entity based on the case position of the declension entity in the command text;

[0052] Step 140, respond to the command text based on the original entity.

[0053] Specifically, the Indo-European language family may include German, Russian, Spanish, Portuguese, and Persian. Case is a grammatical category that represents the structural and semantic relationship between words. The case to which the case entity belongs in the command text is used to represent the structural and semantic relationship between the noun corresponding to the case entity and other words in the command text. For example, in Russian, the case can be one of the nominative case, the genitive case, the accusative case, the instrumental case, and the prepositional case.

[0054] Considering that in existing natural language understanding solutions, the extracted entity information can only be consistent with the entities in the actual sentence, when extracting entities for Indo-European languages, it often happens that the extracted entities are inflected entities that do not correspond to the original form of the entities, thereby reducing the accuracy of the final interaction.

[0055] To address this problem, in an embodiment of the present invention, when the language to which the command text belongs belongs to the Indo-European language family, entity extraction is performed on the command text to obtain an entity after case inflection in the command text, i.e., a case-inflected entity. Then, the case position to which the case-inflected entity belongs in the command text is determined, and based on the case position to which the case-inflected entity belongs in the command text, the entity in the original form corresponding to the case-inflected entity, i.e., the original entity, is restored. On this basis, the command text can be responded to according to the original entity to complete the interaction process, thereby improving the accuracy and effectiveness of the interaction.

[0056] Taking the Russian interaction scenario as an example, the case of a noun is reflected in the change of the ending of the noun. For example, if the command text is "ПозвонитеAнтону" (Chinese meaning: call Anton), the case entity is "Aнтону", and the case position of the case entity in the command text is the giving case. Since the contacts stored in the address book generally use the nominative case, that is, the original form of the noun, "Aнтону" can be restored to the original entity "Aнтон". On this basis, "Aнтон" can be matched with each contact stored in the address book to obtain the phone number of "Aнтон" and dial it, thereby completing the response process of the command text.

[0057] When performing step 130, the corresponding original entity can be determined only according to the case position of the declined entity in the command text, or the corresponding original entity can be determined by combining the case position of the declined entity, as well as other information related to declension such as part-of-speech information and word type. Specifically, it can be correspondingly set according to the application scenario. For example, in the interaction scenario of calling a contact, the type of the declined entity is relatively fixed, all being personal names, and personal names in Russian all belong to the neuter gender. At this time, the corresponding original entity can be determined only according to the case position of the declined entity in the command text.

[0058] The method provided by the embodiments of the present invention performs entity extraction on the command text when the language of the command text belongs to the Indo-European language family, obtains the declined entity in the command text, determines the case position of the declined entity in the command text, restores the original entity corresponding to the declined entity based on this, and then responds to the command text based on the original entity, thereby solving the problem of noun declension in Indo-European language interactions, improving the accuracy and effectiveness of Indo-European language interactions, and greatly enhancing the user's interaction experience.

[0059] Based on the above embodiments, in step 120, determining the case position of the declined entity in the command text includes:

[0060] Based on the token representation of the previous tokenization, select the case position of the declined entity in the command text from each case position in the language of the command text. The previous token is the token arranged before the declined entity in the command text.

[0061] Specifically, considering that for the Indo-European language family, the case position of a noun in a sentence depends on the word arranged before the noun in the sentence. For example, in Russian, when words such as от, до, из, из-за or около appear, the following noun generally uses the second case (i.e., the genitive case). Therefore, the embodiments of the present invention select the case position of the declined entity in the command text from each case position in the language of the command text according to the token representation of the previous token. Here, the previous token is the token arranged before the declined entity in the command text, thereby improving the accuracy of case position determination.

[0062] For example, if the command text is “Позвоните Антону”, and the declined entity is “Антону”, then the token arranged before the declined entity in the command text, that is, the previous token, is “Позвоните”. According to the token representation of this previous token, the case position of the declined entity in the command text can be selected from each case position in Russian as the dative case.

[0063] Here, the token representation of the pre-order token can be the feature representation of the token obtained by encoding. Specifically, it can be directly obtained by encoding the token, or it can be obtained by extracting from the text representation of the command text obtained by the encoder in the entity extraction model used to perform entity extraction. The embodiments of the present invention do not make specific limitations on this.

[0064] Further, select the case position of the declension entity in the command text, which can be specifically implemented by a classification model for performing case classification. The classification model here can be a model common to multiple languages in the Indo-European language family, or a model separately trained for the language of the command text. The structure of the classification model can, for example, adopt CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), etc. The embodiments of the present invention do not make specific limitations on this.

[0065] Based on any of the above embodiments, Figure 2 is one of the flow diagrams of the method for determining the original entity provided by the present invention, as Figure 2 shown, step 130 includes:

[0066] Step 131, select the part of speech to which the declension entity belongs from each part of speech in the language of the command text;

[0067] Step 132, based on the part of speech to which the declension entity belongs and the case position of the declension entity in the command text, determine the original entity corresponding to the declension entity.

[0068] Specifically, the part of speech is a grammatical category reflecting the attribute concept of a word, and the part of speech to which the declension entity belongs is used to reflect the attribute concept of the noun corresponding to the declension entity. For example, in Russian, the part of speech can be one of neuter, feminine, and masculine.

[0069] Considering that the declension of nouns in the Indo-European language family is also related to the part of speech of the nouns, therefore, the embodiments of the present invention select the part of speech to which the declension entity belongs from each part of speech in the language of the command text. On this basis, the original form of the entity corresponding to the declension entity, that is, the original entity, can be restored by combining the part of speech to which the declension entity belongs and the case position of the declension entity in the command text, thereby improving the accuracy of declension restoration.

[0070] Based on any of the above embodiments, Figure 3 is the second flow diagram of the method for determining the original entity provided by the present invention, as Figure 3 shown, step 132 includes:

[0071] Step 1321: Determine the declension rule corresponding to the part of speech and case of the declension entity based on the relationship between various pre-set combinations of part of speech and case and the declension rules.

[0072] Step 1322: Based on the declension rule corresponding to the part of speech and case of the declension entity, perform declension reduction on the declension entity to obtain the original entity corresponding to the declension entity.

[0073] Specifically, the declension rule is the rule for a noun to undergo declension. Taking Russian as an example, the declension rule can specifically be the rule for the ending of a noun to change. After determining the part of speech and case of the declension entity, the declension rule corresponding to the part of speech and case of the declension entity can be determined according to the relationship between various pre-set combinations of part of speech and case and the declension rules. On this basis, the declension entity can be declension-reduced according to the declension rule corresponding to the part of speech and case of the declension entity, so as to restore the entity in the original form, that is, the original entity, for subsequent interaction processes.

[0074] For example, taking a singular noun in Russian as an example, Figure 4 is an example diagram of the relationship between various combinations of part of speech and case and declension rules provided by the present invention. As Figure 4 shown, the parts of speech are divided into feminine, masculine, and neuter, and the cases are divided into six types: nominative, genitive, dative, accusative, instrumental, and prepositional, that is, corresponding to the first, second, third, fourth, fifth, and sixth cases in Figure 4 respectively. The command text is "Позвоните Антону", then the declension entity is "Антону", the case it belongs to in the command text is the dative case (i.e., the third case), and the part of speech it belongs to is neuter. Therefore, according to the relationship between various combinations of part of speech and case and the declension rules, the declension rule corresponding to the part of speech and case of the declension entity can be determined as "-o→-y". Based on this, the reduction rule of the declension entity should be determined as "-y→-o". And since "Антону" is a proper name and no "-o" needs to be added to the ending, the original entity obtained by declension reduction is "Антон".

[0075] Based on any of the above embodiments, Figure 5 is a schematic flowchart of the entity extraction method provided by the present invention. As Figure 5 shown, entity extraction is performed on the command text to obtain the declension entity in the command text, including:

[0076] Step 510: Encode the command text based on a general dictionary to obtain a text vector of the command text. The general dictionary is constructed based on multilingual parallel corpora;

[0077] Step 520: Based on the text vector of the command text, perform entity extraction on the command text to obtain the declension entity in the command text.

[0078] Specifically, multilingual parallel corpora are corpora that express the same meaning but are in different languages. Considering that existing natural language understanding solutions perform entity extraction on sentences in a single language and are not applicable to multiple languages, and with the acceleration of the globalization process, a solution that only supports a single language cannot meet the actual needs. In view of the above problems, in the embodiments of the present invention, a multilingual shared dictionary, i.e., a general dictionary, is constructed based on multilingual parallel corpora, and then, based on the general dictionary, a command text is encoded to obtain a word embedding vector that represents the meaning of the command text and removes language information, and this is used as the text vector of the command text. On this basis, entity extraction is performed on the command text according to the text vector of the command text to obtain the inflected entities in the command text, so as to ensure the universality of the entity extraction method for multiple languages and meet the user's requirement for an interactive experience adapted to multiple languages.

[0079] Furthermore, the general dictionary can specifically be obtained by constructing the embedding of multilingual parallel corpora through the Byte Pair Encoding (BPE) algorithm, and entity extraction can specifically be implemented through a neural network model. It can be understood that by constructing the general dictionary, the alignment effect of different languages in the embedding space can be significantly improved, so that the subsequent neural network model for performing entity extraction pays more attention to the organizational structure of the language rather than the language type.

[0080] Based on any of the above embodiments, step 520 includes:

[0081] Based on the entity extraction model, entity extraction is performed on the text vector to obtain inflected entities;

[0082] The entity extraction model is trained based on the sample text vectors and sample entities of the first sample text in the language of the command text on the basis of the general model, and the general model is trained based on the sample text vectors and sample entities of the second sample text in multiple languages.

[0083] Specifically, considering that for multilingual interaction scenarios, it is rather troublesome to train a separate neural network model for each language, which greatly increases the operation cost. In view of this, in the embodiments of the present invention, sentences in multiple languages of the Indo-European language family with the same meaning are first collected as the second sample text, and the sample entities in the second sample text are labeled. Then, the second sample text is encoded according to the general dictionary to obtain the sample text vectors of the second sample text in the same semantic space. Next, the sample text vectors and sample entities of the second sample text are used to train the initial model to obtain a general model applicable to multiple languages. On this basis, the general model is fine-tuned to obtain the entity extraction model for the language to which the command text belongs.

[0084] Here, the fine-tuning of the general model can be specifically achieved in the following way: First, collect the first sample text in the language of the command text, label the sample entities in the first sample text, and then encode the first sample text according to the general dictionary to obtain the sample text vector of the first sample text; Since the entities output by the general model do not contain language information, a mapping layer can be added on the basis of the general model to obtain an initial entity extraction model that can map back to the entities in this language; Finally, apply the sample text vector and sample entities of the first sample text to train the initial entity extraction model, so as to obtain the entity extraction model corresponding to this language after fine-tuning, and then the accuracy of the entity extraction task in this language can be improved.

[0085] After training the entity extraction model, input the text vector of the encoded command text into the entity extraction model, and the entity extraction model extracts entities from the text vector to extract the inflected entities in the command text.

[0086] Based on any of the above embodiments, step 140 includes:

[0087] Match the original entity with multiple pre-stored candidate entities to obtain a matching result, and based on the matching result, respond to the command text.

[0088] Specifically, the candidate entity is an entity that can be selected in the interaction scenario. For example, in the interaction scenario of a phone, the candidate entity can be each contact stored in the address book. For another example, in the interaction scenario of music, the candidate entity can be each song stored in the playlist. For another example, in the interaction scenario of navigation, the candidate entity can be each address stored in the address library.

[0089] After obtaining the entity in the original form, that is, the original entity, the original entity can be matched with multiple candidate entities in the original form pre-stored, so as to obtain a matching result. The matching result here can represent whether the matching is successful. If the matching is successful, the matching result can also represent the candidate entity that matches the original entity. Immediately, the command text of the user can be responded according to the matching result, such as calling the phone number of contact ×××, playing the song ×××, etc., so as to ensure the smooth completion of the entire interaction process.

[0090] Based on any of the above embodiments, the present invention provides an interaction method applicable to multiple languages. It can be understood that the business requirements of different products are not the same. Therefore, for different products, the business scope to which the interaction method is applied is also different. For example, the business scope supported by smart speakers includes music, reminders, weather, schedules, chatting, etc., and the business scope supported by smart cars includes navigation, music, phone calls, radio stations, command control, etc. Especially in vertical scenarios such as navigation, music, and phone calls, there are usually problems with noun declensions in the Indo-European language family. The embodiments of the present invention will be described in detail by taking the business scope supported by smart cars as an example.

[0091] Figure 6 is the second schematic flowchart of the interaction method provided by the present invention. As Figure 6 shown, the implementation process of this interaction method is specifically as follows:

[0092] S1. Train an end-to-end model by encoding the corpus with the BPE algorithm to obtain a general model suitable for multiple languages:

[0093] Considering that current word vectors are usually trained in a single-language (Chinese or English) corpus and are not universal, and their embedding layer cannot cover the semantic information of multiple languages. Therefore, in order to better integrate the information of multiple languages in a shared vocabulary, the embodiments of the present invention use BPE in text preprocessing, and iteratively replace the most frequent symbol pairs (original bytes) in the given dataset with a single unused symbol, so that the processed vocabulary is insensitive to the type information of the language and more retains the organizational structure information of the language. In the embodiments of the present invention, by constructing the embedding of a multilingual parallel corpus with the BPE algorithm, a general dictionary shared by all languages can be obtained, and this general dictionary can significantly improve the alignment effect of different languages in the embedding space. The embodiments of the present invention sample sentences from a multilingual corpus for BPE learning from a random multinomial distribution. In order to ensure balanced corpus and the universality of the multilingual corpus, the sampling of sentences follows a multinomial distribution:

[0094]

[0095] where: n i : the corpus volume of the i-th language (n k is the same); the sum of the corpora of all languages; p i : the probability of the i-th language corpus in the total language corpus (p j is the same); α: hyperparameter. Preferably, in order to solve the problem of unbalanced corpus distribution among languages, the hyperparameter can be set to 0.5; the sum of the probabilities of all language corpora; q i : the sampling rate of the i-th language.

[0096] In the embodiments of the present invention, the BPE algorithm is used for all languages involved in the Indo-European language family and sentences with the same meaning, that is, the second sample text, to obtain the sample text vector of the second sample text, and a large number of training corpora are constructed therefrom to train a pre-trained model independent of language information, and this pre-trained model is used as a multi-language general model.

[0097] The general model sequentially includes an encoder, a decoder, a fully connected layer, and a softmax. During the training process, the input of the encoder is the sample text vector of the second sample text (each word vector can be a 1024-dimensional BPE embedding), and the output is the intermediate vector of the second sample text. Then, the target vector of the second sample text is obtained through the decoder, and named entities are extracted according to the target vector.

[0098] Optionally, the encoder can adopt a neural network model combining Bi-LSTM (Bi-directional Long-Short Term Memory) and max-pooling; the decoder can adopt LSTM (Long Short-Term Memory), which generates a sequence representing a task chain through loops. Each node in the sequence represents a task, and the generation of a single task depends on the output of the encoder and the previous task, that is, the input of the decoder is the output of the encoder and the previous historical feature.

[0099] During the training process, the objective function of the general model is the maximum conditional likelihood function for the joint training of the encoder and the decoder:

[0100]

[0101] where θ is the parameter of the model, N is the number of samples in the training set, x θ is the sample input to the model, y is the predicted result output by the model, and p θ is the probability that the model prediction is correct.

[0102] S2. For the language of the command text, fine-tune the general model, and extract the inflected entities in the command text based on the entity extraction model obtained by fine-tuning:

[0103] Based on step S1, by fine-tuning the trained general model, an end-to-end entity extraction model for each language can be obtained. The specific fine-tuning method here can be to add a mapping layer on the basis of the general model to obtain an initial entity extraction model that can map back to entities in each language, and then train the initial entity extraction model with the labeled corpus related to each language. For example, for the language to which the command text belongs, a mapping layer corresponding to this language can be added on the basis of the general model first to obtain the initial entity extraction model corresponding to this language, and then the sample text vector obtained by BPE encoding of the first sample text and the sample entity corresponding to the first sample text are used to train the initial entity extraction model, so as to obtain the entity extraction model corresponding to this language after fine-tuning.

[0104] In the application process, input the command text into the entity extraction model of the corresponding language, and the key slot information slot in the command text can be extracted, thus completing the task of named entity extraction. However, in the case where the language of the command text belongs to the Indo-European language family, there is a problem of noun declension. At this time, the key slot information extracted is not what is actually desired, which will cause the subsequent interaction process to be unable to proceed. In the embodiments of the present invention, the key slot information extracted at this time is called a declension entity.

[0105] S3. According to the declension entity and the case information of the language corresponding to the command text, use a classification model to obtain the case to which the declension entity belongs in the command text:

[0106] In the embodiments of the present invention, a classification model is used to classify the case of the declension entity to obtain the case to which the declension entity belongs in the command text. Optionally, the classification model can adopt CNN.

[0107] The training process is as follows:

[0108] 1) Corpus: Reuse the first sample text in a specific language used for fine-tuning before, and label the case to which the sample entity belongs in the first sample text

[0109] 2) Objective function: Mean squared error loss function

[0110]

[0111] where θ is the parameter of the model, k is the number of samples in the training set, PT(θ) i is the labeled case, and PF(θ) i is the case obtained by the classification model.

[0112] The application process is as follows:

[0113] 1) Input: The extracted inflected entity slots, the token representation of the previous token arranged before the inflected entity in the command text output by the encoder in the entity extraction model, and the case information of the corresponding language of the command text (that is, how many inflected forms a language has. For example, Russian has 6 inflected forms: nominative, genitive, dative, accusative, instrumental, prepositional).

[0114] 2) Output: The case ID of the inflected entity in the command text

[0115] It can be understood that the classification model in the embodiments of the present invention is a model common to multiple languages of the Indo-European language family. By inputting the case information of the corresponding language of the command text, each case in that language can be obtained, and then it can be selected which case the inflected entity belongs to in the command text. For example, for Russian corpora, such a classification problem becomes a 6-classification problem.

[0116] S4. According to the inflected entity slot, the belonging case ID, and the specific rules of the corresponding language, restore the original entity:

[0117] Through the entity extraction model and classification model in steps S2 and S3, the inflected entity and the case to which the inflected entity belongs in the command text can be obtained. On this basis, the original form of the entity slot, that is, the original entity, can be conveniently and accurately restored through the specific rules of the corresponding language, including the relationship between various case-form combinations and inflection rules in that language. The restored original entity can match the stored entity, thereby completing the subsequent interaction process.

[0118] The method provided in the embodiments of the present invention performs BPE encoding on multi-language data expressing the same meaning to obtain a unified embedding, thereby constructing a training corpus to train an end-to-end general model to implement a sequence labeling task. For the language of the command text, that is, the target language, an entity extraction model can be obtained after fine-tuning. When the language of the command text belongs to the Indo-European language family, the inflected entity in the command text can be obtained by applying the entity extraction model. Then, through a classification model and adding the inflection information of the target language, the case to which the inflected entity belongs in the command text in the target language can be obtained. Immediately, according to the inflected entity and the case to which it belongs in the command text, and the specific rules of the target language, the original entity can be restored, thereby solving the noun inflection problem existing in Indo-European language interactions, improving the effect of multi-language interactions in the Indo-European language family, and at the same time, there is no need to train models for multiple languages separately, and only fine-tuning for a single language is required, greatly enhancing the user's interaction experience.

[0119] The interactive device provided by the present invention will be described below. The interactive device described below can be mutually referred to the interactive method described above.

[0120] Based on any of the above embodiments, the present invention provides an interaction device. Figure 7 It is a schematic structural diagram of the interaction device provided by the present invention, as Figure 7 shown, the device includes:

[0121] An acquisition unit 710, configured to acquire a command text of a user;

[0122] An extraction unit 720, configured to perform entity extraction on the command text to obtain a declined entity in the command text and determine the case position to which the declined entity belongs in the command text when the language of the command text belongs to the Indo-European language family;

[0123] A restoration unit 730, configured to determine a primitive entity corresponding to the declined entity based on the case position to which the declined entity belongs in the command text;

[0124] A response unit 740, configured to respond to the command text based on the primitive entity.

[0125] The device provided by the embodiment of the present invention performs entity extraction on the command text to obtain a declined entity in the command text and determine the case position to which the declined entity belongs in the command text when the language of the command text belongs to the Indo-European language family. Based on this, the primitive entity corresponding to the declined entity is restored, and then the command text is responded to based on the primitive entity, thereby solving the problem of noun declension existing in Indo-European language interactions, improving the accuracy and effectiveness of Indo-European language interactions, and greatly enhancing the user's interaction experience.

[0126] Based on any of the above embodiments, determining the case position to which the declined entity belongs in the command text includes:

[0127] Based on the token representation of the previous tokenization, select, from each case position under the language of the command text, the case position to which the declined entity belongs in the command text, where the previous token is the token arranged before the declined entity in the command text.

[0128] Based on any of the above embodiments, the restoration unit 730 includes:

[0129] A part-of-speech determination subunit, configured to select the part of speech to which the declined entity belongs from each part of speech under the language of the command text;

[0130] A declension restoration subunit, configured to determine a primitive entity corresponding to the declined entity based on the part of speech to which the declined entity belongs and the case position to which the declined entity belongs in the command text.

[0131] Based on any of the above embodiments, the declension restoration subunit is specifically configured to:

[0132] Determine the declension rule corresponding to the part of speech and case of the declension entity based on the relationship between various pre-set combinations of word part-of-speech cases and declension rules;

[0133] Based on the declension rule corresponding to the part of speech and case of the declension entity, perform declension reduction on the declension entity to obtain the original entity corresponding to the declension entity.

[0134] Based on any of the above embodiments, perform entity extraction on the command text to obtain the declension entities in the command text, including:

[0135] Encode the command text based on a general dictionary to obtain the text vector of the command text, where the general dictionary is constructed based on multilingual parallel corpora;

[0136] Based on the text vector of the command text, perform entity extraction on the command text to obtain the declension entities in the command text.

[0137] Based on any of the above embodiments, based on the text vector of the command text, perform entity extraction on the command text to obtain the declension entities in the command text, including:

[0138] Based on an entity extraction model, perform entity extraction on the text vector to obtain the declension entities;

[0139] The entity extraction model is trained based on the sample text vectors and sample entities of the first sample text in the language of the command text on the basis of a general model, and the general model is trained based on the sample text vectors and sample entities of the second sample text in multiple languages.

[0140] Based on any of the above embodiments, the response unit 740 is used for:

[0141] Match the original entity with multiple pre-stored candidate entities to obtain a matching result, and respond to the command text based on the matching result.

[0142] Figure 8 Illustrates a schematic diagram of the entity structure of an electronic device, such as Figure 8As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute an interaction method, which includes: obtaining a command text of a user; when the language of the command text belongs to the Indo-European language family, performing entity extraction on the command text to obtain a declined entity in the command text, and determining the case position to which the declined entity belongs in the command text; based on the case position to which the declined entity belongs in the command text, determining the original entity corresponding to the declined entity; and based on the original entity, responding to the command text.

[0143] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0144] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the interaction method provided by the above-mentioned various methods, which includes: obtaining a command text of a user; when the language of the command text belongs to the Indo-European language family, performing entity extraction on the command text to obtain a declined entity in the command text, and determining the case position to which the declined entity belongs in the command text; based on the case position to which the declined entity belongs in the command text, determining the original entity corresponding to the declined entity; and based on the original entity, responding to the command text.

[0145] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements an interaction method provided by the above-mentioned various methods. The method includes: obtaining a command text of a user; when the language of the command text belongs to the Indo-European language family, performing entity extraction on the command text to obtain declined entities in the command text, and determining the case position to which the declined entities belong in the command text; based on the case position to which the declined entities belong in the command text, determining the original entities corresponding to the declined entities; and based on the original entities, responding to the command text.

[0146] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0147] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An interaction method, characterized in that, including: obtaining a command text of a user; when the language of the command text belongs to the Indo-European language family, performing entity extraction on the command text to obtain an inflected entity in the command text, and determining a case position to which the inflected entity belongs in the command text; determining, based on the case position to which the inflected entity belongs in the command text, a primitive entity corresponding to the inflected entity, where the primitive entity is an entity in a primitive form corresponding to the inflected entity; the case position to which the inflected entity belongs in the command text is used to characterize a structural and semantic relationship between a noun corresponding to the inflected entity and other words in the command text; responding to the command text based on the primitive entity; the responding to the command text based on the primitive entity includes: matching the primitive entity with a plurality of pre-stored candidate entities to obtain a matching result, and responding to the command text based on the matching result; the determining the case position to which the inflected entity belongs in the command text includes: selecting, based on a word segmentation representation of a previous word segmentation, a case position to which the inflected entity belongs in the command text from each case position in the language of the command text, where the previous word segmentation is a word segmentation arranged before the inflected entity in the command text.

2. The interactive method according to claim 1, wherein the determining, based on the case position to which the inflected entity belongs in the command text, a primitive entity corresponding to the inflected entity includes: selecting a part of speech to which the inflected entity belongs from each part of speech in the language of the command text; determining a primitive entity corresponding to the inflected entity based on the part of speech to which the inflected entity belongs and the case position to which the inflected entity belongs in the command text.

3. The interactive method according to claim 2, wherein the determining a primitive entity corresponding to the inflected entity based on the part of speech to which the inflected entity belongs and the case position to which the inflected entity belongs in the command text includes: determining an inflection rule corresponding to the part of speech and the case position to which the inflected entity belongs based on a relationship between various part-of-speech case combinations and inflection rules set in advance; performing inflection reduction on the inflected entity based on the inflection rule corresponding to the part of speech and the case position to which the inflected entity belongs to obtain a primitive entity corresponding to the inflected entity.

4. The interactive method according to claim 1, wherein the performing entity extraction on the command text to obtain an inflected entity in the command text includes: encoding the command text based on a general dictionary to obtain a text vector of the command text, where the general dictionary is constructed based on parallel corpora of multiple languages; performing entity extraction on the command text based on the text vector of the command text to obtain an inflected entity in the command text.

5. The interactive method according to claim 4, wherein the performing entity extraction on the command text based on the text vector of the command text to obtain an inflected entity in the command text includes: performing entity extraction on the text vector based on an entity extraction model to obtain the inflected entity; the entity extraction model is trained based on a sample text vector and sample entities of a first sample text in the language of the command text on the basis of a general model, and the general model is trained based on a sample text vector and sample entities of a second sample text in multiple languages.

6. An interaction device, characterized in that, including: An acquisition unit for acquiring a user's command text; An extraction unit for performing entity extraction on the command text to obtain inflected entities in the command text and determining the case position to which the inflected entities belong in the command text when the language of the command text belongs to the Indo-European language family; A restoration unit for determining an original entity corresponding to the inflected entity based on the case position to which the inflected entity belongs in the command text, where the original entity is an entity in the original form corresponding to the inflected entity; the case position to which the inflected entity belongs in the command text is used to characterize the structural and semantic relationship between the noun corresponding to the inflected entity and other words in the command text; A response unit for responding to the command text based on the original entity; Specifically, the response unit is configured to: match the original entity with a plurality of pre-stored candidate entities to obtain a matching result, and respond to the command text based on the matching result; Specifically, the extraction unit is configured to: Based on the word segmentation representation of the previous word segmentation, select the case position to which the inflected entity belongs in the command text from each case position in the language of the command text, where the previous word segmentation is the word segmentation arranged before the inflected entity in the command text.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, the interactive method according to any one of claims 1 to 5 is implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the interactive method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Server used for recognizing browser voice commands and browser voice command recognition system

    CN102629246A