Model generation method, word sense disambiguation method, device, medium, and equipment

By automatically constructing samples using parallel corpus and preset interpretation sets, the deep learning model is generated, and the problem of high data dependence in the existing technology is solved and efficient word meaning disambiguation training is achieved.

CN115017986BActive Publication Date: 2025-07-22BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210613094.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-07-22
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

In the prior art, building a deep learning model for word meaning disambiguation requires a large amount of manual annotation training data, resulting in high data dependence and low efficiency.

Method used

By obtaining multiple sets of parallel corpus and preset interpretation sets, multiple samples are automatically constructed, and the first classification model is generated using machine learning, reducing manual annotation, and semi-supervised training is achieved.

Benefits of technology

Reduces the data dependence of training deep learning models, improves training efficiency and accuracy, and reduces the need for manual annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115017986B_ABST
    Figure CN115017986B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a model generation method, a word sense disambiguation method, an apparatus, a medium, and a device. The model generation method includes: obtaining multiple groups of parallel corpora, where each group of parallel corpora includes a first text and a second text that are translations of each other, the first text belongs to a first language, and the second text belongs to a second language; determining multiple first samples according to the multiple groups of parallel corpora and a preset paraphrase set, each first sample including the first text, a first original word in the first text, and a first paraphrase in the preset paraphrase set that matches a translation word of the first original word, the translation word being a word in the second text that matches the first original word, and the first paraphrase belonging to the second language; and generating a first classification model according to the multiple first samples. The present disclosure can reduce the data dependence for generating the first classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligent technologies, and in particular, to a model generation method, a word sense disambiguation method, an apparatus, a medium, and a device. Background Art

[0002] Word sense disambiguation refers to identifying the interpretation of a word or phrase in a given context. There are many words with multiple meanings in many languages. For polysemous words, it is of great significance to distinguish the specific meaning of the word in the context. Due to the complexity of language, it is very challenging to automatically distinguish the interpretations in different contexts by computer algorithms. In related technologies, word sense disambiguation can be based on a deep learning model. However, obtaining a deep learning model requires a large amount of training data as samples, and the construction of the samples requires a lot of manpower. Summary of the Invention

[0003] This part of the content is provided to briefly introduce the concepts, which will be described in detail in the following detailed implementation part. This part of the content is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0004] In a first aspect, the present disclosure provides a model generation method, including:

[0005] Obtaining multiple groups of parallel corpora, each group of the parallel corpora including a first text and a second text that are translations of each other, the first text belonging to a first language, and the second text belonging to a second language;

[0006] Determining multiple first samples according to the multiple groups of the parallel corpora and a preset paraphrase set, each first sample including the first text, a first original word in the first text, and a first paraphrase that matches the translation word of the first original word in the preset paraphrase set, the translation word being a word in the second text that matches the first original word, and the first paraphrase belonging to the second language;

[0007] Generating a first classification model according to the multiple first samples.

[0008] In a second aspect, the present disclosure provides a word sense disambiguation method, including:

[0009] Obtaining a target text and a word to be disambiguated in the target text, the target text belonging to a first language;

[0010] Processing the target text and the word to be disambiguated according to the first classification model to determine a target paraphrase of the word to be disambiguated for the target text, the target paraphrase belonging to a second language, and the first classification model is obtained according to the method described in the first aspect.

[0011] In a third aspect, the present disclosure provides a model generation device, including:

[0012] A first acquisition module, configured to acquire multiple groups of parallel corpora, each group of the parallel corpora including a first text and a second text that are translations of each other, the first text belonging to a first language, and the second text belonging to a second language;

[0013] A first determination module, configured to determine multiple first samples according to the multiple groups of the parallel corpora and a preset paraphrase set, each first sample including the first text, a first original word in the first text, and a first paraphrase that matches the translation word of the first original word in the preset paraphrase set, the translation word being a word in the second text that matches the first original word, and the first paraphrase belonging to the second language;

[0014] A first generation model, configured to generate a first classification model according to the multiple first samples.

[0015] In a fourth aspect, the present disclosure provides a word sense disambiguation device, including:

[0016] A second acquisition module, configured to acquire a target text and a disambiguation word in the target text, the target text belonging to the first language;

[0017] A second determination module, configured to process the target text and the disambiguation word according to the first classification model to determine a target paraphrase of the disambiguation word for the target text, the target paraphrase belonging to the second language, and the first classification model is obtained according to the method described in the first aspect.

[0018] In a fifth aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the method described in any one of the first aspect and the second aspect are implemented.

[0019] In a sixth aspect, the present disclosure provides an electronic device, including:

[0020] A storage device, on which at least one computer program is stored;

[0021] At least one processing device, configured to execute the at least one computer program in the storage device to implement the steps of the method described in any one of the first aspect and the second aspect.

[0022] Through the above technical solution, multiple first samples are automatically constructed by multiple groups of parallel corpora and a preset paraphrase set, that is, the training data required for generating the first classification model is automatically constructed, and the first paraphrase included in each first sample can be used as the label of the first sample. Since the first paraphrase is obtained by matching the translation of the first original word in the preset paraphrase set, the labels of multiple first samples do not need to be manually labeled, and semi-supervised training of the first classification model can be realized, reducing the data dependence of the training of the first classification model.

[0023] Other features and advantages of the present disclosure will be described in detail in the following specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In combination with the accompanying drawings and with reference to the following specific implementation manners, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale. In the drawings:

[0025] Figure 1 is a schematic diagram of an implementation environment shown according to an exemplary embodiment of the present disclosure.

[0026] Figure 2 is a flowchart of a model generation method shown according to an exemplary embodiment of the present disclosure.

[0027] Figure 3 is a flowchart of a word sense disambiguation method shown according to an exemplary embodiment of the present disclosure.

[0028] Figure 4 is a block diagram of a model generation device shown according to an exemplary embodiment of the present disclosure.

[0029] Figure 5 is a block diagram of a word sense disambiguation device shown according to an exemplary embodiment of the present disclosure.

[0030] Figure 6 is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the accompanying drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0032] It should be understood that the various steps recited in the method embodiments of the present disclosure may be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0033] As used herein, the term "comprising" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0034] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0035] It should be noted that the modification of "one" and "a plurality of" mentioned in the present disclosure is illustrative rather than restrictive. Those skilled in the art should understand that unless clearly stated otherwise in the context, it should be understood as "one or more".

[0036] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0037] All actions of obtaining signals, information or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining the authorization given by the corresponding device owner.

[0038] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users and the authorization of users should be obtained through appropriate means in accordance with relevant laws and regulations.

[0039] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.

[0040] As an optional but non-limiting implementation, in response to receiving an active request from a user, the way of sending a prompt message to the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0041] It can be understood that the above notification and the process of obtaining user authorization are only illustrative and do not limit the implementation of the present disclosure. Other ways that comply with relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0042] At the same time, it can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0043] Figure 1 is a schematic diagram of an implementation environment shown according to an exemplary embodiment of the present disclosure. As Figure 1 shown, the implementation environment may include: a model training device 110 and a model using device 120. In some embodiments, the model training device 110 may be a computer device such as a computer or a server, and is used to train a first classification model and a second classification model. The model training device 110 may train the first classification model and the second classification model by means of machine learning. For the training process of the first classification model and the second classification model, reference may be made to step 230 and its related description below, which will not be elaborated here.

[0044] The trained second classification model may be deployed in the model using device 120 for use. The model using device 120 may be a terminal device such as a mobile phone, a tablet computer, a personal computer, a multimedia playback device, etc., or may also be a server. The model using device 120 may process the target text and the disambiguation word in the target text through the second classification model to obtain the target interpretation of the disambiguation word for the target text. For specific details of obtaining the target interpretation, reference may be made to the following Figure 3 and its related description, which will not be elaborated here.

[0045] Figure 2 is a flowchart of a model generation method shown according to an exemplary embodiment of the present disclosure. As Figure 2 shown, the model generation method includes the following steps.

[0046] Step 210, obtain multiple groups of parallel corpora, each group of parallel corpora including a first text and a second text that are translations of each other, the first text belonging to a first language and the second text belonging to a second language.

[0047] In some embodiments, the first text and the second text that are translations of each other may refer to texts in different languages that express the same meaning. The first text and the second text may be sentence-level texts or text at the discourse level. A discourse may be an overall unit of language composed of a series of consecutive words, phrases, clauses, sentences, or paragraphs, and the text at the discourse level may be the text formed by this overall unit of language. For example, papers, books, periodicals, etc. It can be understood that the text at the discourse level may include a large number (e.g., four thousand, ten thousand, etc.) of characters.

[0048] In some embodiments, the specific languages of the first language and the second language may be determined according to the actual situation. For example, the first language is English and the second language is Chinese. Another example is that the first language is English and the second language is German. The present disclosure does not impose any restrictions on this.

[0049] Step 220: Determine a plurality of first samples according to multiple sets of parallel corpora and a preset paraphrase set. Each first sample includes the first text, the first original word in the first text, and the first paraphrase that matches the translation word of the first original word in the preset paraphrase set. The translation word is the word in the second text that matches the first original word, and the first paraphrase belongs to the second language.

[0050] In some embodiments, the preset paraphrase set may include all the paraphrases in the second language corresponding to a large number of words in the first language that have been pre-statistically calculated. It can be understood that each first original word in the first text of each set of parallel corpora can match a corresponding paraphrase in the second language in the preset paraphrase set.

[0051] In some embodiments, the preset paraphrase set may include one, which may include all the paraphrases in the second language corresponding to each first original word in the first text of each set of parallel corpora. Correspondingly, the first paraphrase in each first sample may be the paraphrase that the translation word of the first original word matches in this one preset paraphrase set.

[0052] In some embodiments, the preset paraphrase set may include multiple ones. Each preset paraphrase set may include all the paraphrases in the second language corresponding to each first original word. Correspondingly, the first paraphrase in each first sample may be the paraphrase that the translation word of the first original word matches in the target preset paraphrase set, and the target preset paraphrase set may be the preset paraphrase set that matches the first original word.

[0053] In some embodiments, there are multiple preset paraphrase sets. According to multiple sets of parallel corpora and the preset paraphrase sets, multiple first samples are determined, including: for each first original word in the first text of each set of parallel corpora, a translation word matching the first original word is determined in the second text, and according to the similarity between the translation word and each paraphrase in the target preset paraphrase set, a matching first paraphrase is determined for the translation word. The target preset paraphrase set is the preset paraphrase set that matches the first original word among the multiple preset paraphrase sets; multiple first samples are determined according to the first text in each set of parallel corpora, each first original word in the first text, and the first paraphrase matched by the translation word of each first original word.

[0054] In some embodiments, according to the word alignment result of the first text and the second text, a translation word matching the first original word can be determined in the second text. Specific details regarding word alignment can be found in the relevant description below and will not be elaborated here.

[0055] As described above, when there are multiple preset paraphrase sets, each preset paraphrase set can include all the paraphrases of each first original word in the second language. In some embodiments, the target preset paraphrase set for the translation word can be determined in multiple ways. For example, the target preset paraphrase set matching the identifier can be determined in the multiple preset paraphrase sets through the identifier of the first original word corresponding to the translation word.

[0056] In some embodiments, the similarity between the translation word and each paraphrase in the target preset paraphrase set can be determined according to the string edit distance or the word vector distance. The string distance can refer to the minimum number of edits required to modify the translation word to each paraphrase. Edits can include, but are not limited to, modification, insertion, and deletion, etc.

[0057] In some embodiments, the word vector distance can refer to the distance between the word vector of the translation word and the word vectors of each paraphrase. The word vectors of the translation word and each paraphrase can be obtained through encoding. The distance between word vectors includes, but is not limited to, cosine distance, Euclidean distance, Manhattan distance, Mahalanobis distance, or Minkowski distance, etc.

[0058] In some embodiments, the smaller the string edit distance or the word vector distance, the greater the similarity. The first paraphrase matched by the translation word can include paraphrases with a similarity greater than a preset threshold. The preset threshold can be specifically determined according to the actual situation. For example, the preset threshold can be 0.96 or 0.98, etc. In some embodiments, the first paraphrase can be one or more. When the first paraphrase is one, the first paraphrase can be the paraphrase with the greatest similarity.

[0059] In the embodiments of the present disclosure, by calculating the similarity only between the translation word and each interpretation in the target preset interpretation set, that is, by calculating the similarity only between the translation word and each interpretation in the preset interpretation set matched by the first original word corresponding to the translation word, the calculation amount of the similarity calculation can be greatly reduced, so as to improve the matching efficiency of the first interpretation.

[0060] In some embodiments, the model generation method further includes: for each group of parallel corpora, performing word segmentation on the first text and the second text in the parallel corpus respectively to obtain a first word sequence and a second word sequence; performing word alignment processing on the first word sequence and the second word sequence to determine a first mapping relationship between each first original word in the first text and each translation word in the second text; determining the translation word matching the first original word in the second text, including: according to the first mapping relationship, determining the translation word matching the first original word in the second text.

[0061] In some embodiments, a word segmentation tool can be used to perform word segmentation on the first text and the second text respectively, splitting the texts in two languages into word sequences, that is, splitting the first text in the first language and the second text in the second language into word sequences to obtain a first word sequence and a second word sequence. In some embodiments, a word alignment tool can be used to perform word alignment processing on the first word sequence and the second word sequence. For example, the fast_align tool package or the GIZA++ tool package can be used.

[0062] Exemplarily, taking the first text as the sentence-level text S a and the second text as the sentence-level text S b as an example, if S a is split to obtain the first word sequence S b and the second word sequence obtained by splitting is then the mapping relationship between the first original word a a in S i and the translation word b b in S j can be established through word alignment to obtain the first mapping relationship.

[0063] Step 230, generating a first classification model according to multiple first samples.

[0064] In some embodiments, generating a first classification model according to multiple first samples may include: using each first interpretation in each first sample as the label of the first sample (for example, the first label described later), and generating a first classification model according to multiple first samples. In some embodiments, a first classification model can be generated according to multiple first samples by means of machine learning.

[0065] In the embodiments of the present disclosure, multiple first samples are automatically constructed through multiple groups of parallel corpora and a preset paraphrase set, that is, the training data required for generating the first classification model is automatically constructed, and the first paraphrase included in each first sample can be used as the label of the first sample. Since the first paraphrase is obtained by matching the translation of the first original word in the preset paraphrase set, the labels of the multiple first samples do not need to be manually labeled, and semi-supervised training of the first classification model can be realized, reducing the data dependence on the training of the first classification model.

[0066] Since the first paraphrase is obtained by matching the translation of the first original word in the preset paraphrase set, errors may occur in the matching process. Therefore, when using the first paraphrase as the label of the first sample, label noise will be included in the first sample. To reduce the influence of label noise on the training effect of the first classification model, a second classification model can be introduced during the training process of the first classification model.

[0067] In some embodiments, generating a first classification model according to multiple first samples includes: processing the multiple first samples according to a pre-trained second classification model to determine a second paraphrase of the first original word in each first sample, where the second paraphrase belongs to a second language; generating a first classification model according to the multiple first samples carrying the second paraphrase.

[0068] In some embodiments, the first classification model and the second classification model may refer to a word sense disambiguation model. It should be noted that the goal of the word sense disambiguation model is to identify the meaning of a word in context. For example, let s represent a sentence, w represent a word in sentence s, and G w = {g w,1 , g w,2 , ……, g w,N} represent the set composed of all paraphrases of word w. The word sense disambiguation model can be used to judge which one of the meanings of word w in sentence s and those in G w is equivalent. Therefore, the word sense disambiguation model can be regarded as a text classification model.

[0069] In some embodiments, the second classification model may be a cross - language word - sense disambiguation model, which can identify the second paraphrase in a second language corresponding to the first original word in a first language. When the second classification model is a cross - language word - sense disambiguation model, the second classification model can be trained as follows: Obtain a plurality of second samples according to preset text data, where each second sample includes a third text, a second original word in the third text, and a third paraphrase of the second original word. The third text belongs to the first language, and the third paraphrase belongs to the second language; Iteratively update the parameters of the initial second classification model according to the plurality of second samples to reduce the first loss function value corresponding to each second sample, and obtain a trained second classification model; wherein, the first loss function value corresponding to each second sample is determined through the following process: Process the second sample through the second classification model to obtain a first predicted paraphrase; Determine the first loss function value based at least on the difference between the first predicted paraphrase and the third paraphrase.

[0070] In some embodiments, the preset text data may include dictionary data in the corresponding language. For example, taking the first language as English and the second language as Chinese, the preset text data may be an English - Chinese bilingual dictionary. Multiple second samples can be obtained through the dictionary data in the corresponding language. For example, multiple second samples can be obtained through English sentences in the English - Chinese bilingual dictionary and the Chinese paraphrases of each word in the English sentences. In some embodiments, the third paraphrase in the second sample can be obtained through manual annotation. Correspondingly, the second classification model can be obtained through supervised training.

[0071] In some embodiments, each second sample may include a third text, a second original word in the third text, and a third paraphrase of the second original word. That is, each second sample can be a triple containing a sentence, a word, and a paraphrase. Taking the previous sentence s, word w, and paraphrase set G w and paraphrase g w as an example, the triple form of each second sample can be D=(s, w, g w ).

[0072] Since the second classification model is a cross - language word - sense disambiguation model, the third text and the second original word in the second sample for training the second classification model are in the first language, and the third paraphrase is in the second language. For example, taking the triple form of the previous second sample as an example, the sentence s and the word w in the second sample can be in English, and the paraphrase g w is in Chinese.

[0073] During the training process of the cross - language word - sense disambiguation model (i.e., the second classification model), the parameters of the second classification model can be continuously updated based on multiple second samples. Exemplarily, the parameters of the second classification model can be continuously adjusted to reduce the first loss function values corresponding to each second sample, such that the first loss function values meet the preset conditions. For example, the loss function values converge, or the loss function values are less than a preset value. When the first loss function values meet the preset conditions, the model training is completed, and a trained second classification model is obtained.

[0074] In some embodiments, the second classification model can be a monolingual word - sense disambiguation model, and the monolingual word - sense disambiguation model can identify the second paraphrase of the first language corresponding to the first original word of the first language. When the second classification model is a monolingual word - sense disambiguation model, processing multiple first samples according to the pre - trained second classification model to determine the second paraphrase of the first original word in each first sample includes: processing multiple first samples according to the pre - trained second classification model to determine the fourth paraphrase of each first sample, where the fourth paraphrase belongs to the first language; determining the second paraphrase according to the second mapping relationship for the fourth paraphrase of each first sample.

[0075] In some embodiments, the second mapping relationship can refer to the mapping relationship of bilingual paraphrases, that is, the mapping relationship between the fourth paraphrase of the first language and the second paraphrase of the second language. The second mapping relationship can be obtained by statistical analysis of the corresponding language dictionaries. For example, it can be obtained by statistical analysis of the bilingual paraphrases in an English - Chinese bilingual dictionary. By using the monolingual word - sense disambiguation model as the second classification model, the training speed of the monolingual word - sense disambiguation model is faster, thereby making the training efficiency of the second classification model higher.

[0076] In some embodiments, when the second classification model is a monolingual word - sense disambiguation model, the second classification model can be trained as follows: obtaining multiple third samples according to preset text data, where each third sample includes a fourth text, a third original word in the fourth text, and a fifth paraphrase of the third original word, and the fifth paraphrase belongs to the first language; iteratively updating the parameters of the initial second classification model according to the multiple third samples to reduce the second loss function values corresponding to each third sample, and obtaining a trained second classification model; where the second loss function values corresponding to each third sample are determined through the following process: processing the third sample through the second classification model to obtain a second predicted paraphrase; determining the second loss function value based at least on the difference between the second predicted paraphrase and the fifth paraphrase. The training method of the monolingual word - sense disambiguation model is similar to the training method of the aforementioned cross - language word - sense disambiguation model, and the details here can be referred to the relevant description above and will not be elaborated here.

[0077] In some embodiments, the second paraphrase obtained by the second classification model processing the first sample can be used as the label of the first sample (e.g., the second label described later), so as to generate the first classification model through multiple first samples carrying the second paraphrase. Since the first paraphrase is also the label of the first sample, when training the first classification model according to multiple first samples carrying the second paraphrase, the first sample carries two labels.

[0078] In some embodiments, the second classification model can be used as the word sense disambiguation teacher model, and the first classification model can be used as the word sense disambiguation student model, so as to train and generate the first classification model according to multiple first samples carrying the second paraphrase through the knowledge distillation algorithm. Knowledge distillation can make the output of the word sense disambiguation student model close to (or fit) the output of the word sense disambiguation model.

[0079] In some embodiments, generating the first classification model according to multiple first samples carrying the second paraphrase includes: iteratively updating the parameters of the initial first classification model according to multiple first samples to reduce the target loss function value corresponding to each first sample, and generating the first classification model; wherein, the target loss function value corresponding to each first sample is determined through the following process: processing the first sample through the first classification model to obtain the target predicted paraphrase; determining the target loss function value based at least on the first difference between the target predicted paraphrase and the first label, and the second difference between the target predicted paraphrase and the second label, where the first label and the second label are the first paraphrase and the second paraphrase respectively.

[0080] The target predicted paraphrase can be the output of the first classification model (i.e., the word sense disambiguation student model), the first label can be the true label of the training data of the first classification model, the second label can be the output of the second classification model (i.e., the word sense disambiguation teacher model), and the target loss function value in the training process of the first classification model can be determined through the target predicted paraphrase, the first label, and the second label.

[0081] Exemplarily, the target loss function value can be obtained through the following formula (1):

[0082]

[0083] where λ is the weight, Loss is the target loss function value, is the loss function value determined according to the output of the first classification model and the output of the second classification model, that is, the loss function value determined according to the second difference, for example, the determined cross-entropy value, is the loss function value determined according to the output of the first classification model and the true label, that is, the loss function value determined according to the first difference, for example, the cross-entropy value.

[0084] During the training process of the first classification model, the parameters of the first classification model can be continuously updated based on multiple first samples. Exemplarily, the parameters of the first classification model can be continuously adjusted to reduce the target loss function values corresponding to each first sample, so that the target loss function values meet the preset conditions. For example, the target loss function values converge, or the target loss function values are less than the preset values. When the target loss function values meet the preset conditions, the model training is completed, and the trained first classification model is obtained. The trained first classification model can process the target text in the first language and the disambiguation word in the target text to obtain the target paraphrase in the second language. Understandably, the trained first classification model is used for cross - language word sense disambiguation.

[0085] In the embodiments of the present disclosure, the target loss function value of the first classification model is constructed through the first difference and the second difference, that is, the first classification model is trained by the method of knowledge distillation, and a second classification model is used to guide the training of the first classification model, so that the output of the first classification model is close to the output of the second classification model, which can reduce the influence of the label noise of the automatically generated data, that is, reduce the influence of the incorrect first paraphrase existing in the first samples during the training of the first classification model, and improve the accuracy performance of the first classification model.

[0086] Figure 3 is a flowchart of a word sense disambiguation method shown according to an exemplary embodiment of the present disclosure. As Figure 3 shown, the word sense disambiguation method includes the following steps.

[0087] Step 310, obtain the target text and the disambiguation word in the target text, and the target text belongs to the first language.

[0088] Step 320, process the target text and the disambiguation word according to the first classification model, and determine the target paraphrase of the disambiguation word for the target text, and the target paraphrase belongs to the second language.

[0089] In some embodiments, the disambiguation word can be a word whose cross - language paraphrase needs to be identified. The target text can be the text where the disambiguation word is located, and the context of the disambiguation word can be obtained through the target text. By processing the target text and the disambiguation word through the first classification model, the cross - language paraphrase of the disambiguation word in the context of the target text can be obtained. For example, if the target text and the disambiguation word are in English, the Chinese paraphrase of the disambiguation word can be obtained. The first classification model can be the model trained according to the above - mentioned Figure 2 steps. For the specific details of the first classification model, reference can be made to the above - mentioned Figure 2 and its related descriptions, which will not be elaborated here.

[0090] Figure 4 is a block diagram of a model generation device shown according to an exemplary embodiment of the present disclosure. AsFigure 4 As shown, the model generation device 400 includes:

[0091] A first acquisition module 410, configured to acquire multiple groups of parallel corpora, each group of the parallel corpora including a first text and a second text that are translations of each other, the first text belonging to a first language, and the second text belonging to a second language;

[0092] A first determination module 420, configured to determine multiple first samples according to multiple groups of the parallel corpora and a preset paraphrase set, each first sample including the first text, a first original word in the first text, and a first paraphrase that matches the translation word of the first original word in the preset paraphrase set, the translation word being a word in the second text that matches the first original word, and the first paraphrase belonging to the second language;

[0093] A first generation model 430, configured to generate a first classification model according to the multiple first samples.

[0094] In some embodiments, the first generation model 430 is further configured to:

[0095] Process the multiple first samples according to a pre-trained second classification model to determine a second paraphrase of the first original word in each first sample, the second paraphrase belonging to the second language;

[0096] Generate a first classification model according to the multiple first samples carrying the second paraphrase.

[0097] In some embodiments, there are multiple preset paraphrase sets, and the first determination module 420 is further configured to:

[0098] For each first original word in the first text of each group of the parallel corpora, determine a translation word that matches the first original word in the second text, and determine a matching first paraphrase for the translation word according to the similarity between the translation word and each paraphrase in the target preset paraphrase set, the target preset paraphrase set being the preset paraphrase set that matches the first original word among the multiple preset paraphrase sets;

[0099] Determine multiple first samples according to the first text in each group of the parallel corpora, each first original word in the first text, and the first paraphrase that matches the translation word of each first original word.

[0100] In some embodiments, the model generation device 400 further includes:

[0101] A word segmentation module, configured to perform word segmentation on the first text and the second text in each group of the parallel corpora respectively, to obtain a first word sequence and a second word sequence;

[0102] A word alignment module, configured to perform word alignment processing on the first word sequence and the second word sequence, to determine a first mapping relationship between each first original word in the first text and each translated word in the second text;

[0103] The first determination module 420 is further configured to:

[0104] According to the first mapping relationship, determine the translated word in the second text that matches the first original word.

[0105] In some embodiments, the second classification model is trained as follows:

[0106] Obtain a plurality of second samples according to preset text data, each second sample includes a third text, a second original word in the third text, and a third paraphrase of the second original word, the third text belongs to the first language, and the third paraphrase belongs to the second language;

[0107] Iteratively update the parameters of the initial second classification model according to the plurality of second samples, to reduce the first loss function value corresponding to each second sample, and obtain a trained second classification model;

[0108] Wherein, the first loss function value corresponding to each second sample is determined through the following process:

[0109] Process the second sample through the second classification model to obtain a first predicted paraphrase;

[0110] Determine the first loss function value based at least on the difference between the first predicted paraphrase and the third paraphrase.

[0111] In some embodiments, the first generation model 430 is further configured to:

[0112] Process a plurality of the first samples according to the pre-trained second classification model, to determine a fourth paraphrase of each first sample, the fourth paraphrase belongs to the first language;

[0113] Determine the second paraphrase according to the second mapping relationship for the fourth paraphrase of each first sample.

[0114] In some embodiments, the second classification model is trained as follows:

[0115] Obtain a plurality of third samples according to preset text data, each of the third samples including a fourth text, a third original word in the fourth text, and a fifth interpretation of the third original word, the fifth interpretation belonging to the first language;

[0116] Iteratively update the parameters of the initial second classification model according to the plurality of third samples to reduce the second loss function values corresponding to each third sample, and obtain a trained second classification model;

[0117] Among them, the second loss function values corresponding to each third sample are determined through the following process:

[0118] Process the third sample through the second classification model to obtain a second predicted interpretation;

[0119] Determine the second loss function value based at least on the difference between the second predicted interpretation and the fifth interpretation.

[0120] In some embodiments, the first generation model 430 is further configured to:

[0121] Iteratively update the parameters of the initial first classification model according to the plurality of first samples to reduce the target loss function values corresponding to each first sample, and generate the first classification model;

[0122] Among them, the target loss function values corresponding to each first sample are determined through the following process:

[0123] Process the first sample through the first classification model to obtain a target predicted interpretation;

[0124] Determine the target loss function value based at least on a first difference between the target predicted interpretation and a first label, and a second difference between the target predicted interpretation and a second label, the first label and the second label being the first interpretation and the second interpretation respectively.

[0125] Figure 5 is a block diagram of a word sense disambiguation device shown according to an exemplary embodiment of the present disclosure. As Figure 5 shown, the word sense disambiguation device 500 includes:

[0126] A second acquisition module 510, configured to acquire a target text and a word to be disambiguated in the target text, the target text belonging to the first language;

[0127] A second determination module 520, configured to process the target text and the word to be disambiguated according to the first classification model, and determine a target interpretation of the word to be disambiguated for the target text, the target interpretation belonging to the second language.

[0128] Next, refer to Figure 6, which shows a schematic structural diagram of an electronic device (such as the terminal device or server in Figure 1 ) 600 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0129] As Figure 6 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 602 or the programs loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0130] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0131] Particularly, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.

[0132] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0133] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (for example, a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (for example, the Internet), and end-to-end networks (for example, ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0134] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately and not be assembled into the electronic device.

[0135] The above computer-readable medium stores one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain multiple sets of parallel corpora, each set of the parallel corpora including a first text and a second text that are translations of each other, the first text belonging to a first language and the second text belonging to a second language; determine multiple first samples according to the multiple sets of the parallel corpora and a preset paraphrase set, each first sample including the first text, a first original word in the first text, and a first paraphrase that matches the translation word of the first original word in the preset paraphrase set, the translation word being a word in the second text that matches the first original word, and the first paraphrase belonging to the second language; and generate a first classification model according to the multiple first samples.

[0136] Alternatively, the above computer-readable medium stores one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain a target text and a disambiguation word in the target text, the target text belonging to a first language; process the target text and the disambiguation word according to the first classification model to determine a target paraphrase of the disambiguation word for the target text, the target paraphrase belonging to a second language.

[0137] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0139] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of a module does not constitute a limitation on the module itself.

[0140] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, the types of hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0141] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0142] According to one or more embodiments of the present disclosure, Example 1 provides a method for generating a model, including:

[0143] Obtain multiple groups of parallel corpora, where each group of the parallel corpora includes a first text and a second text that are translations of each other, the first text belongs to a first language, and the second text belongs to a second language;

[0144] According to multiple groups of the parallel corpora and a preset paraphrase set, determine multiple first samples, where each first sample includes the first text, a first original word in the first text, and a first paraphrase in the preset paraphrase set that matches the translation word of the first original word. The translation word is a word in the second text that matches the first original word, and the first paraphrase belongs to the second language;

[0145] Generate a first classification model according to the multiple first samples.

[0146] According to one or more embodiments of the present disclosure, Example 2 provides a model generation method of Example 1. The generating a first classification model according to the multiple first samples includes:

[0147] Process the multiple first samples according to a pre-trained second classification model to determine a second paraphrase of the first original word in each first sample, and the second paraphrase belongs to the second language;

[0148] Generate a first classification model according to the multiple first samples carrying the second paraphrase.

[0149] According to one or more embodiments of the present disclosure, Example 3 provides a model generation method of Example 1. The preset paraphrase set includes multiple ones. The determining multiple first samples according to multiple groups of the parallel corpora and the preset paraphrase set includes:

[0150] For each first original word in the first text of each group of the parallel corpora, determine a translation word in the second text that matches the first original word, and determine a matching first paraphrase for the translation word according to the similarity between the translation word and each paraphrase in the target preset paraphrase set. The target preset paraphrase set is the preset paraphrase set that matches the first original word among the multiple preset paraphrase sets;

[0151] Determine multiple first samples according to the first text in each group of the parallel corpora, each first original word in the first text, and the first paraphrase that matches the translation word of each first original word.

[0152] According to one or more embodiments of the present disclosure, Example 4 provides a model generation method of Example 3. The method further includes:

[0153] For each group of the parallel corpora, perform word segmentation on the first text and the second text in the parallel corpora respectively to obtain a first word sequence and a second word sequence;

[0154] Perform word alignment processing on the first word sequence and the second word sequence to determine a first mapping relationship between each first original word in the first text and each translated word in the second text;

[0155] Determining the translated word in the second text that matches the first original word includes:

[0156] According to the first mapping relationship, determine the translated word in the second text that matches the first original word.

[0157] According to one or more embodiments of the present disclosure, Example 5 provides a model generation method for Example 2, and the second classification model is trained according to the following method:

[0158] Obtain a plurality of second samples according to preset text data, each second sample includes a third text, a second original word in the third text, and a third paraphrase of the second original word, the third text belongs to the first language, and the third paraphrase belongs to the second language;

[0159] Iteratively update the parameters of the initial second classification model according to the plurality of second samples to reduce the first loss function value corresponding to each second sample, and obtain a trained second classification model;

[0160] Wherein, the first loss function value corresponding to each second sample is determined through the following process:

[0161] Process the second sample through the second classification model to obtain a first predicted paraphrase;

[0162] Determine the first loss function value based at least on the difference between the first predicted paraphrase and the third paraphrase.

[0163] According to one or more embodiments of the present disclosure, Example 6 provides a model generation method for Example 6. Processing the plurality of first samples according to the pre-trained second classification model to determine a second paraphrase of the first original word in each first sample includes:

[0164] Process the plurality of first samples according to the pre-trained second classification model to determine a fourth paraphrase of each first sample, and the fourth paraphrase belongs to the first language;

[0165] Determine the second paraphrase according to the second mapping relationship for the fourth paraphrase of each first sample.

[0166] According to one or more embodiments of the present disclosure, Example 7 provides a model generation method for Example 6, and the second classification model is trained according to the following method:

[0167] Obtain a plurality of third samples according to preset text data, each of the third samples including a fourth text, a third original word in the fourth text, and a fifth interpretation of the third original word, the fifth interpretation belonging to the first language;

[0168] Iteratively update the parameters of the initial second classification model according to the plurality of third samples to reduce the second loss function value corresponding to each third sample, and obtain a trained second classification model;

[0169] Among them, the second loss function value corresponding to each third sample is determined through the following process:

[0170] Process the third sample through the second classification model to obtain a second predicted interpretation;

[0171] Determine the second loss function value based at least on the difference between the second predicted interpretation and the fifth interpretation.

[0172] According to one or more embodiments of the present disclosure, Example 8 provides a model generation method for Example 2. The method of generating a first classification model according to the plurality of first samples carrying the second interpretation includes:

[0173] Iteratively update the parameters of the initial first classification model according to the plurality of first samples to reduce the target loss function value corresponding to each first sample, and generate the first classification model;

[0174] Among them, the target loss function value corresponding to each first sample is determined through the following process:

[0175] Process the first sample through the first classification model to obtain a target predicted interpretation;

[0176] Determine the target loss function value based at least on a first difference between the target predicted interpretation and a first label, and a second difference between the target predicted interpretation and a second label, the first label and the second label being the first interpretation and the second interpretation respectively.

[0177] According to one or more embodiments of the present disclosure, Example 9 provides a word sense disambiguation method, including:

[0178] Obtain a target text and a disambiguation word in the target text, the target text belonging to the first language;

[0179] Process the target text and the disambiguation word according to the first classification model to determine a target interpretation of the disambiguation word for the target text, the target interpretation belonging to the second language, and the first classification model is obtained according to the method described in Example 1.

[0180] According to one or more embodiments of the present disclosure, Example 10 provides a model generation device, including:

[0181] A first acquisition module, configured to acquire multiple groups of parallel corpora, each group of the parallel corpora including a first text and a second text that are translations of each other, the first text belonging to a first language, and the second text belonging to a second language;

[0182] A first determination module, configured to determine multiple first samples according to the multiple groups of the parallel corpora and a preset paraphrase set, each first sample including the first text, a first original word in the first text, and a first paraphrase that matches the translation word of the first original word in the preset paraphrase set, the translation word being a word in the second text that matches the first original word, and the first paraphrase belonging to the second language;

[0183] A first generation model, configured to generate a first classification model according to the multiple first samples.

[0184] According to one or more embodiments of the present disclosure, Example 11 provides the model generation device of Example 10, and the first generation model is further configured to:

[0185] Process the multiple first samples according to a pre-trained second classification model to determine a second paraphrase of the first original word in each first sample, the second paraphrase belonging to the second language;

[0186] Generate a first classification model according to the multiple first samples carrying the second paraphrase.

[0187] According to one or more embodiments of the present disclosure, Example 12 provides the model generation device of Example 10, the preset paraphrase set includes multiple, and the first determination module is further configured to:

[0188] For each first original word in the first text of each group of the parallel corpora, determine a translation word that matches the first original word in the second text, and determine a matching first paraphrase for the translation word according to the similarity between the translation word and each paraphrase in the target preset paraphrase set, the target preset paraphrase set being the preset paraphrase set that matches the first original word among the multiple preset paraphrase sets;

[0189] Determine multiple first samples according to the first text in each group of the parallel corpora, each first original word in the first text, and the first paraphrase that matches the translation word of each first original word.

[0190] According to one or more embodiments of the present disclosure, Example 13 provides the model generation device of Example 12, and the model generation device further includes:

[0191] A word segmentation module, configured to perform word segmentation on the first text and the second text in each group of the parallel corpora respectively to obtain a first word sequence and a second word sequence;

[0192] A word alignment module, configured to perform word alignment processing on the first word sequence and the second word sequence to determine a first mapping relationship between each first original word in the first text and each translated word in the second text;

[0193] The first determination module is further configured to:

[0194] According to the first mapping relationship, determine the translated word that matches the first original word in the second text.

[0195] According to one or more embodiments of the present disclosure, Example 14 provides a model generation device of Example 11, and the second classification model is trained according to the following method:

[0196] Obtain a plurality of second samples according to preset text data, each second sample includes a third text, a second original word in the third text, and a third paraphrase of the second original word, the third text belongs to the first language, and the third paraphrase belongs to the second language;

[0197] Iteratively update the parameters of the initial second classification model according to the plurality of second samples to reduce the first loss function value corresponding to each second sample, and obtain a trained second classification model;

[0198] Wherein, the first loss function value corresponding to each second sample is determined through the following process:

[0199] Process the second sample through the second classification model to obtain a first predicted paraphrase;

[0200] Determine the first loss function value based at least on the difference between the first predicted paraphrase and the third paraphrase.

[0201] According to one or more embodiments of the present disclosure, Example 15 provides a model generation device of Example 11, and the first generation model is further configured to:

[0202] Process a plurality of the first samples according to the pre-trained second classification model to determine a fourth paraphrase of each first sample, and the fourth paraphrase belongs to the first language;

[0203] Determine the second paraphrase according to the second mapping relationship for the fourth paraphrase of each first sample.

[0204] According to one or more embodiments of the present disclosure, Example 16 provides the model generation device of Example 15, and the second classification model is trained as follows:

[0205] Obtain a plurality of third samples according to preset text data, each of the third samples includes a fourth text, a third original word in the fourth text, and a fifth paraphrase of the third original word, and the fifth paraphrase belongs to the first language;

[0206] Iteratively update the parameters of the initial second classification model according to the plurality of third samples to reduce the second loss function value corresponding to each third sample, and obtain the trained second classification model;

[0207] Wherein, the second loss function value corresponding to each third sample is determined through the following process:

[0208] Process the third sample through the second classification model to obtain a second predicted paraphrase;

[0209] Determine the second loss function value based at least on the difference between the second predicted paraphrase and the fifth paraphrase.

[0210] According to one or more embodiments of the present disclosure, Example 17 provides the model generation device of Example 11, and the first generation model is further configured to:

[0211] Iteratively update the parameters of the initial first classification model according to the plurality of first samples to reduce the target loss function value corresponding to each first sample, and generate the first classification model;

[0212] Wherein, the target loss function value corresponding to each first sample is determined through the following process:

[0213] Process the first sample through the first classification model to obtain a target predicted paraphrase;

[0214] Determine the target loss function value based at least on a first difference between the target predicted paraphrase and a first label, and a second difference between the target predicted paraphrase and a second label, where the first label and the second label are the first paraphrase and the second paraphrase respectively.

[0215] According to one or more embodiments of the present disclosure, Example 18 provides a word sense disambiguation device, including:

[0216] A second acquisition module, configured to acquire a target text and a disambiguation word to be disambiguated in the target text, and the target text belongs to the first language;

[0217] A second determination module, configured to process the target text and the disambiguation word according to a first classification model, and determine a target interpretation of the disambiguation word for the target text, where the target interpretation belongs to a second language, and the first classification model is obtained according to the method of any one of Examples 1 to 8.

[0218] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the above features are mutually replaced with the technical features (but not limited to) disclosed in the present disclosure that have similar functions to form a technical solution.

[0219] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0220] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims. Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

Claims

1. A model generation method, characterized in that, Including: Obtain multiple groups of parallel corpora, each group of the parallel corpora including a first text and a second text that are translations of each other, the first text belonging to a first language and the second text belonging to a second language; Determine multiple first samples according to the multiple groups of parallel corpora and a preset paraphrase set, each first sample including the first text, a first original word in the first text, and a first paraphrase matched by a translation word of the first original word in the preset paraphrase set, the translation word being a word in the second text that matches the first original word, the first paraphrase belonging to the second language, and the first paraphrase serving as a label of the first sample; Generate a first classification model according to the multiple first samples.

2. The model generation method according to claim 1, wherein The generating a first classification model according to the multiple first samples includes: Process the multiple first samples according to a pre-trained second classification model to determine a second paraphrase of the first original word in each first sample, the second paraphrase belonging to the second language; Generate a first classification model according to the multiple first samples carrying the second paraphrase.

3. The model generation method according to claim 1, wherein The preset paraphrase set includes multiple ones. The determining multiple first samples according to the multiple groups of parallel corpora and the preset paraphrase set includes: For each first original word in the first text of each group of the parallel corpora, determine a translation word in the second text that matches the first original word, and determine a matched first paraphrase for the translation word according to the similarity between the translation word and each paraphrase in a target preset paraphrase set, the target preset paraphrase set being a preset paraphrase set that matches the first original word among the multiple preset paraphrase sets; Determine multiple first samples according to the first text in each group of the parallel corpora, each first original word in the first text, and the first paraphrase matched by the translation word of each first original word.

4. The model generation method according to claim 3, wherein The method further includes: For each group of the parallel corpora, perform word segmentation on the first text and the second text in the parallel corpora respectively to obtain a first word sequence and a second word sequence; Perform word alignment processing on the first word sequence and the second word sequence to determine a first mapping relationship between each first original word in the first text and each translation word in the second text; The determining a translation word in the second text that matches the first original word includes: Determine the translation word in the second text that matches the first original word according to the first mapping relationship.

5. The model generation method according to claim 2, wherein The second classification model is trained according to the following method: Obtain multiple second samples according to preset text data, each second sample including a third text, a second original word in the third text, and a third paraphrase of the second original word, the third text belonging to the first language and the third paraphrase belonging to the second language; Iteratively update the parameters of an initial second classification model according to the multiple second samples to reduce the first loss function value corresponding to each second sample, and obtain a trained second classification model; Wherein, the first loss function value corresponding to each second sample is determined through the following process: Process the second sample through a second classification model to obtain a first predicted paraphrase; Determine a first loss function value based at least on the difference between the first predicted paraphrase and the third paraphrase.

6. The model generation method according to claim 2, wherein The determining the second paraphrase of the first original word in each of the first samples according to the pre-trained second classification model for processing a plurality of the first samples includes: Process a plurality of the first samples according to the pre-trained second classification model to determine a fourth paraphrase of each of the first samples, where the fourth paraphrase belongs to the first language; Determine the second paraphrase according to the fourth paraphrase of each of the first samples according to a second mapping relationship.

7. The model generation method according to claim 6, wherein the second classification model is trained according to the following method: Obtain a plurality of third samples according to preset text data, where each of the third samples includes a fourth text, a third original word in the fourth text, and a fifth paraphrase of the third original word, and the fifth paraphrase belongs to the first language; Iteratively update the parameters of the initial second classification model according to the plurality of third samples to reduce the second loss function value corresponding to each third sample, and obtain a trained second classification model; Among them, The second loss function value corresponding to each third sample is determined through the following process: Process the third sample through the second classification model to obtain a second predicted paraphrase; Determine a second loss function value based at least on the difference between the second predicted paraphrase and the fifth paraphrase.

8. The model generation method according to claim 2, wherein the generating a first classification model according to the plurality of first samples carrying the second paraphrase includes: Iteratively update the parameters of the initial first classification model according to the plurality of first samples to reduce the target loss function value corresponding to each first sample, and generate the first classification model; Wherein, the target loss function value corresponding to each first sample is determined through the following process: Process the first sample through the first classification model to obtain a target predicted paraphrase; Determine a target loss function value based at least on a first difference between the target predicted paraphrase and a first label, and a second difference between the target predicted paraphrase and a second label, where the first label and the second label are the first paraphrase and the second paraphrase respectively.

9. A method for word sense disambiguation, characterized in that, Includes: Obtain a target text and a disambiguation word in the target text, where the target text belongs to the first language; Process the target text and the disambiguation word according to the first classification model to determine a target paraphrase of the disambiguation word for the target text, where the target paraphrase belongs to the second language, and the first classification model is obtained according to the method of claim 1.

10. A model generation device, characterized in that, Includes: A first acquisition module configured to acquire multiple groups of parallel corpora, where each group of parallel corpora includes a first text and a second text that are translations of each other, the first text belongs to the first language, and the second text belongs to the second language; A first determination module, configured to determine a plurality of first samples according to multiple groups of the parallel corpora and a preset paraphrase set, each of the first samples including the first text, a first original word in the first text, and a first paraphrase matched by a translation word of the first original word in the preset paraphrase set, the translation word being a word in the second text that matches the first original word, the first paraphrase belonging to the second language, and the first paraphrase serving as a label of the first sample; A first generation model, configured to generate a first classification model according to the multiple first samples.

11. A word sense disambiguation device, characterized in that Comprising: A second acquisition module, configured to acquire a target text and a disambiguation word in the target text, the target text belonging to the first language; A second determination module, configured to process the target text and the disambiguation word according to the first classification model to determine a target paraphrase of the disambiguation word for the target text, the target paraphrase belonging to the second language, and the first classification model being obtained according to the method described in claim 1.

12. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processing device, the steps of the method described in any one of claims 1-9 are implemented.

13. An electronic device, characterized in that, Comprising: A storage device on which at least one computer program is stored; At least one processing device for executing the at least one computer program in the storage device to implement the steps of the method described in any one of claims 1-9.

Citation Information

Patent Citations

  • Translation method and device based on artificial intelligence, electronic equipment and storage medium

    CN111738025A