Corpus Processing Method, Metaphor Information Processing Method, Apparatus and Electronic Device
By using vacant triplets and language representation models in corpus processing, metaphor triplets are automatically generated, which solves the efficiency and accuracy of metaphorical component extraction relying on labeled data in the existing technology, and achieves efficient and accurate metaphorical component extraction.
Patent Information
- Application Number
- CN202210295700.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-23
AI Technical Summary
The existing metaphor component extraction scheme requires a large amount of labeling data, and the labeling data in real application scenarios is scarce, making it difficult to extract a large number of and accurate metaphor components from actual application scenarios.
By obtaining the vacant triplets in the corpus, determining the corpus-constructing template group based on the metaphor terms corresponding to the vacant, generating a new corpus with vacant, and predicting the probabilities of the vacant through the language representation model, determining the filler words to form a metaphor triplet.
This method does not require manual labeling of data, which can improve the efficiency and accuracy of metaphor component extraction. The generated metaphor triplets have high value in applications such as robot Q&A and literary creation.
Smart Images

Figure CN114662491B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of corpus processing, and in particular, to a corpus processing method, a metaphor information processing method, an apparatus, and an electronic device. Background Art
[0002] Creative language is widely used in literature, poetry, and AI interactions. In particular, metaphor rhetoric can stimulate readers' imagination and enhance the literary grace of sentences. In some tasks of metaphor recognition, there are mainly the following methods: (1) Using a certain amount of labeled data, adopting techniques such as dependency relations, syntactic analysis, and part-of-speech to extract metaphor components by setting rules; (2) Using a certain amount of labeled data sets to train a neural network model to extract metaphor components; (3) Using a pre-trained model to train a certain amount of labeled data sets to obtain a metaphor component extraction model, and then using this model to extract metaphor components.
[0003] In the above metaphor component extraction schemes, labeled data is required to be realized. However, the labeled data in real application scenarios is very scarce, resulting in the difficulty of extracting a large number of accurate metaphor components from actual application scenarios. Summary of the Invention
[0004] The purpose of this application is to provide a corpus processing method, a metaphor information processing method, an apparatus, and an electronic device to improve the efficiency and accuracy of metaphor component extraction.
[0005] In a first aspect, an embodiment of this application provides a corpus processing method, and the method includes: for the corpus in the target corpus dataset, obtaining the corresponding missing triple; where the missing triple includes at least one missing position, and the adjectives and / or nouns in the corpus; the missing position corresponds to one of the metaphor items of the ontology, the metaphor, or the attribute in the metaphor triple, and the attribute represents the common characteristics of the ontology and the metaphor; determining a corpus construction template group according to the metaphor item corresponding to the missing position in the missing triple; generating a new corpus with the missing position according to the corpus construction template in the corpus construction template group and the missing triple; predicting the word probability of the missing position corresponding to the corpus construction template through the new corpus and the trained language representation model; determining the filling word of the missing position according to the word probability of the missing position corresponding to the corpus construction template; adding the filling word to the missing position in the missing triple to obtain the metaphor triple corresponding to the corpus.
[0006] In combination with the first aspect, the embodiments of the present application provide a first possible implementation manner of the first aspect. Among them, the training of the language representation model includes: obtaining a corpus sample set; marking adjectives and / or nouns included in the corpus samples in the corpus sample set according to the inter-word dependency relationship to obtain marked corpus; performing masked language model training and next sentence prediction task training on the pre-trained language representation model according to the marked corpus to obtain a trained language representation model.
[0007] In combination with the first aspect, the embodiments of the present application provide a second possible implementation manner of the first aspect. Among them, each metaphor item in the metaphor triple has a corresponding corpus construction template group, and each corpus construction template group includes multiple different types of corpus construction templates; each type of corpus construction template includes at least two metaphor items; determining the corpus construction template group according to the metaphor item corresponding to the vacancy in the vacancy triple includes: searching for the corpus construction template group of the metaphor item corresponding to the vacancy in the vacancy triple from the corpus construction template groups respectively corresponding to each metaphor item.
[0008] In combination with the first aspect, the embodiments of the present application provide a third possible implementation manner of the first aspect. Among them, the method further includes: determining a metaphor triple set according to the metaphor corpus samples in the metaphor corpus sample set; performing masking processing on the metaphor items of the metaphor triples in the metaphor triple set in a masked manner to obtain the vacancy triple of the metaphor triple; selecting multiple templates including the metaphor items missing in the vacancy triple from multiple different types of corpus construction templates, and constructing multiple vacancy corpus according to the vacancy triple and the selected multiple templates; predicting the prediction result corresponding to each vacancy corpus through the language representation model; where the prediction result is the word list probability distribution information corresponding to the missing metaphor item; determining the optimal template combination corresponding to the missing metaphor item according to the prediction results corresponding to the selected multiple templates, and determining the optimal template combination as the corpus construction template group corresponding to the missing metaphor item.
[0009] In combination with the first aspect, the embodiments of the present application provide a fourth possible implementation manner of the first aspect. Among them, multiple different types of corpus construction templates include: a first type of template including three metaphor items of an ontology, a vehicle, and an attribute, a second type of template including two metaphor items of a vehicle and an attribute, a third type of template including two metaphor items of an ontology and an attribute, and a fourth type of template including two metaphor items of an ontology and a vehicle.
[0010] In combination with the first aspect, an embodiment of the present application provides a fifth possible implementation manner of the first aspect. Among them, determining the filler word for the vacancy corresponding to the template according to the word probability of the vacancy corresponding to the corpus includes: calculating the weighted average of the word probabilities of the vacancies corresponding to each corpus construction template; and determining the word with the largest weighted average word probability as the filler word for the vacancy.
[0011] In a second aspect, an embodiment of the present application further provides a metaphor information processing method. The method includes: listening for core information in the current application scenario; where the application scenario includes robot question answering or literary creation, and the core information is the tendency information characterizing the next application requirement in the application scenario; determining a target metaphor triple from a plurality of metaphor triples according to the core information, where the plurality of metaphor triples are obtained through the above-mentioned corpus processing method and include an ontology, a metaphor, and an attribute; and pushing the prompt information corresponding to the target metaphor triple; where the prompt information includes the target metaphor triple and / or the corpus set corresponding to the target metaphor triple.
[0012] In combination with the second aspect, an embodiment of the present application provides a first possible implementation manner of the first aspect. Among them, determining a target metaphor triple from a plurality of metaphor triples according to the core information includes: determining the metaphor triple containing the core information as the target metaphor triple; and / or determining the related word of the core information and determining the metaphor triple containing the related word as the target metaphor triple.
[0013] In a third aspect, an embodiment of the present application further provides a corpus processing device. The device includes: a vacancy triple acquisition module for obtaining the vacancy triple corresponding to the corpus in the target corpus dataset, where the vacancy triple includes at least one vacancy, as well as adjectives and / or nouns in the corpus; the vacancy corresponds to one of the ontology, metaphor, or attribute in the metaphor triple, and the attribute characterizes the common feature of the ontology and the metaphor; a template group determination module for determining a corpus construction template group according to the metaphor item corresponding to the vacancy in the vacancy triple; a corpus generation module for generating a new corpus with the vacancy according to the corpus construction template in the corpus construction template group and the vacancy triple; a prediction module for predicting the word probability of the vacancy corresponding to the corpus construction template through the new corpus and the trained language representation model; a filler word determination module for determining the filler word for the vacancy according to the word probability of the vacancy corresponding to the corpus construction template; and a metaphor triple determination module for adding the filler word to the vacancy in the vacancy triple to obtain the metaphor triple corresponding to the corpus.
[0014] Fourth aspect, an embodiment of the present application further provides a figurative information processing device, which includes: a listening module, configured to listen for core information in the current application scenario; wherein, the application scenario includes robot question answering or literary creation, and the core information is a tendency information representing the next application requirement in the application scenario; a target figurative triple determination module, configured to determine a target figurative triple from a plurality of figurative triples according to the core information, wherein the plurality of figurative triples are obtained by the above-mentioned corpus processing method and include an ontology, a vehicle, and an attribute; a prompt information pushing module, configured to push the prompt information corresponding to the target figurative triple; wherein, the prompt information includes the target figurative triple and / or the corpus set corresponding to the target figurative triple.
[0015] Fifth aspect, an embodiment of the present application further provides an electronic device, including a processor and a memory, where the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the above-mentioned corpus processing method or the above-mentioned figurative information processing method.
[0016] Sixth aspect, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores computer executable instructions, and when the computer executable instructions are called and executed by a processor, the computer executable instructions cause the processor to implement the above-mentioned corpus processing method or the above-mentioned figurative information processing method.
[0017] An embodiment of the present application provides a corpus processing method, a figurative information processing method, a device, and an electronic device. By using adjectives and / or nouns in the corpus and the vacancy, a vacancy triple is obtained, and then a corpus construction template group is determined according to the figurative item corresponding to the vacancy. A new corpus with a vacancy is generated according to the corpus construction template and the vacancy triple. The word probability of the vacancy corresponding to the corpus construction template is predicted by the new corpus and the language representation model, and the filling word of the vacancy is determined according to the word probability of the vacancy, so as to obtain the figurative triple corresponding to the above-mentioned corpus. By adopting the above technology, a new corpus with a vacancy is generated by a preset corpus construction template, and the filling word corresponding to the vacancy is predicted by the language representation model to obtain a complete triple, and then a figurative triple is obtained. This process does not require manual data annotation, and can improve the efficiency and accuracy of figurative component extraction compared with the existing figurative component extraction methods.
[0018] In addition, a large number of figurative triples can be obtained based on the above method. Whether in robot question answering (such as AI chat interaction) or literary creation, the above figurative triples can be used for relevant information prompts, enriching the content of the prompt information and enhancing the use value of the figurative triples. Description of the Drawings
[0019] To more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the specific embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0020] Figure 1 A flowchart of a corpus processing method provided by an embodiment of the present application;
[0021] Figure 2 Example diagrams of the first, second, third, and fourth types of templates provided by an embodiment of the present application;
[0022] Figure 3 An example diagram of a corpus processing method provided by an embodiment of the present application;
[0023] Figure 4 A flowchart of a metaphor information processing method provided by an embodiment of the present application;
[0024] Figure 5 A structural diagram of a corpus processing device provided by an embodiment of the present application;
[0025] Figure 6 A structural diagram of another corpus processing device provided by an embodiment of the present application;
[0026] Figure 7 A structural diagram of a metaphor information processing device provided by an embodiment of the present application;
[0027] Figure 8 A structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments
[0028] The following will clearly and completely describe the technical solutions of the present application in combination with the embodiments. Obviously, the described embodiments are some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0029] For the sake of convenience of description, taking literary creation through AI as an example, for example, creating literary works such as poems, couplets, lyrics, novels, news releases, etc. through AI.
[0030] With the further improvement of computer computing power and the accumulation of a large amount of text corpora, pre-trained language models (also known as language representation models) have greatly improved the performance of natural language generation. A pre-trained language model is an autoregressive language model trained using a large amount of unsupervised text corpora. Since there are many neural network parameters in the pre-trained language model and the model is pre-trained on a large scale of text, it has learned the most basic statistical features in natural language, can better capture the correlation between long texts, and improve the fluency of the generated text. Therefore, the pre-trained language model has high language fluency and correlation of long texts. Based on the pre-trained language model, the model parameters are fine-tuned with the corpora of specific types of literary works (such as novels), and the fine-tuned model can be used to create literary works of feature types.
[0031] Metaphor is a commonly used rhetorical device in literary creation, which can stimulate readers' imagination and enhance the literary grace of sentences. A metaphor usually includes three components: the tenor (or called the logical object, usually a noun), the vehicle (or called the comparison object, usually also a noun), and the attribute (used to represent the common characteristics of the tenor and the vehicle, usually an adjective). These three components can all be regarded as the metaphorical terms in the metaphorical triple, that is, the metaphorical triple in the embodiments of the present application includes: the tenor, the vehicle, and the attribute. Suppose a metaphorical sentence: "The white clouds are like cotton candies". In this metaphorical sentence, "white" is the attribute in the metaphorical component, used to represent the common characteristics of the clouds and the cotton candies, "clouds" is the tenor in the metaphorical component, and "cotton candies" is the vehicle in the metaphorical component; this metaphorical sentence can be transformed into a triple (tenor, attribute, vehicle), that is, (clouds, white, cotton candies). This triple indicates that the clouds can be compared to cotton candies in terms of the attribute of whiteness. When we have a large number of triples, we can use simple patterns to construct enough metaphorical sentences and apply them to various scenarios. Therefore, identifying the metaphorical components in the corpus to form triples is very important for creating metaphorical sentences through AI.
[0032] Considering that the labeled data in real application scenarios is very scarce, in order to extract a large number of accurate metaphorical components from actual application scenarios, a corpus processing method, a metaphor information processing method, a device, and an electronic device provided in the embodiments of the present application can improve the efficiency and accuracy of metaphorical component extraction. This technology can be applied to the creation process of various texts and literary works, especially to the creation process of literary works with specific text structures. This technology can also be applied to the sentiment analysis of various texts and literary works, especially to the sentiment analysis of literary works with specific text structures. In addition, this technology can also be applied to the human-computer interaction of AI chatbots to enhance the interest of human-computer interaction.
[0033] Figure 1The flowchart shows a corpus processing method provided by an embodiment of this application. This method is applied to electronic devices (such as mobile phones, computers, servers, etc.). The method includes the following steps:
[0034] Step S102: For the corpus in the target corpus dataset, obtain the corresponding vacancy triple. The vacancy triple includes at least one vacancy, as well as adjectives and / or nouns in the corpus. The vacancy corresponds to one of the metaphorical terms of the ontology, metaphor, or attribute in the metaphor triple, and the attribute represents the common characteristics of the ontology and the metaphor.
[0035] Regarding the above step S102, the acquisition method of the above target corpus dataset can be through web crawling, through optical character recognition, directly provided by a third party, etc., and is not limited thereto. The corpus in the above target corpus dataset includes adjectives and / or nouns, and the adjectives and / or nouns included in the corpus can be combined into one or more vacancy triples. For example, a certain corpus includes "cloud" and "cotton candy". If these two nouns are to be used to create a metaphorical sentence, one noun can be used as the ontology in the metaphor triple, and the other noun can be used as the metaphor in the metaphor triple, but the adjective that can be used as the attribute in the metaphor triple is missing. The "cloud" and "cotton candy" in this corpus can be combined into a vacancy triple with one vacancy (cloud, ____, cotton candy), and this vacancy corresponds to the attribute in the metaphor triple.
[0036] The corpus in the above target corpus dataset can be each corpus in the target corpus dataset, or, according to the usage needs, it can also be a part of the corpus in the target corpus dataset. For example: it can be a corpus with certain characteristics in the target corpus dataset, or a corpus randomly selected from the target corpus dataset, etc.
[0037] Step S104: Determine a corpus construction template group according to the metaphorical term corresponding to the vacancy in the vacancy triple.
[0038] Regarding the above step S104, multiple different types of corpus construction templates can be pre-constructed. Each type of corpus construction template includes at least two metaphorical terms of the ontology, metaphor, and attribute, and one of the two metaphorical terms can be vacant. For a vacancy triple corresponding to a certain corpus, since the metaphorical term corresponding to the vacancy in the vacancy triple is known, the corpus construction templates including this metaphorical term can be combined into the corpus construction template group corresponding to this corpus.
[0039] As a possible implementation, it is assumed that each metaphor item in the metaphor triple has a corresponding corpus construction template group, and each corpus construction template group contains multiple corpus construction templates of different types; each type of corpus construction template contains at least two metaphor items; based on this, the above step S104 may specifically include: searching for the corpus construction template group of the metaphor item corresponding to the vacancy in the vacancy triple from the corpus construction template groups corresponding to each metaphor item. Continuing with the previous example, the corpus construction templates are divided into types with attributes and types without attributes. For a vacancy triple (cloud, ____, marshmallow) formed by "cloud" and "marshmallow" in a certain corpus, since the vacancy corresponds to the attribute in the metaphor triple, in order to predict the filling word for this vacancy, the templates of the type with attributes can be found from all the corpus construction templates by searching, and the found corpus construction templates can be combined into the corpus construction template group corresponding to this corpus. By using the above operation method of searching for the corpus construction template group, the prediction requirements of the filling word can be met, and the determination efficiency of the corpus construction template group can be improved.
[0040] Step S106, generate a new corpus with a vacancy according to the corpus construction templates in the corpus construction template group and the vacancy triple.
[0041] For the above step S106, the known metaphor items and the vacancy in the vacancy triple can be combined into a new corpus with a vacancy by using the corpus construction templates in the determined corpus construction template group. Continuing with the previous example, for a vacancy triple (cloud, ____, marshmallow) formed by "cloud" and "marshmallow" in a certain corpus, after combining the corpus construction templates with attributes into the corpus construction template group corresponding to this corpus, the "cloud" and / or "marshmallow" in the vacancy triple can be used as seed words and combined with "____" by using the found corpus construction templates to form sentences containing the vacancy, such as "Clouds are like marshmallows ____", "Clouds are ____", "This ____ marshmallow", etc. Then, these sentences can be further combined into a new corpus, and then the filling word for the vacancy can be predicted from the new corpus.
[0042] Step S108, predict the word probability of the vacancy corresponding to the corpus construction template through the new corpus and the trained language representation model.
[0043] For the above step S108, the new corpus with vacancies can be input into the above trained language representation model, and the model is used to predict the filling words for the vacancies in each sentence in the new corpus, and the predicted words for each vacancy and the word probability of the predicted words are output. The word probability is used to indicate the probability value of the predicted word for the vacancy. The larger the probability value, the greater the probability that the predicted word is a suitable filling word, and the smaller the probability value, the smaller the probability that the predicted word is a suitable filling word. Continuing with the previous example, for the missing triplet (cloud, ____, marshmallow) composed of "cloud" and "marshmallow" in a certain corpus, one of the corpus construction templates can be used to combine "cloud" and "marshmallow" as seed words and "____" into a sentence containing a vacancy "Clouds are like marshmallows____", and input the sentence into the above-mentioned language representation model. The model outputs the predicted words for the vacancy in the sentence (that is, the vacancy corresponding to the corpus construction template), including "white", "soft" and "soft and fluffy", and the word probabilities of the three predicted words are 0.15, 0.1 and 0.75 respectively. Therefore, the possibility that "soft and fluffy" is the filler word in the sentence "Clouds are like marshmallows____" is greater than the possibility of "white" and "soft".
[0044] Step S110, determining a filling word for the vacancy according to the word probability of the vacancy corresponding to the corpus construction template.
[0045] For step S110, after predicting the word probability of the vacancy corresponding to each corpus construction template through the language representation model, the word probability distribution of all predicted words corresponding to the vacancy is obtained, and a word can be selected from the word probability distribution as the filler word for the vacancy. The specific selection method can be determined according to actual needs. For example, the word with the largest word probability is selected from the word probability distribution as the filler word, or based on the preset weight value of each template and the word probability of the vacancy corresponding to each corpus construction template, the word probability of each word is comprehensively calculated, and a word is selected as the filler word according to the comprehensive calculation result, etc., and there is no limitation on this.
[0046] As a possible implementation, the above step S110 may include: performing weighted average calculation on the word probability of the vacancy corresponding to each corpus construction template; and determining the word with the largest word probability after the weighted average as the filling word for the vacancy. Continuing with the previous example, for the missing triplet (cloud, ____, marshmallow) composed of "cloud", "marshmallow" and "____" in a certain corpus, the corpus construction template group includes two corpus construction templates, one of which corresponds to the word probabilities of "white", "soft" and "soft and fluffy" for the missing position, which are 0.15, 0.1 and 0.75 respectively, and the other corresponds to the word probabilities of "white", "soft" and "soft and fluffy" for the missing position, which are 0.2, 0.2 and 0.6 respectively. Certain weights are assigned to "white", "soft" and "soft and fluffy", and the weighted average calculation is performed on the word probabilities of the missing positions corresponding to the two corpus construction templates to obtain the word probabilities of "white", "soft" and "soft and fluffy" for the missing positions corresponding to the entire corpus construction template group, and the word "soft and fluffy" with the largest word probability is determined as the filling word for the missing position. The above-mentioned operation mode of taking weighted average of the word probabilities of the vacancy corresponding to each corpus construction template to determine the filling word of the vacancy can further improve the efficiency and accuracy of determining the filling word of the vacancy.
[0047] Step S112, adding a filler word to the vacant position in the vacant triplet to obtain a metaphor triplet corresponding to the corpus.
[0048] Continuing with the previous example, for the missing triple composed of "clouds" and "marshmallows" in a certain corpus (clouds, ____, marshmallows), after determining that the filler word for the missing position in the missing triple is "soft and mianmian", "soft and mianmian" can be added to the missing position to obtain the metaphor triple corresponding to the corpus (clouds, soft and mianmian, marshmallows).
[0049] The embodiment of the present application provides a corpus processing method, which obtains a vacant triple through adjectives and / or nouns and vacancies in the corpus, and then determines a corpus construction template group according to the metaphor item corresponding to the vacant position, and generates a new corpus with a vacant position according to the corpus construction template and the vacant triple, predicts the word probability of the vacant position corresponding to the corpus construction template through the new corpus and the language representation model, and determines the filler word of the vacant position according to the word probability of the vacant position, and then obtains the metaphor triple corresponding to the above corpus. Using the above technology, a new corpus with a vacant position is generated using a preset corpus construction template, and the filler word corresponding to the vacant position is predicted by the language representation model to obtain a complete triple, and then obtains a metaphor triple. This process does not require manual data annotation, and can improve the efficiency and accuracy of metaphor component extraction compared to the existing metaphor component extraction method.
[0050] Based on the above corpus processing method, in order to further meet the actual prediction needs of corresponding filler words for different metaphor items, the type classification method of the corpus construction template can be optimized. Specifically, multiple different types of corpus construction templates can be divided into four types, namely: the first type of template that includes three metaphor items: the tenor, the vehicle, and the attribute; the second type of template that includes two metaphor items: the vehicle and the attribute; the third type of template that includes two metaphor items: the tenor and the attribute; and the fourth type of template that includes two metaphor items: the tenor and the vehicle. Figure 2 An exemplary introduction to these four types of templates is given. Among them, T (tensor) represents the tenor, V (vehicle) represents the vehicle, A (attribute) represents the attribute, and p i represents the weight value corresponding to each template respectively. Figure 2 The corresponding in Figure 2 Take the first type and the second type of templates in
[0051] I: (11) This cloud is like cotton candy in terms of being ____; (12) Cotton candy is very ____, and so is the cloud; (13) The cloud is like cotton candy because they are both ____;
[0052] II: (21) This ____ cotton candy; (22) Cotton candy is very ____; (23) Cotton candy is ____;
[0053] Among them, the above "____" represents the blank position.
[0054] By adopting the above type classification method of the corpus construction template, the types of the corpus construction template cover the common combination ways of metaphor items in the corpus in the actual application scenario, thereby ensuring that the corpus construction template can meet the usage requirements of predicting corresponding filler words for different metaphor items.
[0055] In order to further improve the reliability of the above corpus construction template group, the above corpus processing method may further include:
[0056] (1) Determine the metaphor triple set according to the metaphor corpus samples in the metaphor corpus sample set.
[0057] Each metaphor corpus sample in the above metaphor corpus sample set is usually a single metaphor sentence that includes the tenor, the vehicle, and the attribute. The tenor, the vehicle, and the attribute in the single metaphor sentence can be combined into the metaphor triple of this metaphor sentence. In this way, the metaphor triples of each metaphor corpus sample in the metaphor corpus sample set can be obtained, and all the obtained metaphor triples are combined into the above metaphor triple set.
[0058] (2) Mask the metaphorical terms of the metaphorical triples in the metaphorical triple set in a masked manner to obtain the missing triples of the metaphorical triples.
[0059] A part of the metaphorical triples can be selected from the above metaphorical triple set as the objects for metaphorical term masking. The selection method can be determined according to actual needs. For example, a certain proportion of metaphorical triples can be randomly selected, or metaphorical triples can be selected at intervals of a certain number of sentences, etc. There is no limitation on this; then, the metaphorical terms of the selected part of the metaphorical triples are masked to obtain the corresponding missing triples. For example, assume that a certain metaphorical corpus sample is a single metaphorical sentence "The lake is like a transparent mirror", and the corresponding metaphorical triple of this sentence is (lake, transparent, mirror). After random masking, a new sentence "The lake is like [MASK][MASK] mirror" is obtained, and the new sentence corresponds to a missing triple (lake, [MASK][MASK], mirror).
[0060] (3) Select multiple templates containing the missing metaphorical terms in the missing triples from multiple different types of corpus construction templates, and construct multiple missing corpora according to the missing triples and the selected multiple templates.
[0061] Continuing with the previous example, for the case where the above-mentioned multiple different types of corpus construction templates are divided into the first type of template, the second type of template, the third type of template, and the fourth type of template, for the missing triple (lake, [MASK][MASK], mirror), the missing metaphorical term is an attribute. The first type of template, the second type of template, and the third type of template containing attributes can be selected for combination to obtain an initial template combination set containing multiple template combinations; among them, the combination method can be determined according to actual needs. For example, all possible template combinations can be listed by an exhaustive method, or specified by a third-party institution, specified manually according to experience, randomly combined, etc. There is no limitation on this. Use each corpus construction template in all the obtained template combinations to combine "lake", "[MASK][MASK]", and "mirror" into new sentences. Each new sentence can be considered a missing corpus, and in this way, multiple missing corpora are obtained.
[0062] (4) Through a language representation model, predict the prediction results corresponding to each missing corpus; among them, the prediction result is the word list probability distribution information corresponding to the missing metaphorical term.
[0063] After obtaining multiple pieces of the above-mentioned vacant corpus, all the obtained vacant corpus are input into the above-mentioned language representation model, and through the output of the model, the word list probability distribution information corresponding to the vacant metaphor item in each vacant corpus is obtained. This information is used to represent the word probability distribution of the predicted word corresponding to the vacant metaphor item in each vacant corpus. Its form can specifically be a table, a graph, etc. Its content includes the predicted word corresponding to the vacant metaphor item in each vacant corpus and the word probability of each filler word. The sum of the word probabilities of all the predicted words corresponding to the metaphor item in the same vacant corpus in this information is 1.
[0064] (5) Determine the optimal template combination corresponding to the vacant metaphor item according to the prediction results corresponding to the selected multiple templates, and determine the optimal template combination as the corpus construction template group corresponding to the vacant metaphor item.
[0065] Continuing with the previous example, after obtaining the set of initial template combinations corresponding to the vacant triple (lake water, [MASK][MASK], mirror), it is necessary to determine one of the template combinations in this set of template combinations as the optimal template combination. The determination method of the optimal template combination can be determined according to actual needs, such as being specified by a third-party institution, being specified by a person according to experience, being determined according to the recognition results of the metaphor components of a certain number (such as 100) of metaphorical sentences, etc. This is not limited. Then, determine the optimal template combination as the corpus construction template group corresponding to the vacant metaphor item.
[0066] Adopting the above-mentioned operation method for determining the corpus construction template group corresponding to the vacant metaphor item, the vacant triple can be obtained by the masking method, and multiple vacant corpus can be generated by selecting appropriate templates according to the vacant metaphor item in the vacant triple. Then, the word list probability distribution information corresponding to the vacant metaphor item is predicted by the model, so as to determine the optimal template combination corresponding to the vacant metaphor item as the corpus construction template group corresponding to the vacant metaphor item. This corpus construction template group has a relatively high reliability and can meet the actual prediction needs of the filler words corresponding to the vacant metaphor item, thereby further ensuring the accuracy of the metaphor component extraction result.
[0067] As a possible implementation manner, the training of the above-mentioned language representation model includes:
[0068] (1) Obtain a corpus sample set.
[0069] Various channels can be utilized to collect corpus, mainly including review data from various social media websites and e-commerce websites, People's Daily, Chinese Wikipedia, Baidu Encyclopedia, etc.; then, the collected corpus is cleaned, filtering out web links, tag information, etc., and retaining the pure text information containing nouns and adjectives; the filtered corpus is converted into a format that can be parsed by the language representation model. For paragraphs, they can be divided into combinations of multiple sentences in the form of a window size of 2, and each combination has exactly two sentences. Mask the nouns and / or adjectives in each sentence, and form the above-mentioned corpus sample set with the multiple sentences (i.e., multiple corpus samples) obtained after the masking process.
[0070] (2) According to the dependency relationship between words, mark the adjectives and / or nouns contained in the corpus samples in the corpus sample set to obtain the marked corpus.
[0071] Perform dependency syntax analysis on all sentences (i.e., all corpus samples) in the above-mentioned corpus sample set. For example, analyze the syntactic dependency relationship between words in a sentence in the way of a dependency syntax tree to obtain the dependency relationship between words in each sentence; according to the dependency relationship between words, such as the amod (adjectival modifier) relationship, etc., find and mark the nouns and / or adjectives in each sentence. For example, use the character <n>< / n> to mark nouns and use the character to mark adjectives. After marking, the above-mentioned marked corpus is obtained. That is, mark <n>and< / n> before and after the nouns in the corpus so that the noun is represented as <n>noun< / n> ; and mark and before and after the adjectives in the corpus so that the adjective is represented as adjective , and then the marked corpus is obtained.
[0072] (3) According to the marked corpus, perform masked language model (MLM, Mask Language Model) training and next sentence prediction task (NSP, Next Sentence Prediction) training on the pre-trained language representation model to obtain the trained language representation model.
[0073] As a possible implementation manner, the above-mentioned language representation model may include a BERT (Bidirectional Encoder Representation from Transformer) model. The BERT model usually contains a 12-layer bidirectional Transformer structure, the dimension of the initialized word vectors is 768, there are 12 attention layers, and the vocabulary size is 21128.
[0074] The pre-training of the language representation model includes two pre-training tasks, namely, the masked language model task and the next sentence prediction task. The masked language model task is used to predict the next word according to the previous context, and the next sentence prediction task is used to determine whether the current sentence is the next sentence of the previous sentence. The above-mentioned tokenized corpus is input into the language representation model for training. During the training process, the losses of the two tasks are added up to obtain the final loss, and the training is stopped until the model converges or the number of iterations reaches the preset number. After the training is completed, a pre-trained language representation model is obtained. Inputting a text into this model can obtain the feature representation of this text.
[0075] To facilitate the understanding of the corpus processing method provided by the embodiments of the present application, here Figure 3 an exemplary description of the above corpus processing method is as follows: (1) Collect the comment data of various social media websites and e-commerce websites, People's Daily, Chinese Wikipedia, Baidu Encyclopedia, etc. as the corpus, and clean the collected corpus to filter out web links, tag information, etc. Then, perform masking processing on the nouns and / or adjectives in each sentence to obtain the above-mentioned corpus sample set; (2) Analyze the syntactic dependency relationship between words in the sentence by using the dependency syntax tree, and find and mark the nouns and / or adjectives in each sentence according to the amod relationship to obtain the above-mentioned tokenized corpus; (3) Input the above-mentioned tokenized corpus into the BERT model for MLM training and NSP training to obtain a trained language representation model; (4) The above-mentioned target corpus dataset is directly provided by a third party. Combine the adjectives and / or nouns included in each corpus in the above-mentioned target corpus dataset into one or more missing triples; From the above-mentioned first type of template, second type of template, third type of template, and fourth type of template pre-constructed, find the template containing the missing metaphor item in the missing triple and the optimal template combination; (5) Use each template in the optimal template combination to combine the word corresponding to the existing metaphor item in the missing triple and the missing position corresponding to the missing metaphor item into a new corpus with a missing position; (6) Input the new corpus into the above-mentioned language representation model, predict the filling words for the missing positions in each sentence of the new corpus through this model, and output the word probability distribution of all the predicted words for the missing positions corresponding to each corpus construction template. Perform weighted average calculation on the word probability distributions corresponding to all the corpus construction templates obtained (this process can be regarded as a template fusion process) to obtain the word probability distribution corresponding to the missing position of the optimal template combination. The filling word for the missing position can be determined from this word probability distribution, and the filling word is correspondingly added to the missing position in the missing triple to obtain the corresponding metaphor triple.
[0076] Based on the above corpus processing method, the embodiments of the present application further provide a metaphor information processing method. Figure 4The figure is a schematic flowchart of a metaphor information processing method provided by an embodiment of the present application. This method is applied to an electronic device (such as a mobile phone, a computer, a server, etc.), and the method includes the following steps:
[0077] Step S402, monitor the core information in the current application scenario; wherein, the application scenario includes robot question answering or literary creation, and the core information is the tendency information representing the next application requirement in the application scenario.
[0078] For example, when a creator needs to create a literary work describing clouds, the core information is "clouds".
[0079] Step S404, determine a target metaphor triple from multiple metaphor triples according to the core information, wherein the multiple metaphor triples are obtained by the above-mentioned corpus processing method and include an ontology, a metaphor, and an attribute.
[0080] To ensure that the target metaphor triple meets the usage requirements of actual applications, the above step S404 may include: determining the metaphor triple containing the above core information as the target metaphor triple; and / or, determining the related words of the above core information, and determining the metaphor triple containing the related words as the target metaphor triple.
[0081] For example, if the core information is "clouds", the metaphor triple containing "clouds" can be determined as the target metaphor triple; after determining that the related word of "clouds" is "white", the metaphor triple containing "white" is determined as the target metaphor triple. By adopting the above method of determining the target metaphor triple by including a specific word, the determination efficiency of the target metaphor triple is improved.
[0082] Step S406, push the prompt information corresponding to the target metaphor triple; wherein, the prompt information includes the target metaphor triple and / or the corpus set corresponding to the target metaphor triple.
[0083] For example, after determining the target metaphor triple, the target metaphor triple itself can be directly pushed as the prompt information; for another example, a certain sentence template can be used to combine the metaphor components in the target metaphor triple into a sentence, and then the combined sentence can be pushed as the prompt information.
[0084] The metaphor information processing method provided by the embodiment of the present application can push a specified metaphor triple or the corpus corresponding to the specified metaphor triple according to the application requirement, so as to prompt the creator to create, bring inspiration to the creator and improve the creation efficiency of the creator, or enrich the interest of interaction during the robot question answering process.
[0085] Based on the above corpus processing method, the embodiment of the present application also provides a corpus processing device. Refer to Figure 5 as shown, the device includes:
[0086] The missing triple acquisition module 501 is configured to acquire the missing triples corresponding to the corpus in the target corpus dataset, where the missing triples include at least one missing position, and the adjectives and / or nouns in the corpus; the missing position corresponds to one of the metaphor terms in the metaphor triple, namely the noumenon, the vehicle, or the attribute, and the attribute characterizes the common features of the noumenon and the vehicle.
[0087] The template group determination module 502 is configured to determine a corpus construction template group according to the metaphor term corresponding to the missing position in the missing triples.
[0088] The corpus generation module 503 is configured to generate a new corpus with the missing position according to the corpus construction template in the corpus construction template group and the missing triples.
[0089] The prediction module 504 is configured to predict the word probability of the missing position corresponding to the corpus construction template through the new corpus and the trained language representation model.
[0090] The filler word determination module 505 is configured to determine the filler word of the missing position according to the word probability of the missing position corresponding to the corpus construction template.
[0091] The metaphor triple determination module 506 is configured to add the filler word to the missing position in the missing triples to obtain the metaphor triple corresponding to the corpus.
[0092] A corpus processing device provided by an embodiment of the present application obtains missing triples through adjectives and / or nouns and missing positions in the corpus, then determines a corpus construction template group according to the metaphor term corresponding to the missing position, and generates a new corpus with the missing position according to the corpus construction template and the missing triples. The word probability of the missing position corresponding to the corpus construction template is predicted through the new corpus and the language representation model, and the filler word of the missing position is determined according to the word probability of the missing position, so as to obtain the metaphor triple corresponding to the above-mentioned corpus. By adopting the above technology, a new corpus with a missing position is generated by using a preset corpus construction template, and the filler word corresponding to the missing position is predicted through the language representation model to obtain a complete triple, and then the metaphor triple is obtained. This process does not require manual annotation of data, and can improve the efficiency and accuracy of metaphor component extraction compared with the existing metaphor component extraction methods.
[0093] In the above metaphorical triple, each metaphorical item has a corresponding corpus construction template group, and each corpus construction template group contains multiple corpus construction templates of different types; each type of corpus construction template contains at least two metaphorical items; based on this, the above template group determination module 502 is further configured to: from the corpus construction template groups corresponding to each metaphorical item, find the corpus construction template group of the metaphorical item corresponding to the missing position in the missing triple.
[0094] The above multiple corpus construction templates of different types include: the first type of template containing three metaphorical items of ontology, vehicle, and attribute, the second type of template containing two metaphorical items of vehicle and attribute, the third type of template containing two metaphorical items of ontology and attribute, and the fourth type of template containing two metaphorical items of ontology and vehicle.
[0095] The above filler determination module 505 is further configured to: perform a weighted average calculation on the word probabilities of the missing positions corresponding to each corpus construction template; determine the word with the largest weighted average word probability as the filler for the missing position.
[0096] The above missing triple acquisition module 501 is further configured to: determine a metaphorical triple set according to the metaphorical corpus samples in the metaphorical corpus sample set; perform a masking process on the metaphorical items of the metaphorical triples in the metaphorical triple set in a masking manner to obtain the missing triples of the metaphorical triples.
[0097] The above corpus generation module 503 is further configured to: select multiple templates containing the missing metaphorical items in the missing triple from multiple corpus construction templates of different types, and construct multiple missing corpora according to the missing triple and the selected multiple templates.
[0098] The above prediction module 504 is further configured to: predict the prediction result corresponding to each missing corpus through the language representation model; wherein, the prediction result is the word list probability distribution information corresponding to the missing metaphorical item.
[0099] The above template group determination module 502 is further configured to: determine the optimal template combination corresponding to the missing metaphorical item according to the prediction results corresponding to the selected multiple templates, and determine the optimal template combination as the corpus construction template group corresponding to the missing metaphorical item.
[0100] Based on the above corpus processing device, an embodiment of the present application further provides another corpus processing device. Refer to Figure 6 As shown, this device further includes:
[0101] The training module 507 is used to obtain a corpus sample set; mark the adjectives and / or nouns contained in the corpus samples in the corpus sample set according to the inter-word dependency relationship to obtain a marked corpus; perform masked language model training and next sentence prediction task training on the pre-trained language representation model according to the marked corpus to obtain a trained language representation model.
[0102] Based on the above metaphor information processing method, the present application embodiment also provides a metaphor information processing device, see Figure 7 As shown, the device comprises:
[0103] The monitoring module 701 is used to monitor the core information in the current application scenario; wherein the application scenario includes robot question answering or literary creation, and the core information is the tendency information representing the next application demand in the application scenario.
[0104] The target metaphor triple determination module 702 is used to determine the target metaphor triple from multiple metaphor triples according to the core information, wherein the multiple metaphor triples are obtained by the above-mentioned corpus processing method and include a subject, a metaphor and an attribute.
[0105] The prompt information pushing module 703 is used to push the prompt information corresponding to the target metaphor triple; wherein the prompt information includes the target metaphor triple and / or the corpus corresponding to the target metaphor triple.
[0106] A metaphor information processing device provided in an embodiment of the present application can push specified metaphor triples or corpus corresponding to specified metaphor triples according to application requirements, thereby prompting creators to create, which can not only bring inspiration to creators but also improve the creator's creative efficiency, or enrich the fun of interaction during the robot question and answer process.
[0107] The above-mentioned determining the target metaphor triple from multiple metaphor triples based on the core information includes: determining the metaphor triple containing the core information as the target metaphor triple; and / or determining the associated words of the core information, and determining the metaphor triple containing the associated words as the target metaphor triple.
[0108] The implementation principle and technical effects of the above-mentioned device embodiment are the same as those of the corresponding method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding contents in the above-mentioned method embodiment.
[0109] The present application also provides an electronic device, such as Figure 8As shown, it is a schematic structural diagram of the electronic device. Among them, the electronic device 100 includes a processor 81 and a memory 80. The memory 80 stores computer-executable instructions that can be executed by the processor 81. The processor 81 executes the computer-executable instructions to implement the above-mentioned corpus processing method or the above-mentioned metaphor information processing method.
[0110] In Figure 8 the illustrated embodiment, the electronic device further includes a bus 82 and a communication interface 83. Among them, the processor 81, the communication interface 83, and the memory 80 are connected through the bus 82.
[0111] Among them, the memory 80 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 83 (which can be wired or wireless), a communication connection is realized between the system network element and at least one other network element. The Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 82 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 82 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 8 only a single bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0112] The processor 81 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 81 or the instructions in the form of software. The above-mentioned processor 81 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor 81 reads the information in the memory and combines its hardware to complete the steps of the corpus processing method or the metaphor information processing method in the foregoing embodiments.
[0113] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions cause the processor to implement the above-mentioned corpus processing method or the above-mentioned metaphor information processing method. For the specific implementation, reference can be made to the foregoing method embodiments, and details are not described herein again.
[0114] The computer program products of the corpus processing method, the metaphor information processing method, the device and the electronic device provided by the embodiments of the present application include a computer-readable storage medium storing program codes. The instructions included in the program codes can be used to execute the methods described in the foregoing method embodiments. For the specific implementation, reference can be made to the method embodiments, and details are not described herein again.
[0115] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present application.
[0116] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0117] In the description of this application, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to this application. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0118] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of this application, used to illustrate the technical solutions of this application, rather than limiting it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in this application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A corpus processing method, characterized in that, the method includes: For the corpus in the target corpus dataset, obtain the corresponding missing triple; wherein, the missing triple includes at least one missing position, as well as adjectives and / or nouns in the corpus; the missing position corresponds to one of the metaphor terms of the ontology, metaphor, or attribute in the metaphor triple, and the attribute represents the common characteristics of the ontology and the metaphor; Determine a corpus construction template group according to the metaphor term corresponding to the missing position in the missing triple; each type of corpus construction template in the corpus construction template group includes at least two metaphor terms of the ontology, metaphor, and attribute; Generate a new corpus with the missing position according to the corpus construction template in the corpus construction template group and the missing triple; Predict the word probability of the missing position corresponding to the corpus construction template through the new corpus and the trained language representation model; Determine the filling word for the missing position according to the word probability of the missing position corresponding to the corpus construction template; Add the filling word to the missing position in the missing triple to obtain the metaphor triple corresponding to the corpus.
2. The method according to claim 1, characterized in that, the training of the language representation model includes: Obtain a corpus sample set; Mark the adjectives and / or nouns included in the corpus samples in the corpus sample set according to the inter-word dependency relationship to obtain marked corpora; Perform masked language model training and next sentence prediction task training on the pre-trained language representation model according to the marked corpora to obtain a trained language representation model.
3. The method according to claim 1, characterized in that, Each metaphor term in the metaphor triple has a corresponding corpus construction template group, and each corpus construction template group includes multiple different types of corpus construction templates; Determining the corpus construction template group according to the metaphor term corresponding to the missing position in the missing triple includes: Search for the corpus construction template group of the metaphor term corresponding to the missing position in the missing triple from the corpus construction template groups respectively corresponding to each metaphor term.
4. The method according to claim 1 or 3, characterized in that, the method further includes: Determine a metaphor triple set according to the metaphor corpus samples in the metaphor corpus sample set; Perform masking processing on the metaphor terms of the metaphor triples in the metaphor triple set in a masking manner to obtain the missing triple of the metaphor triple; Select multiple templates including the missing metaphor terms in the missing triple from multiple different types of corpus construction templates, and construct multiple missing corpora according to the missing triple and the selected multiple templates; Predict the prediction result corresponding to each missing corpus through the language representation model; wherein, the prediction result is the word list probability distribution information corresponding to the missing metaphor term; Determine the optimal template combination corresponding to the missing metaphor term according to the prediction results corresponding to the selected multiple templates, and determine the optimal template combination as the corpus construction template group corresponding to the missing metaphor term.
5. The method according to claim 1 or 3, characterized in that, Multiple corpus construction templates of different types include: the first type of template containing three metaphorical terms, namely the noumenon, the vehicle, and the attribute; the second type of template containing two metaphorical terms, namely the vehicle and the attribute; the third type of template containing two metaphorical terms, namely the noumenon and the attribute; and the fourth type of template containing two metaphorical terms, namely the noumenon and the vehicle.
6. The method according to claim 1, wherein, determining the filler word for the vacancy according to the word probability of the vacancy corresponding to the corpus construction template includes: performing weighted average calculation on the word probability of the vacancy corresponding to each corpus construction template; determining the word with the largest weighted average word probability as the filler word for the vacancy.
7. A metaphor information processing method, wherein, the method includes: monitoring the core information in the current application scenario; wherein, the application scenario includes robot question answering or literary creation, and the core information is the tendency information representing the next application requirement in the application scenario; determining a target metaphor triple from multiple metaphor triples, wherein the multiple metaphor triples are obtained by the corpus processing method according to any one of claims 1-6 and include the noumenon, the vehicle, and the attribute; pushing the prompt information corresponding to the target metaphor triple; wherein, the prompt information includes the target metaphor triple and / or the corpus set corresponding to the target metaphor triple.
8. The method according to claim 7, wherein, determining a target metaphor triple from multiple metaphor triples according to the core information includes: determining the metaphor triple containing the core information as the target metaphor triple; and / or, determining the correlative word of the core information, and determining the metaphor triple containing the correlative word as the target metaphor triple.
9. A corpus processing device, wherein, the device includes: A vacancy triple acquisition module for obtaining the vacancy triple corresponding to the corpus in the target corpus dataset; wherein, the vacancy triple includes at least one vacancy, as well as the adjectives and / or nouns in the corpus; the vacancy corresponds to one metaphorical term among the noumenon, the vehicle, or the attribute in the metaphor triple, and the attribute represents the common characteristics of the noumenon and the vehicle; A template group determination module for determining a corpus construction template group according to the metaphorical term corresponding to the vacancy in the vacancy triple; each type of corpus construction template in the corpus construction template group includes at least two metaphorical terms among the noumenon, the vehicle, and the attribute; A corpus generation module for generating a new corpus with the vacancy according to the corpus construction template in the corpus construction template group and the vacancy triple; A prediction module for predicting the word probability of the vacancy corresponding to the corpus construction template through the new corpus and the trained language representation model; A filler word determination module for determining the filler word for the vacancy according to the word probability of the vacancy corresponding to the corpus construction template; A metaphor triple determination module for adding the filler word to the vacancy in the vacancy triple to obtain the metaphor triple corresponding to the corpus.
10. A metaphor information processing device, characterized in that, the device includes: a listening module, configured to listen for core information in the current application scenario; wherein, the application scenario includes robot question answering or literary creation, and the core information is the tendency information representing the next application requirement in the application scenario; a target metaphor triple determination module, configured to determine a target metaphor triple from a plurality of metaphor triples according to the core information, wherein the plurality of metaphor triples are obtained by the corpus processing method according to any one of claims 1-6, and include an ontology, a vehicle, and an attribute; a prompt information pushing module, configured to push prompt information corresponding to the target metaphor triple; wherein, the prompt information includes the target metaphor triple and / or the corpus set corresponding to the target metaphor triple.
11. An electronic device, characterized in that, it includes a processor and a memory, the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the corpus processing method according to any one of claims 1 to 6 or the metaphor information processing method according to any one of claims 7 to 8.
12. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores computer executable instructions, and when the computer executable instructions are called and executed by a processor, the computer executable instructions cause the processor to implement the corpus processing method according to any one of claims 1 to 6 or the metaphor information processing method according to any one of claims 7 to 8.
Citation Information
Patent Citations
Metaphor sentence recognition method, device and apparatus and storage medium
CN111914544A
Statement acquisition method and device
CN112307754A