A method, related device and equipment for corpus processing
By mining the expanded corpus and obtaining target corpus with similar semantics, the problems of insufficient training corpus and poor generalization effect in the existing technology are solved, and sufficient expansion of corpus and quality improvement of model training are achieved.
Patent Information
- Application Number
- CN202110774306.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-07-08
AI Technical Summary
When processing request statements, the existing semantic analytics platform lacks training corpus or overfits the model, resulting in poor generalization effect of the model.
By corpus mining to expand corpus, a large number of candidate corpus are obtained, and target corpus similar to the semantics of the corpus to be expanded are mined from it, and the corpus set is expanded to meet the needs of model training.
It realizes the acquisition of more corpus through corpus mining, enriching the corpus set to be expanded, and improving the quantity and quality of corpus trained by the model, thereby improving the generalization ability of the model.
Smart Images

Figure CN113821593B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a method for corpus processing, related devices and equipment. Background Art
[0002] With the popularization of artificial intelligence, more and more artificial intelligence technologies can bring convenience to people's lives. For example, users input some request statements through a smart assistant, and the smart assistant analyzes and processes these request statements, and transmits the processing results to subsequent services to make corresponding feedback, thus completing an interaction process with the user by voice.
[0003] Currently, when processing request statements, the existing semantic parsing platform usually uses an intent classification model to parse the statements. The intent classification model needs to be trained with a certain amount of training corpus. However, since only a small amount of seed corpus can be provided in the existing training corpus, it is not enough to support the normal training of the model, or the trained model is too overfitted to the training data, resulting in poor generalization effect of the model. Summary of the Invention
[0004] Embodiments of the present application provide a method for corpus processing, related devices and equipment. By mining the corpus of the corpus to be expanded, a large number of candidate corpora are obtained, and target corpora semantically similar to the corpus to be expanded are further mined from the large number of candidate corpora, so that the corpus to be expanded is sufficiently expanded, thereby meeting the requirement of the model training for the quantity of the corpus.
[0005] In view of this, on the one hand, the present application provides a method for corpus processing, including:
[0006] Obtain the corpus to be expanded;
[0007] Obtain K candidate corpora according to the corpus to be expanded, where the semantic similarity between each candidate corpus and the corpus to be expanded is greater than or equal to a similarity threshold, and K is an integer greater than 1;
[0008] Input the multiple candidate corpora and the corpus to be expanded into a semantic recognition model to obtain K semantic recognition results, where each semantic recognition result is a similarity score or a similarity classification. The similarity score represents the degree of semantic similarity between the candidate corpus and the corpus to be expanded, and the similarity classification represents the semantic category to which the candidate corpus and the corpus to be expanded belong;
[0009] If at least one of the K semantic recognition results meets the corpus extraction condition, determine the candidate corpus corresponding to at least one of the semantic recognition results as the target corpus, so as to obtain at least one target corpus belonging to the corpus to be expanded.
[0010] Another aspect of the present application provides an apparatus for corpus processing, including:
[0011] An acquisition unit, configured to acquire a corpus to be augmented;
[0012] The acquisition unit is further configured to acquire K candidate corpora according to the corpus to be augmented, where the semantic similarity between each candidate corpus and the corpus to be augmented is greater than or equal to a similarity threshold, and K is an integer greater than 1;
[0013] A processing unit, configured to input the K candidate corpora and the corpus to be augmented into a semantic recognition model to obtain K semantic recognition results, where each semantic recognition result is a similarity score or a similarity classification, the similarity score represents the degree of semantic similarity between the candidate corpus and the corpus to be augmented, and the similarity classification represents the semantic category to which the candidate corpus and the corpus to be augmented belong;
[0014] A determination unit, configured to, if at least one of the K semantic recognition results satisfies the corpus extraction condition, determine the candidate corpus corresponding to at least one of the semantic recognition results as the target corpus, so as to obtain at least one target corpus belonging to the corpus to be augmented.
[0015] In a possible design, in an implementation manner of another aspect of the embodiments of the present application,
[0016] The acquisition unit is further configured to acquire a corpus sample set, where the corpus sample set includes positive sample corpora and negative sample corpora, and the positive sample corpora correspond to labeled tags;
[0017] The processing unit is further configured to perform feature extraction on the positive sample corpora to obtain positive sample corpus features, and perform feature extraction on the negative sample corpora to obtain negative sample corpus features;
[0018] The processing unit is further configured to input the positive sample corpus features and the negative sample corpus features into the semantic recognition model to obtain a semantic prediction result;
[0019] The processing unit is further configured to train the semantic recognition model according to the semantic prediction result and the labeled tag.
[0020] In a possible design, in an implementation manner of another aspect of the embodiments of the present application, the processing unit may specifically be configured to:
[0021] Perform word segmentation on the positive sample corpora and the negative sample corpora respectively to obtain at least two to-be-processed positive sample words and at least two to-be-processed negative sample words;
[0022] Convert at least two to-be-processed positive sample words and at least two to-be-processed negative sample words into at least two positive sample word vectors and at least two negative sample word vectors;
[0023] Concatenate at least two positive sample word vectors to obtain the positive sample corpus feature, and concatenate at least two negative sample word vectors to obtain the negative sample corpus feature.
[0024] In a possible design, in an implementation manner of another aspect of the embodiments of the present application,
[0025] The processing unit is further configured to, if there is no semantic recognition result among the K semantic recognition results that meets the corpus extraction condition, translate the corpus to be augmented into N language corpora corresponding to N languages.
[0026] The processing unit is further configured to translate the N language corpora into at least N back-translated corpora according to the language of the corpus to be augmented.
[0027] In a possible design, in an implementation manner of another aspect of the embodiments of the present application,
[0028] The determining unit is further configured to determine at least one target corpus and at least N back-translated corpora as a plurality of corpora to be labeled.
[0029] The processing unit is further configured to perform slot matching on each corpus to be labeled to obtain a slot matching result corresponding to each corpus to be labeled, where the slot matching result is a matching similarity score, and the matching similarity score represents the semantic similarity degree between a preset slot and each to-be-labeled word in each corpus to be labeled.
[0030] The determining unit is further configured to, if the matching similarity score is greater than or equal to a preset matching threshold, determine the preset slot corresponding to the matching similarity score as the target slot, and determine the to-be-labeled word corresponding to the matching similarity score as the target slot value, where the target slot is used to represent the attribute of the target slot value.
[0031] The processing unit is further configured to perform deduplication processing on the corpus to be labeled, the target slot corresponding to the corpus to be labeled, and the target slot value to obtain the target labeled corpus.
[0032] In a possible design, in an implementation manner of another aspect of the embodiments of the present application,
[0033] The processing unit is further configured to, if there is no semantic recognition result among the K semantic recognition results that meets the corpus extraction condition, classify the K candidate corpora to obtain a first candidate corpus set and a second candidate corpus set, where the first candidate corpus set includes i first candidate corpora, the second candidate corpus set includes j second candidate corpora, and both i and j are integers greater than 1 and less than K.
[0034] The processing unit is further configured to pairwisely combine the i first candidate corpora with the j second candidate corpora respectively to obtain i*j target corpus pairs.
[0035] In a possible design, in an implementation manner of another aspect of the embodiments of the present application, the processing unit may specifically be configured to:
[0036] If each semantic recognition result is a similarity score, determine the candidate corpus corresponding to the similarity score greater than or equal to a preset similarity threshold as the target corpus;
[0037] If each semantic recognition result is a similarity classification, determine the candidate corpus corresponding to the similarity classification greater than or equal to a preset classification probability threshold as the target corpus.
[0038] Another aspect of the present application provides a computer device, including: a memory, a transceiver, a processor, and a bus system;
[0039] Wherein, the memory is used to store programs;
[0040] The processor is configured to implement the methods in the above aspects when executing the programs in the memory;
[0041] The bus system is used to connect the memory and the processor to enable the memory and the processor to communicate with each other.
[0042] Another aspect of the present application provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a computer, the computer is enabled to execute the methods in the above aspects.
[0043] Another aspect of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the network device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the network device executes the methods provided in the above aspects.
[0044] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0045] By obtaining the corpus to be expanded, K candidate corpora are obtained according to the corpus to be expanded, and the multiple candidate corpora and the corpus to be expanded are input into the semantic recognition model to obtain K semantic recognition results. Then, when at least one of the K semantic recognition results meets the corpus extraction condition, the candidate corpus corresponding to at least one semantic recognition result is determined as the target corpus, so as to obtain at least one target corpus belonging to the corpus to be expanded. Through the above method, it is realized to obtain a large number of candidate corpora by mining the corpus to be expanded, and further mine the target corpora with similar semantics to the corpus to be expanded from the large number of candidate corpora, which can realize expanding a larger number of target corpora based on the corpus to be expanded, enriching the corpus to be expanded, so that the corpus set can be sufficiently expanded to meet the requirement of the number of corpora for model training. Brief Description of the Drawings
[0046] Figure 1 is a schematic diagram of the architecture of the corpus control system in an embodiment of the present application;
[0047] Figure 2 is a schematic diagram of an embodiment of the method for corpus processing in an embodiment of the present application;
[0048] Figure 3 is another schematic diagram of an embodiment of the method for corpus processing in an embodiment of the present application;
[0049] Figure 4 is another schematic diagram of an embodiment of the method for corpus processing in an embodiment of the present application;
[0050] Figure 5 is another schematic diagram of an embodiment of the method for corpus processing in an embodiment of the present application;
[0051] Figure 6 is another schematic diagram of an embodiment of the method for corpus processing in an embodiment of the present application;
[0052] Figure 7 is another schematic diagram of an embodiment of the method for corpus processing in an embodiment of the present application;
[0053] Figure 8 is another schematic diagram of an embodiment of the method for corpus processing in an embodiment of the present application;
[0054] Figure 9 is a schematic diagram of the principle process of the method for corpus processing in an embodiment of the present application;
[0055] Figure 10 is a schematic diagram of the principle of model training of the method for corpus processing in an embodiment of the present application;
[0056] Figure 11It is a schematic diagram of a corpus annotation interface for the corpus processing method in an embodiment of the present application;
[0057] Figure 12 It is a schematic diagram of the principle of corpus parsing for the corpus processing method in an embodiment of the present application;
[0058] Figure 13 It is a schematic diagram of an embodiment of the corpus processing apparatus in an embodiment of the present application;
[0059] Figure 14 It is a schematic diagram of an embodiment of a computer device in an embodiment of the present application. Detailed implementation manners
[0060] The embodiment of the present application provides a corpus processing method, which is used for corpus mining of the corpus to be expanded to obtain a large number of candidate corpora, and further mining target corpora with semantics similar to the corpus to be expanded from the large number of candidate corpora, so that the corpus to be expanded can be sufficiently expanded to meet the requirement of the model training for the quantity of the corpus.
[0061] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present invention are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0062] It should be understood that the corpus processing method provided by the present application can be applied to the scenario of intelligent voice interaction by parsing statements. As an example, for example, a smart speaker parses the received request statement to play or pause music. As another example, for example, a smart TV parses the received request statement to turn on or switch TV channels, etc. As still another example, for example, a smart watch parses the received request statement to set an alarm. In the above various scenarios, in order to complete the parsing of the statement, the solution provided in the prior art is to parse the statement through an intent classification model, and the intent classification model needs to be trained with a certain amount of training corpus. However, since only a small amount of seed corpus can be provided in the existing training corpus, it is not enough to support the normal training of the model, or the trained model is too overfitted to the training data, resulting in too poor generalization effect of the model.
[0063] To solve the above problems, the present application proposes a method for corpus processing, which is applied to Figure 1 the corpus control system shown in, please refer to Figure 1 , Figure 1 As shown in, it is a schematic architecture diagram of the corpus control system in an embodiment of the present application. As shown in the figure, the server obtains the corpus to be expanded provided by the client, obtains K candidate corpora according to the corpus to be expanded, and inputs the multiple candidate corpora and the corpus to be expanded into the semantic recognition model to obtain K semantic recognition results. Then, when at least one of the K semantic recognition results meets the corpus extraction condition, the candidate corpus corresponding to at least one semantic recognition result is determined as the target corpus, so as to obtain at least one target corpus belonging to the corpus to be expanded. Through the above method, it is realized to obtain a large number of candidate corpora through corpus mining of the corpus to be expanded, and further mine the target corpora semantically similar to the corpus to be expanded from the large number of candidate corpora, which can realize the expansion of a larger number of target corpora based on the corpus to be expanded, so as to sufficiently expand the corpus set and meet the requirement of the model training for the quantity of the corpus.
[0064] It can be understood that the client and the server are communicatively connected. Figure 1 One server is shown in, but in the actual scenario, multiple servers may also participate. Especially in the scenario of multi-model training interaction, the number of servers depends on the actual scenario and is not specifically limited here. In addition, Figure 1 the number of clients shown in is only an example and is not used to limit the number of clients. The number should be flexibly determined according to the actual situation.
[0065] It should be noted that in this embodiment, the server or the trading server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The client can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto.
[0066] To solve the above problems, the present application proposes a method for corpus processing, which is generally executed by a server or a terminal device. Correspondingly, the device for corpus processing is generally arranged in the server or the terminal device.
[0067] It can be understood that, for the corpus processing methods, related devices, and apparatuses disclosed in this application, multiple servers / terminal devices can form a blockchain, and the servers / terminal devices are nodes on the blockchain. In practical applications, data sharing between nodes may be required in the blockchain, and a corpus set and a corpus to be expanded can be stored on each node.
[0068] Next, the corpus processing method in this application will be introduced. Please refer to Figure 2 , an embodiment of the corpus processing method in the embodiments of this application includes:
[0069] In step S101, obtain the corpus to be expanded;
[0070] In this embodiment, the corpus to be expanded generally refers to skill corpus. As Figure 11 shown, when creating the skills of a dialogue system, the creator needs to provide a certain amount of training corpus. However, in actual use, especially when third-party creators build their own skills, usually only a very small number of skill corpus can be provided, such as "Is today suitable for moving to a new home?", "What things are suitable to do today?", or "What things can be done today?", etc. To avoid overfitting the corpus recognition model based on a very small number of corpus, resulting in poor usage effects of the skills created based on the corpus recognition model, in this embodiment, each skill corpus in a small amount of skill corpus corresponding to each skill is obtained from the database as the corpus to be expanded, so that the obtained corpus to be expanded can be mined subsequently to enrich the skill corpus, thereby meeting the requirement of the model training for the quantity of the corpus to a certain extent.
[0071] Among them, a skill refers to the abstraction of specific capabilities in a task-based dialogue system, which can specifically be manifested as a music skill. This music skill can be used to represent that the dialogue system can understand short texts (queries) related to music, perform operations such as domain landing and parameter extraction on the short texts, and then express the key information in the short texts in structured information, and then transmit it to subsequent services for the dialogue system to make corresponding feedback, thereby realizing an interaction process between the dialogue system and the user's voice. It can also be manifested as other skills, such as story skills, telephone skills, or time skills, etc., which are not specifically limited here.
[0072] Among them, the short text can specifically be manifested as a request statement input by the user, which is usually used to represent an intention expectation of the user. For example, "Play Zhang San's It's Raining", "Tell me the story of The Foolish Old Man Removes the Mountains", or "I want to watch the movie Infernal Affairs", etc.
[0073] Specifically, as Figure 9As shown in the figure, in this embodiment, each piece of skill corpus in the small amount of skill corpora corresponding to each skill is obtained from the database as the corpus to be expanded, so that the corpus to be expanded can be mined through natural language processing (NLP) subsequently to enrich the skill corpus.
[0074] Among them, natural language processing is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use daily, so it has some close connections with linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, pointing maps, and other technologies.
[0075] In step S102, K candidate corpora are obtained according to the corpus to be expanded, where the semantic similarity between each candidate corpus and the corpus to be expanded is greater than or equal to the similarity threshold, and K is an integer greater than 1;
[0076] In this embodiment, after obtaining the corpus to be expanded, as Figure 9 shown, some candidate corpora can be retrieved from the index library of the distributed full-text search engine (ElasticSearch, ES). Specifically, the semantic similarity between the corpus to be expanded and the corpora in the index library can be calculated, and then the candidate corpora with higher similarity to the corpus to be expanded can be selected by comparing the semantic similarity with the preset similarity threshold. For example, the corpus corresponding to the semantic similarity greater than or equal to the similarity threshold is determined as the candidate corpus, where the similarity threshold can be specifically set according to actual application requirements and is not specifically limited here.
[0077] Among them, the candidate corpus can be understood as the relevant corpus obtained from the index library from the perspective of literal hit. These candidate corpora may be literally similar to the corpus to be expanded, but these candidate corpora are not necessarily semantically similar to the corpus to be expanded. It can be understood that the candidate corpus contains the target corpus with a relatively similar semantic expression to the corpus to be expanded, and also contains a large number of corpora that are semantically related to the corpus to be expanded but may not be semantically similar.
[0078] Among them, the calculation method for calculating the semantic similarity between the corpus to be expanded and the corpora in the index library can specifically adopt Hamming distance, cosine similarity, Pearson correlation coefficient, etc., or other calculation methods can also be used for calculation, such as Euclidean distance, Manhattan distance, etc., and it is not specifically limited here.
[0079] Among them, the semantic similarity between the corpus to be augmented and the corpus in the index library is calculated based on cosine similarity. Specifically, the feature words in the corpus to be augmented are extracted, and then the extracted feature words are converted into text vectors. Similarly, the text vectors of each corpus in the index library can be obtained. Then, by substituting the text vectors of the corpus to be augmented and each corpus in the index library into the cosine similarity calculation formula respectively, the semantic similarity between the corpus to be augmented and the corpus in the index library can be obtained.
[0080] Specifically, after obtaining the corpus to be augmented, through the retrieval service provided by ES, the corpus with a semantic similarity greater than or equal to the similarity threshold between a large number of corpora to be augmented can be retrieved in the index library. Then, these retrieved corpora are determined as candidate corpora, and the corpus to be augmented can be expanded into a large number of candidate corpora, which can enrich the corpus set to a certain extent and thus meet the requirement of the model training for the quantity of the corpus to a certain extent.
[0081] For example, assume that a corpus to be augmented is "What things are suitable to do today". Through ES, the candidate corpora retrieved in the index library can be such as "Things that can be done today", "What popular things are there today", and "What things happened today", etc.
[0082] In step S103, the K candidate corpora and the corpus to be augmented are input into the semantic recognition model to obtain K semantic recognition results. Among them, each semantic recognition result is a similarity score or a similarity classification. The similarity score represents the degree of semantic similarity between the candidate corpus and the corpus to be augmented, and the similarity classification represents the semantic category to which the candidate corpus and the corpus to be augmented belong.
[0083] In this embodiment, since the candidate corpora include the target corpora with relatively similar semantic expressions to the corpus to be augmented, and also include a large number of corpora that are semantically related to the corpus to be augmented but may not be semantically similar, after obtaining multiple candidate corpora, this embodiment can input the K candidate corpora and the corpus to be augmented into the semantic recognition model to obtain K semantic recognition results, so that in the follow-up, multiple corpora with semantic similarity to the semantic expression of the corpus to be augmented can be accurately obtained according to the K semantic recognition results. It can make the augmentation of the corpus to be augmented meaningful and at the same time enrich the corpus to be augmented, thus meeting the requirement of the model training for the quantity of the corpus to a certain extent.
[0084] Among them, the semantic recognition model can specifically be a Logistic Regression (LR) model, a Linear Regression (LR) model, or other semantic recognition models, such as a recurrent neural network model, a feedforward neural network model, etc. There is no specific limitation here.
[0085] Specifically, as Figure 9 shown, after obtaining K candidate corpora, the K candidate corpora and the corpus to be expanded can be input into a semantic recognition model. For example, when input into a logistic regression model, the category probabilities of the semantic categories to which the candidate corpora and the corpus to be expanded belong can be obtained, that is, similarity classification. Or when input into a linear regression model, the semantic similarity degree between the candidate corpora and the corpus to be expanded can be scored, that is, similarity score, so that subsequently, the K candidate corpora obtained can be filtered according to the similarity score or similarity classification to accurately obtain multiple corpora with semantics similar to those expressed by the corpus to be expanded, enrich the corpus to be expanded, and thus meet the requirement of the model training for the quantity of corpora to a certain extent.
[0086] For example, assume that there is a corpus to be expanded as "What songs are suitable to listen to on rainy days", and assume that there are 3 candidate corpora as "Music related to rainy days", "Listening to songs on rainy days", and "What happened on rainy days", etc. Inputting these candidate corpora and the corpus to be expanded into the logistic regression model, the semantic recognition results including the category probabilities of the semantic categories to which the corpus to be expanded and these candidate corpora belong are respectively 0.91, 0.74, and 0.48.
[0087] In step S104, if at least one of the K semantic recognition results meets the corpus extraction condition, the candidate corpus corresponding to at least one of the K semantic recognition results is determined as the target corpus to obtain at least one target corpus belonging to the corpus to be expanded.
[0088] In this embodiment, after obtaining the K semantic recognition results, the semantic recognition results that meet the corpus extraction condition can be screened out from the K semantic recognition results, that is, at least one of the K semantic recognition results meets the corpus extraction condition. Furthermore, the candidate corpus corresponding to the semantic recognition result that meets the corpus extraction condition can be used as the candidate corpus with a relatively high semantic similarity to the corpus to be expanded screened out from the K candidate corpora, that is, the target corpus, which can make the expansion of the corpus to be expanded meaningful and at the same time enrich the corpus to be expanded, thus meeting the requirement of the model training for the quantity of corpora to a certain extent.
[0089] Among them, screening out semantic recognition results that meet the corpus extraction conditions from K semantic recognition results can be specifically manifested as K similarity classifications, that is, among the K category probabilities representing the semantic categories to which the candidate corpus and the corpus to be augmented belong, the category probabilities greater than or equal to the preset probability threshold can be determined as the semantic recognition results that meet the corpus extraction conditions. Conversely, the category probabilities less than the preset probability threshold are the semantic recognition results that do not meet the corpus extraction conditions. Among them, the probability threshold can be set according to actual application requirements, such as 0.7, and no specific limitation is made here. Or the category probability approaching 1 can be determined as the semantic recognition result that meets the corpus extraction conditions. Conversely, the category probability approaching 0 is the semantic recognition result that does not meet the corpus extraction conditions. There can also be other forms, and no specific limitation is made here.
[0090] Specifically, as Figure 9 shown, after obtaining K semantic recognition results, the K candidate corpora can be filtered according to the K semantic recognition results to screen out the semantic recognition results that meet the corpus extraction conditions from the K semantic recognition results. For example, multiple semantic recognition results with higher probability values can be selected from the K semantic recognition results, and then the candidate corpora corresponding to the selected multiple semantic recognition results with high probability values can be determined as multiple target corpora that can represent the semantics similar to the corpus to be augmented, which can enrich the corpus to be augmented and thus meet the requirement of the model training for the quantity of the corpus to a certain extent.
[0091] For example, assume that the category probabilities of the semantic categories to which a corpus to be augmented, "What songs are suitable to listen to on a rainy day", belongs to three candidate corpora, "Music related to rainy days", "Listening to music on a rainy day", and "What happened on a rainy day", are 0.91, 0.74, and 0.48 respectively. Assume that a preset probability threshold is 0.7. Then, by comparing these classification probabilities with the preset probability threshold, it can be clearly obtained that the category probabilities 0.91 and 0.74 are greater than the probability threshold of 0.7. Then, the candidate corpora, "Music related to rainy days" and "Listening to music on a rainy day", can be determined as the target corpora of the corpus to be augmented, "What songs are suitable to listen to on a rainy day".
[0092] In the embodiment of the present application, a method for corpus processing is provided. Through the above method, a large number of candidate corpora are obtained by mining the corpus to be augmented, and further target corpora semantically similar to the corpus to be augmented are mined from the large number of candidate corpora, which can realize the expansion of a larger number of target corpora based on the corpus to be augmented, so as to sufficiently expand the corpus set and thus meet the requirement of the model training for the quantity of the corpus.
[0093] Optionally, in the above Figure 2Based on the corresponding embodiments, in another alternative embodiment of the corpus processing method provided in the embodiments of the present application, as Figure 3 shown, the method further includes:
[0094] In step S301, a corpus sample set is obtained, where the corpus sample set includes positive sample corpora and negative sample corpora, and the positive sample corpora correspond to labeled tags;
[0095] In step S302, feature extraction is performed on the positive sample corpora to obtain positive sample corpus features, and feature extraction is performed on the negative sample corpora to obtain negative sample corpus features;
[0096] In step S303, the positive sample corpus features and the negative sample corpus features are input into a semantic recognition model to obtain a semantic prediction result;
[0097] In step S304, the semantic recognition model is trained according to the semantic prediction result and the labeled tag.
[0098] In this embodiment, after obtaining K candidate corpora, in order to accurately obtain multiple targets with semantics similar to those expressed by the corpus to be expanded from the K candidate corpora, this embodiment can obtain a corpus sample set, perform feature extraction on the corpus sample set, and then use the extracted features and the labeled tags to train the semantic recognition model to improve the learning ability of the semantic recognition model, so that when the K candidate corpora and the corpus to be expanded are input into the trained and optimized semantic recognition model subsequently, a more accurate semantic recognition result can be obtained.
[0099] Among them, the corpus sample set is a test sample for testing the semantic recognition model, including positive sample corpora and negative sample corpora. Among them, the positive sample corpora are seed corpora with labeled tags, that is, skill corpora, and the labeled tags are tags that can be used to represent the semantic attributes in the corpora. The negative sample corpora are common negative sample corpora, which can be understood as corpora that are not related or semantically dissimilar to the positive sample corpora. Further, it can be understood that in the actual use process, since the skill corpora that users can provide are less, the positive sample corpora are usually fewer than the negative sample corpus set.
[0100] Among them, as Figure 9As shown in the figure, for the feature extraction of positive sample corpus and negative sample corpus, specifically, the Natural Language Understanding (NLU) service can be used to extract features from the positive sample corpus or negative sample corpus respectively to obtain positive sample corpus features or negative sample corpus features. Among them, natural language understanding technology includes technologies in multiple fields such as sentence detection, word segmentation, part-of-speech tagging, syntactic analysis, text classification / clustering, character perspective, information extraction / auto summarization, machine translation, automatic question answering, and text generation.
[0101] Among them, the processing method of inputting the positive sample corpus features and negative sample corpus features into the semantic recognition model to obtain the semantic prediction result is similar to the processing method in step S103 of inputting K candidate corpora and the corpus to be expanded into the semantic recognition model to obtain K semantic recognition results, which will not be elaborated here.
[0102] Among them, according to the semantic prediction result and the annotation label, the semantic recognition model is trained. Specifically, it can be based on the cross-entropy loss function, and the semantic recognition model is trained by backpropagation iteration using the semantic prediction result and the annotation label.
[0103] It should be noted that when extracting features from the positive sample corpus and negative sample corpus, when considering complex features, there may be a situation where the feature extraction takes a long time, resulting in a long training process of the model, and thus the entire target corpus parsing process will be relatively complex and time-consuming. Therefore, in this embodiment, the processes such as feature extraction of the corpus sample set and training the semantic recognition model can be executed offline to reduce the pressure caused by online execution.
[0104] Specifically, as Figure 10 shown, the corpus sample set for training the semantic recognition model can be obtained from the database. Furthermore, the NLU technology can be used to extract features from the positive sample corpus and negative sample corpus in the corpus sample set respectively to accurately and fully obtain the positive sample corpus features and negative sample corpus features that can be used as model inputs. Furthermore, the positive sample corpus features and negative sample corpus features can be input into the semantic recognition model to obtain the semantic prediction result. Then, based on the cross-entropy loss function, and using the semantic prediction result and the annotation label to perform backpropagation iteration training on the semantic recognition model, it is possible to adjust the values of various parameters according to the error during the backpropagation process, and continuously iterate the above process until convergence to achieve the optimization of the model parameters.
[0105] Optionally, on the basis of the above Figure 2 corresponding embodiment, in another optional embodiment of the corpus processing method provided by the embodiment of the present application, as Figure 4As shown, feature extraction is performed on the positive sample corpus to obtain positive sample corpus features, and feature extraction is performed on the negative sample corpus to obtain negative sample corpus features, including:
[0106] In step S401, word segmentation is respectively performed on the positive sample corpus and the negative sample corpus to obtain at least two positive sample words to be processed and at least two negative sample words to be processed;
[0107] In step S402, at least two positive sample words to be processed and at least two negative sample words to be processed are converted into at least two positive sample word vectors and at least two negative sample word vectors;
[0108] In step S403, at least two positive sample word vectors are vector-concatenated to obtain positive sample corpus features, and at least two negative sample word vectors are vector-concatenated to obtain negative sample corpus features.
[0109] In this embodiment, after obtaining the positive sample corpus and the negative sample corpus, preprocessing can be first performed on the positive sample corpus and the negative sample corpus. Specifically, punctuation removal processing and word segmentation processing can be respectively performed on the positive sample corpus and the negative sample corpus. Since the purpose of word segmentation is to divide a continuous sentence into individual word units, the understanding of the corpus is thus transformed into the processing of words, thereby improving the efficiency of corpus processing. Among them, when performing word segmentation processing on the positive sample corpus and the negative sample corpus respectively, specifically, it can be performed based on a dictionary method, a statistical method, or a rule-based method, or other word segmentation algorithms can also be used, and no specific limitation is made here. This embodiment can use a general binary word model to segment sentences to obtain at least two positive sample words to be processed and at least two negative sample words to be processed.
[0110] Furthermore, as Figure 10As shown, in order for the computer to better understand the features in the positive sample corpus and the negative sample corpus, so that the subsequent model can be better trained based on the extracted features, in this embodiment, at least two positive sample words to be processed and at least two negative sample words to be processed are converted into digital features that can be used for machine learning, that is, at least two positive sample word vectors and at least two negative sample word vectors. Specifically, at least two positive sample words to be processed and at least two negative sample words to be processed can be respectively input into a variety of word vector extraction models to obtain word vectors of various different dimensions. For example, when input into the Bidirectional Encoder Representations from Transformers (BERT) model, words can be converted into vectors that can represent basic features and can consider the context features of the entire sentence to a certain extent. And when input into the fasttext model, entities and some specific attributes in the words can be represented by word vectors, and the semantic features of the entire sentence can be considered to a certain extent. It can also be other word vector extraction models, such as the ELMo model, the word2vec model, or the glove model, etc. There is no specific limitation here. In this way, various different dimensional vector representations of the same word can be obtained, that is, various different dimensional vector representations of the positive sample words to be processed and the negative sample words to be processed can be obtained to represent various different dimensional features of the positive sample corpus and the negative sample corpus.
[0111] Further, after converting the positive sample words to be processed and the negative sample words to be processed into various different word vector representations by using a variety of different types of word vector extraction models, at least two positive sample word vectors can be vector concatenated to obtain the positive sample corpus features, and at least two negative sample word vectors can be vector concatenated to obtain the negative sample corpus features. The feature dimension can be increased through the way of vector concatenation, so that the model can fully learn various different dimensional features to achieve a more accurate prediction effect.
[0112] Specifically, after obtaining the positive sample corpus and the negative sample corpus, in order to enable the semantic recognition model to more accurately and quickly recognize and process the positive sample corpus and the negative sample corpus, through word segmentation processing, the positive sample corpus and the negative sample corpus can be converted into word processing, which can improve the efficiency of corpus processing. Furthermore, by inputting at least two to-be-processed positive sample words and at least two to-be-processed negative sample words obtained through processing into the BERT model and the fasttext model respectively, various vector representations of different dimensions of the same to-be-processed positive sample word and various vector representations of different dimensions of the same to-be-processed negative sample word can be obtained. Then, by concatenating the vector representations of different dimensions of all the to-be-processed positive sample words and concatenating the vector representations of different dimensions of all the to-be-processed negative sample words, positive sample corpus features and negative sample corpus features with increased feature dimensions can be obtained, enabling the model to more fully and quickly learn various features of different dimensions, thus achieving a more accurate prediction effect.
[0113] Optionally, based on the corresponding embodiment above Figure 2 In another optional embodiment of the corpus processing method provided by the embodiment of the present application, as Figure 5 shown, the method further includes:
[0114] In step S501, if there is no semantic recognition result among the K semantic recognition results that meets the corpus extraction condition, the to-be-expanded corpus is translated into N language corpora corresponding to N languages;
[0115] In step S502, according to the language of the to-be-expanded corpus, the N language corpora are translated into at least N back-translated corpora.
[0116] In this embodiment, when there is no semantic recognition result among the K semantic recognition results that meets the corpus extraction condition, this embodiment can obtain a large number of back-translated corpora similar in semantics to the to-be-expanded corpus through corpus back-translation processing to enhance the to-be-expanded corpus and enrich the to-be-expanded corpus, thereby meeting the requirement of the model training for the quantity of the corpus to a certain extent.
[0117] Specifically, as Figure 9As shown, after obtaining the corpus to be expanded, this embodiment can translate the corpus to be expanded into N language corpora according to the preset N languages, which can be understood as converting the corpus to be expanded into N language representations. For example, a Chinese corpus to be expanded is converted into N language representations, such as English, German, Spanish, French, Russian, Korean and Japanese. Then, according to the language of the corpus to be expanded, the N language corpora are translated into at least N back-translated corpora, which can be understood as converting the N language representations into one language representation respectively. For example, English corpus is translated into Chinese, and one or more Chinese corpora similar to or consistent with the corpus to be expanded are obtained, that is, N language corpora can be translated to obtain at least N back-translated corpora similar to the corpus to be expanded, which can enrich the corpus to be expanded, thereby meeting the demand for the number of corpora for model training to a certain extent.
[0118] For example, suppose there is a Chinese corpus to be expanded, "美好生活" is translated into an English corpus, namely "Beautiful Life". Then, by translating the English corpus "Beautiful Life" into Chinese, at least three back-translated corpora with similar semantics to the corpus to be expanded can be obtained, such as "美丽人生", "美丽人生" or "美好生活".
[0119] It should be noted that this embodiment may also adopt a synonym replacement method to obtain a large amount of synonymous data that is semantically similar to the corpus to be expanded, thereby enhancing the corpus to be expanded and enriching the corpus to be expanded.
[0120] Optionally, in the above Figure 2 Based on the corresponding embodiment, in another optional embodiment of the method for corpus processing provided in the embodiment of the present application, as Figure 6 As shown, the method also includes:
[0121] In step S601, at least one target corpus and at least N back-translated corpora are determined as a plurality of corpora to be annotated;
[0122] In step S602, slot matching is performed on each corpus to be annotated to obtain a slot matching result corresponding to each corpus to be annotated, wherein the slot matching result is a matching similarity score, and the matching similarity score indicates the semantic similarity between the preset slot and each word to be annotated in each corpus to be annotated;
[0123] In step S603, if the matching similarity score is greater than or equal to the preset matching threshold, the preset slot corresponding to the matching similarity score is determined as the target slot, and the to-be-annotated word corresponding to the matching similarity score is determined as the target slot value, wherein the target slot is used to represent the attribute of the target slot value;
[0124] In step S604, the to-be-annotated corpus, the target slot corresponding to the to-be-annotated corpus, and the target slot value are de-duplicated to obtain the target annotated corpus.
[0125] In this embodiment, after obtaining at least one target corpus and at least N back-translated corpora, slot matching can be performed on the at least one target corpus and the at least N back-translated corpora. It can be understood that semantic information annotation is performed on the at least one target corpus and the at least N back-translated corpora. Similar to label annotation, it can convert unstructured corpora into structured ones. Through slot annotation, it can help the model better learn during training or prediction, thereby improving the prediction effect of the model, and can reduce the workload of manual annotation and lower the labor cost, thus improving the efficiency of corpus processing to a certain extent. Among them, a slot refers to an explicitly defined attribute of an entity, which is usually used in a task-based dialogue system to represent the slot design under a specific intention and can be used to express important information in a query. For example, when a skill corpus of a music skill is "I want to listen to Zhang San's It's Raining", the entity "Zhang San" can be represented by the slot "singer", such as "singer = Zhang San", and "It's Raining" can be represented by the slot "song", such as "song = It's Raining".
[0126] Specifically, as Figure 9 shown, after obtaining the target corpus or the back-translated corpus, in order to enable the corpus recognition model to better and more accurately recognize or parse the corpus, this embodiment can first use the at least one target corpus and the at least N back-translated corpora as multiple to-be-annotated corpora. Furthermore, the similarity score between each word in each to-be-annotated corpus and each preset slot in the preset slot library can be calculated to obtain the matching similarity score between each word and the preset slot. Among them, as Figure 11 shown, the preset slot is configured according to the entity library and is used to represent the explicitly defined attributes of the entity. The preset slot can specifically be expressed as "time", "zodiac sign", etc., and can also be other preset slots, such as "departure point", "destination", or "singer", etc. There is no specific limitation here. Then, the preset slot corresponding to the matching similarity score greater than or equal to the preset matching threshold and the word can be understood as being adaptable, and then the preset slot can be determined as the target slot, and the word can be determined as the slot value. The attribute of the slot value can be clearly represented through the slot, and the semantic recognition model can be helped to better learn through the labeled slot and slot value, thereby improving the recognition performance of the semantic recognition model. On the contrary, if the matching similarity score between each word and the preset slot is less than the preset matching threshold, it can be understood that the word is not an entity word or does not have a marking meaning, or other slot information needs to be filled in later.
[0127] Furthermore, the to-be-annotated corpus, the target slots corresponding to the to-be-annotated corpus, and the target slot values can be de-duplicated and sorted to obtain a neat and orderly target annotated corpus, which can not only make the target annotated corpus meaningful, but also avoid duplication and redundancy of the target annotated corpus and waste of resources.
[0128] For example, assuming a target corpus "Query flights from Guangzhou to Beijing on October 4th", by performing slot matching on the to-be-annotated words "Query / October 4th / from / Guangzhou / to / Beijing / of / flights", the corresponding target slots and target slot values can be obtained, such as the target slot "departure time" and the target slot value "October 4th", the target slot "departure place" and the target slot value "Guangzhou", the target slot "destination" and the target slot value "Beijing", etc.
[0129] It should be noted that as Figure 12 shown, after obtaining the target annotated corpus, the skill corpus required for creating skills can be fully expanded. Furthermore, the corpus sample set added with the target annotated corpus can be used to iteratively train the corpus recognition model, and then it can be in the online state. The trained model can be used to parse newly obtained queries, which can improve the parsing ability of new skills.
[0130] For example, as shown in Table 1, when using Method A, such as the template matching method, to parse the test sets of 5 skills, such as "calendar retrieval", it can be observed that for Method A, its accuracy P is very high, but the recall rate R is very low, resulting in an unsatisfactory final F1 value; while for Method B, that is, by adding the to-be-expanded corpus and the target annotated corpus corresponding to the to-be-expanded corpus to the corpus sample set, when the trained corpus recognition model is used to test the test sets of 5 skills, its recall rate is greatly improved, making the overall F1 value increase significantly.
[0131] Table 1
[0132]
[0133] Optionally, based on the above Figure 2 corresponding embodiment, in another optional embodiment of the corpus processing method provided by the embodiments of the present application, as Figure 7 shown, the method further includes:
[0134] In step S701, if there is no semantic recognition result among the K semantic recognition results that meets the corpus extraction condition, the K candidate corpora are classified to obtain a first candidate corpus set and a second candidate corpus set, where the first candidate corpus set includes i first candidate corpora, the second candidate corpus set includes j second candidate corpora, and both i and j are integers greater than 1 and less than K;
[0135] In step S702, the i first candidate corpora are respectively combined pairwise with the j second candidate corpora to obtain i*j target corpus pairs.
[0136] In this embodiment, when there is no semantic recognition result among the K semantic recognition results that meets the corpus extraction condition, this embodiment can, through the processing method of corpus clustering, first divide the K candidate corpora into a first candidate corpus set and a second candidate corpus set, and then combine each corpus in the first candidate corpus set and the second candidate corpus set pairwise, so as to obtain a plurality of target corpus pairs similar to the candidate corpora, that is, target corpus pairs similar to the corpus to be expanded, to enrich the corpus to be expanded, thereby meeting the requirement of the model training for the quantity of the corpus to a certain extent.
[0137] Among them, the first candidate corpus set can be represented as a corpus set that is relatively similar to the corpus to be expanded, that is, the i first candidate corpora included in the first candidate corpus set are corpora that are relatively similar to the corpus to be expanded, such as the same number of words in the corpus or the number of the same characters in the corpus exceeding half of the total number of words in the corpus, etc. The second candidate corpus set can be represented as a corpus set that is not very similar to the corpus to be expanded, that is, the j second candidate corpora included in the second candidate corpus set are corpora that are not very similar to the corpus to be expanded, such as different numbers of words in the corpus or very few same characters in the corpus.
[0138] Specifically, after obtaining the K candidate corpora similar to the corpus to be expanded, the K candidate corpora can be classified. Specifically, it can be classified according to the number of the same characters in the corpus or the number of the same words in the corpus, or other classification methods can also be used, such as the number of words in the corpus, or a combination of several classification methods. There is no specific limitation here. A first candidate corpus set that is relatively similar to the corpus to be expanded and a second candidate corpus set that is not very similar to the corpus to be expanded can be obtained. Then, by combining the i first candidate corpora in the first candidate corpus set and the j first candidate corpora in the second candidate corpus set pairwise respectively, i*j target corpus pairs greater than K can be obtained, that is, target corpus pairs that are more than the K candidate corpora and are similar to the candidate corpora can be obtained, and more target corpus pairs similar to the corpus to be expanded can be mined, so that the corpus to be expanded is enriched, thereby meeting the requirement of the model training for the quantity of the corpus to a certain extent.
[0139] For example, assume there are 5 candidate corpora. If a corpus has more than half of the total number of characters in common with the corpus to be expanded, and more than half of the total number of words in common with the corpus to be expanded, it is taken as the first candidate corpus; otherwise, it is taken as the second candidate corpus. Assume 2 first candidate corpora and 3 second candidate corpora are obtained. Then, by combining the 2 first candidate corpora with the 3 second candidate corpora respectively, 6 target corpus pairs can be obtained.
[0140] Optionally, based on the above Figure 2 corresponding embodiment, in another optional embodiment of the corpus processing method provided by the embodiments of the present application, as Figure 8 shown, if at least one semantic recognition result among the K semantic recognition results meets the corpus extraction condition, the candidate corpus corresponding to at least one semantic recognition result is determined as the target corpus, so as to obtain at least one target corpus belonging to the corpus to be expanded, including:
[0141] In step S801, if each semantic recognition result is a similarity score, the candidate corpus corresponding to the similarity score greater than or equal to the preset similarity threshold is determined as the target corpus;
[0142] In step S802, if each semantic recognition result is a similarity classification, the candidate corpus corresponding to the similarity classification greater than or equal to the preset classification probability threshold is determined as the target corpus.
[0143] In this embodiment, after obtaining the K semantic recognition results, in order to more quickly and accurately obtain the target corpus, when each semantic recognition result is a similarity score, the candidate corpus corresponding to the similarity score greater than or equal to the preset similarity threshold can be determined as the target corpus, where the preset similarity threshold is set according to actual application requirements and is not specifically limited here. Or, when each semantic recognition result is a similarity classification, the candidate corpus corresponding to the similarity classification greater than or equal to the preset classification probability threshold is determined as the target corpus, so as to quickly screen out the target corpus similar to the corpus to be expanded, thereby improving the processing efficiency of the corpus to be expanded to a certain extent. The preset classification probability threshold is set according to actual application requirements and is not specifically limited here.
[0144] Specifically, when the K semantic recognition results are K similarity scores, it can be understood that the larger the similarity score, the higher the similarity between the candidate corpus and the corpus to be expanded. Then, the similarity scores can be compared with a preset similarity threshold, and the candidate corpus corresponding to the similarity score greater than or equal to the similarity threshold can be determined as the target corpus similar to the corpus to be expanded. Or, when the K semantic recognition results are K similarity classifications, since the similarity classification is the semantic category to which the candidate corpus and the corpus to be expanded belong, it can be expressed as K category probabilities. That is, it can be understood that the larger the category probability, the higher the similarity of the semantic category to which the candidate corpus and the corpus to be expanded belong. Then, the similarity classification can be compared with a preset classification probability threshold, and the candidate corpus corresponding to the category probability greater than or equal to the preset classification probability threshold can be determined as the target corpus similar to the corpus to be expanded.
[0145] For example, assume that the similarity scores between a corpus to be expanded "What songs are suitable to listen to on a rainy day" and three candidate corpora "Music related to rainy days", "Songs one wants to listen to on a rainy day", and "What to do on a rainy day" are 94, 91, and 37 respectively. Assume a preset similarity threshold is 72. Then, by comparing these similarity scores with the preset similarity threshold, it can be clearly seen that the similarity scores 94 and 91 are greater than the similarity threshold of 72. Then, the candidate corpora "Music related to rainy days" and "Songs one wants to listen to on a rainy day" can be determined as the target corpora of the corpus to be expanded "What songs are suitable to listen to on a rainy day".
[0146] For example, assume that the category probabilities of the semantic categories to which the corpus to be expanded "What songs are suitable to listen to on a rainy day" and three candidate corpora "Music related to rainy days", "Songs one wants to listen to on a rainy day", and "What to do on a rainy day" belong are 0.91, 0.89, and 0.39 respectively. Assume a preset probability threshold is 0.7. Then, by comparing these classification probabilities with the preset classification probability threshold, it can be clearly seen that the category probabilities 0.91 and 0.89 are greater than the probability threshold of 0.7. Then, the candidate corpora "Music related to rainy days" and "Songs one wants to listen to on a rainy day" can be determined as the target corpora of the corpus to be expanded "What songs are suitable to listen to on a rainy day".
[0147] The apparatus for corpus processing in the present application will be described in detail below. Please refer to Figure 13 , Figure 13 which is a schematic diagram of an embodiment of the apparatus for corpus processing in an embodiment of the present application. The apparatus for corpus processing 20 includes:
[0148] An acquisition unit 201, configured to acquire the corpus to be expanded;
[0149] The obtaining unit 201 is further configured to obtain K candidate corpora according to the corpus to be augmented, where the semantic similarity between each candidate corpus and the corpus to be augmented is greater than or equal to a similarity threshold, and K is an integer greater than 1;
[0150] The processing unit 202 is configured to input the K candidate corpora and the corpus to be augmented into a semantic recognition model to obtain K semantic recognition results, where each semantic recognition result is a similarity score or a similarity classification. The similarity score represents the degree of semantic similarity between the candidate corpus and the corpus to be augmented, and the similarity classification represents the semantic category to which the candidate corpus and the corpus to be augmented belong;
[0151] The determining unit 203 is configured to, if at least one of the K semantic recognition results meets the corpus extraction condition, determine the candidate corpora corresponding to the at least one semantic recognition result as target corpora, so as to obtain at least one target corpus belonging to the corpus to be augmented.
[0152] Optionally, on the basis of the above Figure 13 corresponding embodiment, in another embodiment of the corpus processing apparatus provided in the embodiments of the present application,
[0153] The obtaining unit 201 is further configured to obtain a corpus sample set, where the corpus sample set includes positive sample corpora and negative sample corpora, and the positive sample corpora correspond to labeled tags;
[0154] The processing unit 202 is further configured to perform feature extraction on the positive sample corpora to obtain positive sample corpus features, and perform feature extraction on the negative sample corpora to obtain negative sample corpus features;
[0155] The processing unit 202 is further configured to input the positive sample corpus features and the negative sample corpus features into a semantic recognition model to obtain a semantic prediction result;
[0156] The processing unit 202 is further configured to train the semantic recognition model according to the semantic prediction result and the labeled tags.
[0157] Optionally, on the basis of the above Figure 13 corresponding embodiment, in another embodiment of the corpus processing apparatus provided in the embodiments of the present application, the processing unit 202 may specifically be configured to:
[0158] Perform word segmentation on the positive sample corpora and the negative sample corpora respectively to obtain at least two to-be-processed positive sample words and at least two to-be-processed negative sample words;
[0159] Convert the at least two to-be-processed positive sample words and the at least two to-be-processed negative sample words into at least two positive sample word vectors and at least two negative sample word vectors;
[0160] Concatenate at least two positive sample word vectors to obtain positive sample corpus features, and concatenate at least two negative sample word vectors to obtain negative sample corpus features.
[0161] Optionally, based on the corresponding embodiment above, in another embodiment of the corpus processing apparatus provided by the embodiments of the present application, Figure 13 The processing unit 202 is further configured to, if there is no semantic recognition result among the K semantic recognition results that satisfies the corpus extraction condition, translate the corpus to be augmented into N language corpora corresponding to N languages;
[0162] The processing unit 202 is further configured to translate the N language corpora into at least N back-translated corpora according to the language of the corpus to be augmented.
[0163] Optionally, based on the corresponding embodiment above, in another embodiment of the corpus processing apparatus provided by the embodiments of the present application,
[0164] The determination unit 203 is further configured to determine at least one target corpus and at least N back-translated corpora as multiple corpora to be annotated; Figure 13 The processing unit 202 is further configured to perform slot matching on each corpus to be annotated to obtain a slot matching result corresponding to each corpus to be annotated, where the slot matching result is a matching similarity score, and the matching similarity score represents the semantic similarity degree between a preset slot and each word to be annotated in each corpus to be annotated;
[0165] The determination unit 203 is further configured to, if the matching similarity score is greater than or equal to a preset matching threshold, determine the preset slot corresponding to the matching similarity score as the target slot, and determine the word to be annotated corresponding to the matching similarity score as the target slot value, where the target slot is used to represent the attribute of the target slot value;
[0166] The processing unit 202 is further configured to perform deduplication processing on the corpus to be annotated, the target slot corresponding to the corpus to be annotated, and the target slot value to obtain the target annotated corpus.
[0167] Optionally, based on the corresponding embodiment above, in another embodiment of the corpus processing apparatus provided by the embodiments of the present application,
[0168] The processing unit 202 is further configured to perform deduplication processing on the corpus to be annotated, the target slot corresponding to the corpus to be annotated, and the target slot value to obtain the target annotated corpus.
[0169] Optionally, based on the corresponding embodiment above, in another embodiment of the corpus processing apparatus provided by the embodiments of the present application, Figure 13 The processing unit 202 is further configured to perform deduplication processing on the corpus to be annotated, the target slot corresponding to the corpus to be annotated, and the target slot value to obtain the target annotated corpus.
[0170] The processing unit 202 is further configured to classify the K candidate corpora into a first candidate corpus set and a second candidate corpus set if there is no semantic recognition result among the K semantic recognition results that meets the corpus extraction condition, where the first candidate corpus set includes i first candidate corpora, the second candidate corpus set includes j second candidate corpora, and both i and j are integers greater than 1 and less than K;
[0171] The processing unit 202 is further configured to pair the i first candidate corpora with the j second candidate corpora pairwise to obtain i*j target corpus pairs.
[0172] Optionally, based on the corresponding embodiment above Figure 13 In another embodiment of the corpus processing device provided by the embodiment of the present application, the processing unit 202 may specifically be configured to:
[0173] If each semantic recognition result is a similarity score, determine the candidate corpus corresponding to the similarity score greater than or equal to a preset similarity threshold as the target corpus;
[0174] If each semantic recognition result is a similarity classification, determine the candidate corpus corresponding to the similarity classification greater than or equal to a preset classification probability threshold as the target corpus.
[0175] On the other hand, the present application provides another schematic diagram of a computer device, as Figure 14 shown, Figure 14 is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 300 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more mass storage devices) for storing application programs 331 or data 332. Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device 300. Further, the central processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the computer device 300.
[0176] The computer device 300 may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or, one or more operating systems 333, such as Windows Server TM, Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0177] The above computer device 300 is also used to execute the steps in the corresponding embodiments as Figures 2 to 8 shown.
[0178] On the other hand, the present application provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps in the method described in the embodiments as Figures 2 to 8 shown.
[0179] On the other hand, the present application provides a computer program product including instructions. When the computer program product runs on a computer or a processor, the computer or the processor is caused to execute the steps in the method described in the embodiments as Figures 2 to 8 shown.
[0180] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0181] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0182] In several embodiments provided by the present application, it should be understood that the disclosed system, device, and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.
[0183] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0184] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0185] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
Claims
1. A corpus processing method, characterized in that, Including: Obtain the corpus to be expanded; Obtain K candidate corpora according to the corpus to be expanded, where the semantic similarity between each candidate corpus and the corpus to be expanded is greater than or equal to a similarity threshold, and K is an integer greater than 1; Input the K candidate corpora and the corpus to be expanded into a semantic recognition model to obtain K semantic recognition results, where each semantic recognition result is a similarity score or a similarity classification. The similarity score represents the degree of semantic similarity between the candidate corpus and the corpus to be expanded, and the similarity classification represents the semantic category to which the candidate corpus and the corpus to be expanded belong; If at least one of the K semantic recognition results meets the corpus extraction condition, determine the candidate corpus corresponding to the at least one semantic recognition result as the target corpus to obtain at least one target corpus belonging to the corpus to be expanded; If none of the K semantic recognition results meets the corpus extraction condition, translate the corpus to be expanded into N language corpora corresponding to N languages; Translate the N language corpora into at least N back-translated corpora according to the language of the corpus to be expanded; If none of the K semantic recognition results meets the corpus extraction condition, classify the K candidate corpora to obtain a first candidate corpus set and a second candidate corpus set, where the first candidate corpus set includes i first candidate corpora, and the second candidate corpus set includes j second candidate corpora. Both i and j are integers greater than 1 and less than K; the similarity between the first candidate corpus set and the corpus to be expanded is greater than the similarity between the second candidate corpus set and the corpus to be expanded; Combine the i first candidate corpora with the j second candidate corpora pairwise to obtain i*j target corpus pairs.
2. The method according to claim 1, characterized in that, The method further includes: Obtain a corpus sample set, where the corpus sample set includes positive sample corpora and negative sample corpora, and the positive sample corpora correspond to labeled tags; Extract features from the positive sample corpora to obtain positive sample corpus features, and extract features from the negative sample corpora to obtain negative sample corpus features; Input the positive sample corpus features and the negative sample corpus features into the semantic recognition model to obtain a semantic prediction result; Train the semantic recognition model according to the semantic prediction result and the labeled tag.
3. The method according to claim 2, wherein The extracting features from the positive sample corpora to obtain positive sample corpus features and extracting features from the negative sample corpora to obtain negative sample corpus features includes: Perform word segmentation on the positive sample corpora and the negative sample corpora respectively to obtain at least two positive sample words to be processed and at least two negative sample words to be processed; Convert the at least two positive sample words to be processed and the at least two negative sample words to be processed into at least two positive sample word vectors and at least two negative sample word vectors; Concatenate the at least two positive sample word vectors to obtain the positive sample corpus feature, and concatenate the at least two negative sample word vectors to obtain the negative sample corpus feature.
4. The method according to claim 1, characterized in that, After translating the N language corpora into N back-translated corpora according to the language of the corpus to be augmented, the method further includes: Determine the at least one target corpus and the at least N back-translated corpora as multiple corpora to be annotated; Perform slot matching on each corpus to be annotated to obtain a slot matching result corresponding to each corpus to be annotated, where the slot matching result is a matching similarity score, and the matching similarity score represents the semantic similarity degree between a preset slot and each word to be annotated in each corpus to be annotated; If the matching similarity score is greater than or equal to a preset matching threshold, determine the preset slot corresponding to the matching similarity score as the target slot, and determine the word to be annotated corresponding to the matching similarity score as the target slot value, where the target slot is used to represent the attribute of the target slot value; Perform deduplication processing on the corpus to be annotated, the target slot corresponding to the corpus to be annotated, and the target slot value to obtain the target annotated corpus.
5. The method according to claim 1, wherein The step of, if at least one of the K semantic recognition results satisfies the corpus extraction condition, determining the candidate corpus corresponding to the at least one semantic recognition result as the target corpus to obtain at least one target corpus belonging to the corpus to be augmented, includes: If each semantic recognition result is a similarity score, determine the candidate corpus corresponding to the similarity score greater than or equal to a preset similarity threshold as the target corpus; If each semantic recognition result is a similarity classification, determine the candidate corpus corresponding to the similarity classification greater than or equal to a preset classification probability threshold as the target corpus.
6. A corpus processing device, characterized in that, Including: An acquisition unit for acquiring a corpus to be augmented; The acquisition unit is further configured to acquire K candidate corpora according to the corpus to be augmented, where the semantic similarity between each candidate corpus and the corpus to be augmented is greater than or equal to a similarity threshold, and K is an integer greater than 1; A processing unit for inputting the K candidate corpora and the corpus to be augmented into a semantic recognition model to obtain K semantic recognition results, where each semantic recognition result is a similarity score or a similarity classification, the similarity score represents the semantic similarity degree between the candidate corpus and the corpus to be augmented, and the similarity classification represents the semantic category to which the candidate corpus and the corpus to be augmented belong; A determination unit for, if at least one of the K semantic recognition results satisfies the corpus extraction condition, determining the candidate corpus corresponding to the at least one semantic recognition result as the target corpus to obtain at least one target corpus belonging to the corpus to be augmented; The processing unit is further configured to, if none of the K semantic recognition results satisfies the corpus extraction condition, translate the corpus to be augmented into N language corpora corresponding to N languages. The processing unit is further configured to translate the N language corpora into at least N back-translated corpora according to the language of the corpus to be augmented; The processing unit is further configured to, if there is no semantic recognition result among the K semantic recognition results that meets the corpus extraction condition, classify the K candidate corpora to obtain a first candidate corpus set and a second candidate corpus set, where the first candidate corpus set includes i first candidate corpora, the second candidate corpus set includes j second candidate corpora, and both i and j are integers greater than 1 and less than K; the similarity between the first candidate corpus set and the corpus to be augmented is greater than the similarity between the second candidate corpus set and the corpus to be augmented; The processing unit is further configured to combine the i first candidate corpora with the j second candidate corpora pairwise to obtain i*j target corpus pairs.
7. The apparatus according to claim 6, characterized in that, The obtaining unit is further configured to obtain a corpus sample set, where the corpus sample set includes positive sample corpora and negative sample corpora, and the positive sample corpora correspond to labeled tags; The processing unit is further configured to perform feature extraction on the positive sample corpora to obtain positive sample corpus features, and perform feature extraction on the negative sample corpora to obtain negative sample corpus features; The processing unit is further configured to input the positive sample corpus features and the negative sample corpus features into the semantic recognition model to obtain a semantic prediction result; The processing unit is further configured to train the semantic recognition model according to the semantic prediction result and the labeled tag.
8. The device according to claim 7, characterized in that, Specifically, the processing unit may be configured to: Perform word segmentation on the positive sample corpora and the negative sample corpora respectively to obtain at least two to-be-processed positive sample words and at least two to-be-processed negative sample words; Convert the at least two to-be-processed positive sample words and the at least two to-be-processed negative sample words into at least two positive sample word vectors and at least two negative sample word vectors; Perform vector splicing on the at least two positive sample word vectors to obtain the positive sample corpus features, and perform vector splicing on the at least two negative sample word vectors to obtain the negative sample corpus features.
9. The device according to claim 6, characterized in that, The determining unit is further configured to determine the at least one target corpus and the at least N back-translated corpora as a plurality of to-be-labeled corpora; The processing unit is further configured to perform slot matching on each to-be-labeled corpus to obtain a slot matching result corresponding to each to-be-labeled corpus, where the slot matching result is a matching similarity score, and the matching similarity score represents the semantic similarity degree between a preset slot and each to-be-labeled word in each to-be-labeled corpus; The determining unit is further configured to, if the matching similarity score is greater than or equal to a preset matching threshold, determine the preset slot corresponding to the matching similarity score as the target slot, and determine the to-be-labeled word corresponding to the matching similarity score as the target slot value, where the target slot is used to represent the attribute of the target slot value; The processing unit is further configured to perform deduplication processing on the to-be-annotated corpus, the target slot corresponding to the to-be-annotated corpus, and the target slot value, so as to obtain a target annotated corpus.
10. The device according to claim 6, characterized in that, Specifically, the processing unit may be configured to: If each semantic recognition result is a similarity score, determine the candidate corpus corresponding to the similarity score greater than or equal to a preset similarity threshold as the target corpus; If each semantic recognition result is a similarity classification, determine the candidate corpus corresponding to the similarity classification greater than or equal to a preset classification probability threshold as the target corpus.
11. A computer device, characterized in that, It includes: A memory, a transceiver, a processor, and a bus system; Wherein, the memory is used to store programs; The processor is configured to implement the method according to any one of claims 1 to 5 when executing the programs in the memory; The bus system is used to connect the memory and the processor, so that the memory and the processor can communicate with each other.
12. A computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the method according to any one of claims 1 to 5.
13. A computer program product, characterized in that, The computer program product includes instructions, which when running on a computer device, cause the computer device to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and apparatus for identifying user intention, terminal, and computer-readable storage medium
CN109376847A
Entity recognition method and device and computer equipment
CN109918680A
Semantic matching model training method, semantic matching method and answer acquisition method
CN110895553A
Extended corpus generation method and device in target field and electronic equipment
CN112541076A