Recognition model training method and device, text recognition method and device

Through the training method of identifying models, using joint learning of fill modules and intention modules, the problems of time-consuming and resource consumption in the existing technology are solved, and the recognition accuracy and training speed are improved.

CN113971399BActive Publication Date: 2025-05-06BEIJING KINGSOFT DIGITAL ENTERTAINMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010716481.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-23
Publication Date
2025-05-06
Estimated Expiration
2040-07-23

AI Technical Summary

Technical Problem

The prior art relies on Attention RNN model in terms of speech information understanding and intention reasoning. Although semantic understanding can be achieved, the preliminary preparation process of the model is time-consuming and resource-consuming, and requires retraining, resulting in inefficiency.

Method used

A training method for identifying models is proposed. By obtaining training text, inputting the language module in the recognized model to be trained for encoding processing, obtaining feature vectors, and inputting the feature vectors into the fill module and the intention module for processing, determining the relationship weight between the two, adjusting the module to obtain the loss value, and iterative training is performed until the training stop condition is reached.

Benefits of technology

Through the joint learning of the fill module and the intent module, the accuracy and training speed of the identification model are improved, the analysis accuracy of intention understanding and slot filling is enhanced, and the time and resource consumption of model training are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113971399B_ABST
    Figure CN113971399B_ABST
Patent Text Reader

Abstract

The present application provides a training method and device for a recognition model, and a text recognition method and device, wherein the training method for the recognition model includes: obtaining a training text; inputting the training text into a recognition model to be trained, encoding the training text through a language module in the recognition model to be trained, and obtaining a feature vector; inputting the feature vector into a filling module and an intention module in the recognition model to be trained respectively for processing, and determining a relationship weight between the filling module and the intention module according to the processing result; adjusting the filling module and the intention module based on the relationship weight, and obtaining a first loss value of the filling module and a second loss value of the intention module according to the adjustment result; iteratively training the recognition model to be trained based on the first loss value and the second loss value until a training stop condition is reached, and obtaining a target recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language processing, and in particular to a method and device for training a recognition model, and a method and device for text recognition. Background Art

[0002] Natural Language Processing (NLP) is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. In the existing technology, most of the RNN (Recurrent Neural Network, RNN) models based on Attention are used to understand speech information and infer internal relationships. Although the purpose of semantic understanding can be achieved, the preliminary preparation process of the model consumes a lot of time, and the prepared model needs to be retrained after a period of application, which is not only time-consuming but also resource-consuming. Therefore, an effective solution is urgently needed to solve the above problems. Summary of the invention

[0003] In view of this, the embodiment of the present application provides a method for training a recognition model to solve the technical defects existing in the prior art. The embodiment of the present application also provides a text recognition method, a training device for a recognition model, a text recognition device, a computing device, and a computer-readable storage medium.

[0004] According to a first aspect of an embodiment of the present application, a method for training a recognition model is provided, comprising:

[0005] Get training text;

[0006] Inputting the training text into the recognition model to be trained, encoding the training text through the language module in the recognition model to be trained to obtain a feature vector;

[0007] The feature vectors are respectively input into a filling module and an intention module in the recognition model to be trained for processing, and a relationship weight between the filling module and the intention module is determined according to the processing result;

[0008] Adjusting the filling module and the intention module based on the relationship weight, and obtaining a first loss value of the filling module and a second loss value of the intention module according to the adjustment result;

[0009] The recognition model to be trained is iteratively trained based on the first loss value and the second loss value until a training stop condition is reached to obtain a target recognition model.

[0010] Optionally, the step of inputting the feature vectors into a filling module and an intention module in the recognition model to be trained for processing, and determining a relationship weight between the filling module and the intention module according to the processing result, comprises:

[0011] Inputting the feature vector into the filling module in the recognition model to be trained for processing to obtain a first intermediate vector, and inputting the feature vector into the intention module in the recognition model to be trained for processing to obtain a second intermediate vector;

[0012] The relationship weight between the filling module and the intention module is calculated based on the first intermediate vector and the second intermediate vector.

[0013] Optionally, the adjusting the filling module and the intention module based on the relationship weight, and obtaining a first loss value of the filling module and a second loss value of the intention module according to the adjustment result, includes:

[0014] Obtaining a first text vector obtained by the filling module processing the feature vector, and a second text vector obtained by the intention module processing the feature vector;

[0015] Calculating the product of the relationship weight and the first text vector to obtain a first target text vector, and calculating the product of the relationship weight and the second text vector to obtain a second target text vector;

[0016] Extracting the entity vector and the intention vector corresponding to the training text in the training set to which the training text belongs;

[0017] The first loss value of the filling module is determined according to the first target text vector and the entity vector, and the second loss value of the intent module is determined according to the second target text vector and the intent vector.

[0018] Optionally, the iteratively training the recognition model to be trained based on the first loss value and the second loss value includes:

[0019] Calculate the target loss value of the recognition model to be trained according to the first loss value and the second loss value;

[0020] The recognition model to be trained is iteratively trained according to the target loss value.

[0021] Optionally, the training stop condition is that the target loss value is less than or equal to a loss value threshold;

[0022] Correspondingly, if the target loss value is less than or equal to the loss value threshold, the target recognition model is obtained;

[0023] The loss value threshold is determined based on each target loss value generated by the recognition model to be trained during the training process.

[0024] Optionally, calculating a target loss value of the recognition model to be trained according to the first loss value and the second loss value includes:

[0025] Determining a weight value of the first loss value and a weight value of the second loss value;

[0026] A weighted summation process is performed on the weight value of the first loss value and the weight value of the second loss value to obtain the target loss value.

[0027] Optionally, before the step of inputting the training text into the recognition model to be trained, encoding the training text through the language module in the recognition model to be trained, and obtaining the feature vector, the step further includes:

[0028] Performing word segmentation processing on the training text to obtain a word unit set;

[0029] Correspondingly, the step of inputting the training text into the recognition model to be trained, encoding the training text through the language module in the recognition model to be trained, and obtaining a feature vector includes:

[0030] The word unit set is input into the recognition model to be trained, and the word unit set is encoded by the language module to obtain the feature vector.

[0031] According to a second aspect of an embodiment of the present application, a text recognition method is provided, comprising:

[0032] Get the text to be recognized;

[0033] Inputting the text to be recognized into a target recognition model for entity extraction and intent understanding, and obtaining a target entity and a target intent corresponding to the text to be recognized;

[0034] Wherein, the target recognition model is obtained by training using any of the recognition model training methods described above.

[0035] Optionally, also include:

[0036] Selecting a target template from the preset filling templates according to the target intention;

[0037] The target entity is structurally processed according to the filling rule of the target template and the target template to obtain the target text.

[0038] Optionally, the step of inputting the text to be recognized into a target recognition model for entity extraction and intent understanding to obtain a target entity and a target intent corresponding to the text to be recognized includes:

[0039] Inputting the text to be recognized into the target recognition model, encoding the text to be recognized through the language module in the target recognition model to obtain a text feature vector;

[0040] The text feature vector is input into the filling module and the intention module in the target recognition model for processing, and the target entity and the target intention are obtained according to the processing results.

[0041] Optionally, before the step of inputting the text to be recognized into the target recognition model, encoding the text to be recognized through a language module in the target recognition model, and obtaining a text feature vector, the step further includes:

[0042] Performing word segmentation processing on the text to be recognized to obtain a word unit set;

[0043] Correspondingly, the text to be recognized is input into the target recognition model, and the language module in the target recognition model encodes the text to be recognized to obtain a text feature vector, including:

[0044] The word unit set is input into the language module for encoding processing to obtain the text feature vector.

[0045] Optionally, the step of inputting the text feature vector into a filling module and an intention module in the target recognition model for processing, and obtaining the target entity and the target intention according to the processing results, comprises:

[0046] Inputting the text feature vector into the filling module for processing to obtain a first intermediate vector, and inputting the text feature vector into the intention module for processing to obtain a second intermediate vector;

[0047] Calculate the product of the first intermediate vector and the first relationship weight of the intention module to obtain a first text vector, and calculate the product of the second intermediate vector and the second relationship weight of the filling module to obtain a second text vector;

[0048] The first text vector is converted through the output layer of the filling module to obtain the target entity, and the second text vector is converted through the output layer of the intention module to obtain the target intention.

[0049] According to a third aspect of an embodiment of the present application, a training device for a recognition model is provided, comprising:

[0050] An acquisition unit, configured to acquire training text;

[0051] An encoding unit is configured to input the training text into the recognition model to be trained, and perform encoding processing on the training text through the language module in the recognition model to be trained to obtain a feature vector;

[0052] A determination unit is configured to input the feature vector into a filling module and an intention module in the recognition model to be trained for processing, and determine a relationship weight between the filling module and the intention module according to the processing result;

[0053] an adjusting unit, configured to adjust the filling module and the intention module based on the relationship weight, and obtain a first loss value of the filling module and a second loss value of the intention module according to the adjustment result;

[0054] The training unit is configured to iteratively train the recognition model to be trained based on the first loss value and the second loss value until a training stop condition is reached to obtain a target recognition model.

[0055] According to a fourth aspect of an embodiment of the present application, a text recognition device is provided, including:

[0056] A text acquisition unit is configured to acquire text to be recognized;

[0057] A model processing unit is configured to input the text to be recognized into a target recognition model for entity extraction and intent understanding, and obtain a target entity and a target intent corresponding to the text to be recognized;

[0058] Wherein, the target recognition model is obtained by training using any of the recognition model training methods described above.

[0059] According to a fifth aspect of an embodiment of the present application, a computing device is provided, including:

[0060] Memory and processor;

[0061] The memory is used to store computer-executable instructions, and the processor implements the steps of the recognition model training method and the text recognition method when executing the computer-executable instructions.

[0062] According to a sixth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the training method of the recognition model and the text recognition method are implemented.

[0063] The training method of the recognition model provided in the present application, in the process of training the recognition model to be trained, uses a language module to encode the training text to obtain a feature vector, then processes the feature vector according to the filling module and the intention module in the recognition model to be trained, and determines the relationship weight between the two according to the processing result, then adjusts the filling module and the intention module according to the relationship weight, and obtains the first loss value of the filling module and the second loss value of the intention module, and finally iteratively trains the recognition model to be trained according to the first loss value and the second loss value, so as to improve the accuracy of the recognition model to be trained by joint learning of the filling module and the intention module, and by combining the relationship weight, the filling module and the intention module assist each other in semantic understanding, thereby further improving the analysis accuracy of intent understanding and slot filling, and training the model in parallel can effectively improve the model training speed.

[0064] The text recognition method provided in the present application can effectively improve the recognition accuracy by adopting a target recognition model to perform entity extraction and intent understanding on the text to be recognized, combined with the checks and balances mechanism between the filling module and the intent module in the target recognition model. The target recognition model uses a language module combined with a joint learning method of intent understanding and slot filling to extract target entities and relationships, further improving the natural language structuring effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is a flow chart of a method for training a recognition model provided in one embodiment of the present application;

[0066] Figure 2 is a schematic diagram of determining a relationship weight in a recognition model provided in an embodiment of the present application;

[0067] Figure 3 is a flowchart of a text recognition method provided by an embodiment of the present application;

[0068] Figure 4 is a schematic diagram of a text recognition method provided by an embodiment of the present application;

[0069] Figure 5 This is a processing flow chart for a text recognition scenario provided by an embodiment of the present application;

[0070] Figure 6 It is a structural schematic diagram of a training device for a recognition model provided in one embodiment of the present application;

[0071] Figure 7 is a structural schematic diagram of a text recognition device provided by an embodiment of the present application;

[0072] Figure 8It is a structural block diagram of a computing device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0073] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.

[0074] The terms used in one or more embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present application. The singular forms of "a", "said" and "the" used in one or more embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more associated listed items.

[0075] It should be understood that, although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first.

[0076] First, the terms involved in one or more embodiments of the present invention are explained.

[0077] BERT model: (Bidirectional Encoder Representation from Transformers, BERT) is a bidirectional attention neural network model. The BERT model can predict the current word through the left and right contexts and predict the next sentence through the current sentence. The goal of the BERT model is to use large-scale unlabeled corpus training to obtain the semantic representation of text containing rich semantic information, and then fine-tune the semantic representation of the text in a specific NLP task, and finally apply it to the NLP task.

[0078] Token: Before any actual processing is done on the input text, it needs to be segmented into language units such as words, punctuation marks, numbers or letters. These units are called tokens. For English text, a token can be a word, a punctuation mark, a number, etc. For Chinese text, the smallest token can be a word, a character, a punctuation mark, a number, etc.

[0079] Word segmentation: The process of recombining continuous character sequences into word sequences according to certain specifications.

[0080] Intent understanding: refers to understanding the intention expressed by the text. Intent understanding can be regarded as a classification problem, that is, determining what the text wants to express, such as whether the text expresses navigational meaning, informational meaning, or transactional meaning. Among them, the intent module is used for intent understanding.

[0081] Slot filling: (slot filling) refers to sequence labeling. Slot filling can be regarded as an entity extraction, that is, labeling entities in the text; among them, the filling module is used for slot filling.

[0082] Loss function: A function that maps the value of a random event or its related random variables to a non-negative real number to represent the loss of the random event. In applications, loss functions are usually associated with optimization problems as learning criteria, that is, solving and evaluating models by minimizing loss functions; for example, they are used for parameter estimation of models in statistics and machine learning.

[0083] Iteration: It is the activity of repeating a feedback process, usually to approach the desired goal or result. Each repetition of the process is called an iteration, and the result of each iteration will be used as the initial value of the next iteration.

[0084] Encoding: is the process of converting information from one form or format to another.

[0085] Weight: (weight) refers to the importance of a factor or indicator relative to a certain thing. It is different from the general proportion. It reflects not only the percentage of a factor or indicator, but emphasizes the relative importance of the factor or indicator, tending to the contribution or importance.

[0086] RNN: (Recurrent Neural Network) is a type of recursive neural network that takes sequence data as input, performs recursion in the direction of sequence evolution, and all nodes (recurrent units) are connected in a chain.

[0087] Attention: is a means of dealing with information overload. It is manifested in how different areas of an image or related words in a sentence are paid different attention. Usually, a lot of attention is allocated to the parts of interest.

[0088] In the present application, a method for training a recognition model is provided. The present application also relates to a text recognition method, a training device for a recognition model, a text recognition device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0089] Figure 1 is a flow chart of a method for training a recognition model provided in an embodiment of the present application. Figure 2 is a schematic diagram of determining relationship weights in a recognition model provided in an embodiment of the present application, wherein Figure 2 including (a) and (b), Figure 1 The specific steps include:

[0090] Step S102, obtaining training text.

[0091] In practical applications, natural language structuring enables robots to effectively store and understand semantic information and reason about internal connections. This is usually done by using an Attention-based RNN model combined with intent understanding / slot filling joint learning to achieve optimal results on multiple data sets. This model focuses on learning the relationship between intent understanding and slot filling. Although it can achieve the effect of semantic understanding, the training speed is slow during model training, the recognition accuracy is low, and the natural language structuring effect is poor.

[0092] The training method of the recognition model provided in this embodiment, in order to improve the training speed and recognition accuracy of the recognition model, uses a language module to encode the training text to obtain a feature vector during the training of the recognition model to be trained, and then processes the feature vector according to the filling module and the intention module in the recognition model to be trained, and determines the relationship weight between the two according to the processing result, and then adjusts the filling module and the intention module according to the relationship weight, and obtains the first loss value of the filling module and the second loss value of the intention module, and finally iteratively trains the recognition model to be trained according to the first loss value and the second loss value, so as to improve the accuracy of the recognition model to be trained by joint learning of the filling module and the intention module, and by combining the relationship weight, the filling module and the intention module assist each other in semantic understanding, thereby further improving the analysis accuracy of intention understanding and slot filling, and training the model in parallel can effectively improve the model training speed.

[0093] In specific implementation, the training text specifically refers to the text used to train the recognition model to be trained. The training text can be a sentence, a paragraph, or an article, etc. This embodiment will take the training text as a sentence as an example to illustrate the training method of the recognition model. It should be noted that the training process of the training text as a paragraph or an article can refer to the corresponding description content of this embodiment, and this embodiment will not be elaborated in detail here.

[0094] In actual applications, the training text can be created according to the training requirements of the recognition model, or it can be a collected real text, or the training text for training the recognition model can be composed of mixed training text and collected real text. The creation method of the training text can be set according to actual needs, and this embodiment does not make too many restrictions here.

[0095] Step S104: input the training text into the recognition model to be trained, and encode the training text through the language module in the recognition model to be trained to obtain a feature vector.

[0096] Specifically, the recognition model to be trained refers to a model that can understand the intent of the text and fill in the slots. Intent understanding specifically refers to the ability to recognize the intent expressed by the text. Intent understanding can be regarded as a classification problem, that is, determining the meaning that the text wants to express, such as whether the text represents navigation, information or transactional meanings. Slot filling specifically refers to sequence labeling, which can be understood as extracting entities from the text, that is, performing entity labeling in the text to extract important information from the text.

[0097] Based on this, on the basis of obtaining the training text as mentioned above, further, it will be necessary to train the recognition model to be trained according to the training text, specifically, inputting the training text into the recognition model to be trained, and encoding the training text through the language module in the recognition model to be trained, so as to obtain the feature vector corresponding to the training text, which is used to train the recognition model to be trained.

[0098] In practical applications, the language module can be a BERT model or other trained embeddings, both of which can be used to encode the training text. In this embodiment, the BERT model is used as the language module for encoding, which can achieve a good encoding effect. For example, the training text is "He went to the mountain to see the sunrise". The training text "He went to the mountain to see the sunrise" is input into the BERT model in the recognition model to be trained for encoder, so that the feature vector Sn corresponding to the training text can be obtained. 11 , S 12 , S 13 , S 14 , S 15 , S 16 , S 17 , S 18 , S 19 ] for subsequent training of the recognition model to be trained.

[0099] Furthermore, before the language module in the to-be-trained recognition model is encoded, in order to speed up the encoding process of the training text, the training text may also be segmented, thereby improving the encoding efficiency of the language module. In this embodiment, the specific implementation method is as follows:

[0100] Performing word segmentation processing on the training text to obtain a word unit set;

[0101] The word unit set is input into the recognition model to be trained, and the word unit set is encoded by the language module to obtain the feature vector.

[0102] Specifically, the word segmentation processing specifically refers to processing the training text into multiple word units according to certain specifications. The word units contained in the word unit set can constitute the training text. After obtaining the word unit set, the word unit set is input into the recognition model to be trained. The language module in the model encodes each word unit in the word unit set to obtain the feature vector.

[0103] For example, the training sample obtained is "He went to the mountain to see the sunrise". By performing word segmentation on the training sample, a word unit set consisting of "He", "went", "to", "the", "mountain", "to", "see", "the" and "sunrise" is obtained. Then, the word unit set is input into the recognition model to be trained. The language module in the recognition model to be trained encodes each word unit in the word unit set to obtain a sub-feature vector He = S 11 、went=S 12 、to=S 13 、the=S 14 、mountain=S 15 、to=S 16 、see=S 17 、the=S 18 and sunrise = S 19 The characteristic vector Sn = [S 11 , S 12 , S 13 , S 14 , S 15 , S 16 , S 17 , S 18 , S 19 ], the feature vector Sn is the vector expression obtained after encoding the training sample.

[0104] In summary, by performing word segmentation on the training text before encoding the training text, the encoding processing efficiency of the language module can be effectively improved, thereby further improving the training speed of the recognition model to be trained.

[0105] Step S106, inputting the feature vector into the filling module and the intention module in the recognition model to be trained for processing, and determining the relationship weight between the filling module and the intention module according to the processing result.

[0106] Specifically, on the basis of the feature vector corresponding to the training text obtained above, further, in order to enable the recognition model to be trained to simultaneously complete intent understanding and slot filling, it is also necessary to train the filling module and the intent module in the recognition model to be trained, and since the filling module and the intent module are arranged in the recognition model to be trained at the same time, it is necessary to consider the influence of the filling module on the intent module, and the influence of the intent module on the filling module, so the feature vector is input into the filling module and the intent module in the recognition model to be trained for processing, and the relationship weight between the filling module and the intent module is determined according to the processing result, and the relationship weight specifically refers to the weight of the mutual influence between the two.

[0107] In specific implementation, the relationship weight of the filling module relative to the intention module, as well as the relationship weight of the intention module relative to the filling module, can be determined according to the processing results, so that when the intention module and the filling module are trained, the effect of joint learning can be achieved, thereby improving the accuracy of the recognition model to be trained.

[0108] Furthermore, in the process of determining the relationship weight between the filling module and the intent module, in order to effectively improve the recognition effect of the recognition model to be trained, it is necessary to simultaneously consider the influence of the intent module on the filling module, and the influence of the filling module on the intent module, so as to simultaneously improve the intent understanding effect of the intent module and the entity extraction effect of the filling module through checks and balances. In this embodiment, the specific implementation method is as follows:

[0109] Inputting the feature vector into the filling module in the recognition model to be trained for processing to obtain a first intermediate vector, and inputting the feature vector into the intention module in the recognition model to be trained for processing to obtain a second intermediate vector;

[0110] The relationship weight between the filling module and the intention module is calculated based on the first intermediate vector and the second intermediate vector.

[0111] Specifically, the filling module plays a role in the recognition model to be trained to extract entities in the text. The extraction form is to fill in the entity label, that is, to add slot filling labels to the entity in the text. In this process, the filling module actually learns the weight of the current hidden layer unit through each hidden layer unit. The first intermediate vector corresponding to the slot context at each moment is obtained by weighted summation. The first intermediate vector can be obtained by formula (1):

[0112]

[0113] In formula (1) represents the first intermediate vector, represents the weight of the current hidden unit in the filling module, Si represents the feature vector corresponding to the i-th word in the text; among them, the weight It can be obtained by the following formula (2) and formula (3):

[0114]

[0115]

[0116] Among them, σ represents the activation function, represents the W matrix of a forward network, h k represents the hidden layer state, e i,j represents the intermediate calculation weight; based on this, usually, the filling module can finally be Extract entities from text ( is the slot label of the i-th word in the input, is a weight matrix), and is annotated by filling labels. However, since the recognition model to be trained also includes an intent module, if the filling module needs to achieve a better output effect, the influence of the intent module on the filling module needs to be considered. This embodiment determines the influence of the intent module on the filling module and the influence of the filling module on the intent module by joint learning, thereby improving the recognition of intent understanding and the accuracy of slot filling at the same time. The specific implementation is as follows:

[0117] First, we need to determine the second intermediate vector c of the intention module I , wherein the calculation process of the second intermediate vector can refer to the calculation process of the first intermediate vector, and this embodiment will not be described in detail here. Next, the second intermediate vector c I Added to the filling module, in order to reduce the influence of the intention module on the filling module, and can improve the prediction ability of the filling module, in order to achieve such an effect, finally through the second intermediate vector c I The relationship weight g of the intention module relative to the filling module is calculated, that is, by introducing the relationship weight g of the intention module on the filling module into the filling module, the filling module can take into account the influence of the intention module while improving its own prediction ability; the determination process of the relationship weight g introduced into the filling module is as follows Figure 2 As shown in (a) in the figure, it can be obtained by formula (4):

[0118]

[0119] Among them, v and W are trainable vectors and matrices, c Irepresents the second intermediate vector, represents the first intermediate vector, g represents the relationship weight, which can be understood as the first intermediate vector And the weighted features of the second intermediate vector, so as to introduce the relationship weight g of the intention module into the filling module to improve the prediction ability of the filling module, so that it can take into account the influence of the intention module, and further improve the ability of the recognition model to be trained.

[0120] Similarly, when determining the impact of the intent module on the filling module, it is also necessary to determine the impact of the filling module on the intent module, so as to simultaneously improve the recognition of intent understanding and the accuracy of slot filling. The specific implementation is as follows:

[0121] First, we need to determine the first intermediate vector of the filling module Secondly, the first intermediate vector is added to the intent module to reduce the impact of the filling module on the intent module and effectively improve the prediction ability of the intent module. In order to achieve such an effect, the first intermediate vector is finally The relationship weight l of the filling module relative to the intention module is calculated, that is, by introducing the relationship weight l of the filling module to the intention module into the intention module, so that the intention module can take into account the influence of the intention module while improving its own prediction ability; the determination process of the relationship weight l introduced into the intention module is as follows Figure 2 As shown in (b), it can be obtained by formula (5):

[0122]

[0123] Among them, v and W are trainable vectors and matrices, c I represents the second intermediate vector, represents the first intermediate vector, l represents the relationship weight, which can be understood as the first intermediate vector And the weighted features of the second intermediate vector, so as to introduce the relationship weight l of the filling module into the intention module to improve the prediction ability of the intention module, so that it can take into account the influence of the filling module, and further improve the ability of the recognition model to be trained.

[0124] In practical applications, the relationship weight calculation process of the intent module and the filling module is carried out simultaneously to achieve the purpose of joint learning and maintain the same progress, which can not only improve the coordination ability of the filling module and the intent module, but also effectively improve the recognition ability of the recognition model to be trained.

[0125] In summary, in order to effectively improve the recognition ability of the recognition model to be trained, the filling module and the intent module in the model are jointly learned to check and balance each other, which can not only improve the entity extraction ability of the model but also effectively improve the intent understanding ability, so that the recognition model can be applied to more scenarios.

[0126] Step S108, adjusting the filling module and the intention module based on the relationship weight, and obtaining a first loss value of the filling module and a second loss value of the intention module according to the adjustment result.

[0127] Specifically, on the basis of the above-mentioned determination of the relationship weight between the filling module and the intention module, the filling module and the intention module will be further adjusted simultaneously according to the relationship weight, so that the filling module and the intention module can check and balance each other, and fully combine with each other to obtain a recognition model with better recognition effect. After the adjustment of the filling module and the intention module is completed, it is necessary to continue to train the recognition model to be trained. In order to obtain a model with better training effect, the first loss value of the filling module and the second loss value of the intention module can be obtained for subsequent iterative training of the recognition model to be trained according to the loss value.

[0128] In specific implementation, since the filling module and the intention module are carried out through joint learning, in the subsequent process of continuing to train the recognition model to be trained, it is necessary to simultaneously consider the first loss value of the filling module and the second loss value of the intention module, so that when training the model, the first loss value and the second loss value are simultaneously reduced to an equilibrium position, thereby training a recognition model that meets the recognition requirements. In actual applications, reducing the first loss value and the second loss value to an equilibrium position at the same time means that after the first loss value is reduced, the second loss value will not rise, and after the second loss value is reduced, the first loss value will not rise, until the sum of the two reaches a minimum value, and the equilibrium position can be determined.

[0129] Furthermore, in the process of adjusting the filling module and the intention module according to the relationship weight, since the filling module can complete entity labeling independently and the intention module can complete intention understanding independently, and the recognition model to be trained includes both the filling module and the intention module, in order to enable the recognition model to be trained to satisfy entity labeling while accurately understanding the intention of the text, it is necessary to introduce the influence of the filling module into the intention module, and introduce the influence of the intention module into the filling module, so as to achieve the purpose of common progress. In this embodiment, the specific implementation method is as follows:

[0130] Obtaining a first text vector obtained by the filling module processing the feature vector, and a second text vector obtained by the intention module processing the feature vector;

[0131] Calculating the product of the relationship weight and the first text vector to obtain a first target text vector, and calculating the product of the relationship weight and the second text vector to obtain a second target text vector;

[0132] Extracting the entity vector and the intention vector corresponding to the training text in the training set to which the training text belongs;

[0133] The first loss value of the filling module is determined according to the first target text vector and the entity vector, and the second loss value of the intent module is determined according to the second target text vector and the intent vector.

[0134] Specifically, the first text vector specifically refers to the vector obtained after the filling module processes the feature vector, and specifically represents the vector expression after the filling module annotates the entities in the training text. The second text vector specifically refers to the vector obtained after the intention module processes the feature vector, and specifically represents the vector expression after the intention module understands the intention of the training text. At this time, the relationship weight of the intention module relative to the filling module is introduced into the first text vector, so as to obtain the entity annotation result after the filling module is affected by the intention module, and express it in the form of the first target text vector. At the same time, the relationship weight of the filling module relative to the intention module is introduced into the second text vector, so as to obtain the intention understanding result after the intention module is affected by the filling module, and express it in the form of the second target text vector.

[0135] In practical applications, when the filling module completes entity annotation alone, To achieve, It represents the result of entity annotation for the i-th word in the text (other combination letters and Chinese characters can refer to the same description in the above embodiment). In the recognition model to be trained, it is necessary to introduce the influence of the intention module when the filling module performs entity annotation in order to meet the recognition requirements of the recognition model to be trained. Therefore, the filling module is adjusted according to the relationship weight g of the intention module relative to the filling module. The adjusted filling module can be implemented by the following formula (6) when performing entity recognition:

[0136]

[0137] Accordingly, when the intention module completes the intention understanding alone, it can Implementation, where yI It represents the result of intent understanding after the intent understanding of the text. In the recognition model to be trained, it is necessary to introduce the influence of the filling module when the intent module understands the intent, so as to meet the recognition requirements of the recognition model to be trained. Therefore, the intent module is selected according to the relationship weight l of the filling module. The adjusted adjustment module can be implemented by the following formula (7) when performing intent understanding:

[0138]

[0139] Based on this, after the output results of the filling module and the intention module are adjusted, the recognition model to be trained will need to be subsequently deeply trained according to the loss value. Prior to this, the first loss value and the second loss value need to be determined, and the calculation of the loss value needs to be combined with the entity vector and intention vector of the training text. The entity vector specifically refers to the vector expression corresponding to the real entity of the training text, and the intention vector specifically refers to the vector expression corresponding to the real intention of the training text. The first target text vector refers to the vector expression corresponding to the entity predicted by the filling module, and the second target text vector refers to the vector expression corresponding to the text intention predicted by the intention module.

[0140] Furthermore, a first loss value corresponding to the filling module can be determined based on the entity vector and the first target text vector, and a second loss value corresponding to the intention module can be determined based on the intention vector and the second target text vector, for subsequent iterative training of the model.

[0141] Using the above example, we determine the feature vector Sn = [S 11 , S 12 , S 13 , S 14 , S 15 , S 16 , S 17 , S 18 , S 19 ], the relationship weight g of the filling module introduced into the recognition model to be trained and the relationship weight l of the intention module introduced are determined. Further, the feature vector Si is input into the filling module, and the relationship weight g is introduced to obtain the first target text vector So1=[S 21 , S 22 , S 23 , S 24 , S 25 , S 26 , S 27 , S 28 , S29 ], which is used to express the annotation results of each entity in the training text in the form of a vector; at the same time, the feature vector Sn is input into the intention module, and the relationship weight l is introduced to obtain the second target text vector St1 output by the filling module = [S 31 , S 32 , S 33 , S 34 , S 35 , S 36 , S 37 , S 38 , S 39 ] is used to express the intent understanding result after the intent understanding of the training text, expressed in the form of a vector.

[0142] In order to further train the recognition model to be trained, the entity vector and intent vector corresponding to the training text will be extracted from the training set to which the training text belongs, that is, the vector expression of the real entity corresponding to the training text is So2=[S 221 , S 222 , S 223 , S 224 , S 225 , S 226 , S 227 , S 228 , S 229 ], the vector expression of the true intention St2 = [S 331 , S 332 , S 333 , S 334 , S 335 , S 336 , S 337 , S 338 , S 339 ], and then determine the first loss value of the filling module through So1 and So2, and determine the second loss value of the intent module according to St1 and St2, so as to continue training the recognition model to be trained.

[0143] In actual applications, the first loss value represents the degree of difference between the predicted result output by the filling module and the actual result, and the second loss value represents the degree of difference between the predicted result output by the intention module and the actual result, so as to decide whether to continue training the recognition model to be trained. The larger the loss value, the worse the model output result, and vice versa. The smaller the loss value, the better the model output result.

[0144] In summary, in the process of training the recognition model to be trained, the method of joint learning of the filling module and the intention module can effectively improve the degree of coordination between the filling module and the intention module, further improving the intention understanding ability and slot filling ability of the recognition model to be trained.

[0145] Step S110, iteratively training the recognition model to be trained based on the first loss value and the second loss value until a training stop condition is reached to obtain a target recognition model.

[0146] Specifically, on the basis of obtaining the first loss value and the second loss value as mentioned above, in order to enable the recognition model to be trained to achieve a high level of slot filling and intent understanding after training, the recognition model to be trained can be iteratively trained based on the first loss value and the second loss value until the recognition model to be trained reaches the training stop condition, so as to obtain a target recognition model that meets the recognition requirements; wherein, the training stop condition specifically refers to that the first loss value and the second loss value meet the training requirements for stopping training the recognition model to be trained, and the target recognition model specifically refers to a recognition model that meets the recognition requirements, which can achieve good entity labeling effects and intent understanding effects.

[0147] Furthermore, in the process of training the recognition model to be trained based on the first loss value and the second loss value, in order to enable the model to achieve a better prediction effect, iterative training can be performed by calculating the target loss value. In this embodiment, the specific implementation method is as follows:

[0148] Calculate the target loss value of the recognition model to be trained according to the first loss value and the second loss value;

[0149] The recognition model to be trained is iteratively trained according to the target loss value.

[0150] Specifically, the target loss value will be calculated based on the first loss value and the second loss value. The calculation method can be through a weighted sum method or a summation method. This embodiment does not make too many restrictions here. Furthermore, after determining the target loss value, the recognition model to be trained can be iteratively trained according to the target loss value until the target loss value is less than or equal to the loss value threshold, so as to obtain a target recognition model that meets the recognition requirements. In addition, the loss value threshold is determined based on the various target loss values ​​generated by the recognition model to be trained during the training process, that is, the loss value threshold will be set according to actual needs. The setting method is to collect the various target loss values ​​generated by the recognition model to be trained during the training process, and then determine them through manual screening, that is, the loss value corresponding to the above-mentioned equilibrium position can be used as the loss value threshold.

[0151] Furthermore, in this embodiment, the target loss value of the recognition model to be trained is calculated according to the first loss value and the second loss value, and the specific implementation method is as follows:

[0152] Determining a weight value of the first loss value and a weight value of the second loss value;

[0153] A weighted summation process is performed on the weight value of the first loss value and the weight value of the second loss value to obtain the target loss value.

[0154] Continuing with the above example, when the first loss value loss1=30% of the filling module is determined based on So1 and So2, and the second loss value loss2=60% of the intent module is determined based on St1 and St2, it is necessary to determine whether the recognition model to be trained at this time meets the conditions for stopping training. According to the weight value P1 of the first loss value loss1, the weight value P2 of the second loss value, and the first loss value loss1=30% and the second loss value loss2=60%, after weighted and calculation, it is determined that the target loss value is 45%, which does not meet the conditions for stopping training. It is necessary to continue training the recognition model to be trained. Then, a large number of training samples continue to be selected in the training set, and the model continues to be trained in the above manner until the target loss value is less than the loss value threshold. It can be determined that the filling module and the intent module in the recognition model to be trained have achieved relatively good prediction capabilities, and the target recognition model can be obtained for use in actual application scenarios.

[0155] In summary, by using the weighted sum calculation method to determine the target loss value, the prediction capabilities of the intent module and the filling module can be further balanced, thereby improving the recognition effect of the target recognition model and meeting the needs of more application scenarios.

[0156] The training method of the recognition model provided in the present application, in the process of training the recognition model to be trained, uses a language module to encode the training text to obtain a feature vector, then processes the feature vector according to the filling module and the intention module in the recognition model to be trained, and determines the relationship weight between the two according to the processing result, then adjusts the filling module and the intention module according to the relationship weight, and obtains the first loss value of the filling module and the second loss value of the intention module, and finally iteratively trains the recognition model to be trained according to the first loss value and the second loss value, so as to improve the accuracy of the recognition model to be trained by joint learning of the filling module and the intention module, and by combining the relationship weight, the filling module and the intention module assist each other in semantic understanding, thereby further improving the analysis accuracy of intent understanding and slot filling, and training the model in parallel can effectively improve the model training speed.

[0157] The following is an embodiment of the text recognition method provided by this application:

[0158] Figure 3 is a flowchart of a text recognition method provided by an embodiment of the present application. Figure 4 is a schematic diagram of a text recognition method provided by an embodiment of the present application; wherein Figure 3 The specific steps include:

[0159] Step S302, obtaining the text to be recognized.

[0160] Step S304: input the text to be recognized into a target recognition model for entity extraction and intent understanding, and obtain the target entity and target intent corresponding to the text to be recognized.

[0161] The target recognition model can be obtained by training using the above-mentioned recognition model training method.

[0162] Specifically, the text to be recognized refers to text that requires entity annotation and intent understanding.

[0163] Furthermore, in the process of performing intent understanding and entity labeling through the target recognition model, in order to achieve more accurate intent understanding and entity labeling, the target recognition model will extract and label entities through the mutually balanced filling module and intent module in the model, and recognize the intent of the text. In this embodiment, the specific implementation method is as follows:

[0164] Inputting the text to be recognized into the target recognition model, encoding the text to be recognized through the language module in the target recognition model to obtain a text feature vector;

[0165] The text feature vector is input into the filling module and the intention module in the target recognition model for processing, and the target entity and the target intention are obtained according to the processing results.

[0166] Specifically, after the text to be recognized is input into the target recognition model, the language module in the target recognition model will encode the text to be recognized to obtain the text feature vector of the text to be recognized, and then pass it through the intention module and filling module in the target recognition model to determine the target intention and target entity of the text to be recognized.

[0167] Furthermore, in order to improve the recognition efficiency of the target recognition model, the text to be recognized may be segmented before being input into the target recognition model. In this embodiment, the specific implementation is as follows:

[0168] Performing word segmentation processing on the text to be recognized to obtain a word unit set;

[0169] The word unit set is input into the language module for encoding processing to obtain the text feature vector.

[0170] Specifically, after obtaining the text to be recognized, the text to be recognized can be segmented, the text to be recognized can be processed into multiple word units, and the multiple word units can be combined into the word unit set. In the process of model recognition, the word unit set can be input into the target recognition model. The word unit set can be encoded and processed into a text feature vector corresponding to the text to be recognized through the language module in the target recognition model, so as to be used for subsequent intention understanding and entity labeling.

[0171] In addition, in the process of identifying the text feature vector, the filling module and the intention module will perform calculation processing in a mutually balanced manner in order to output a more accurate target intention and target entity. In this embodiment, the specific implementation method is as follows:

[0172] Inputting the text feature vector into the filling module for processing to obtain a first intermediate vector, and inputting the text feature vector into the intention module for processing to obtain a second intermediate vector;

[0173] Calculate the product of the first intermediate vector and the first relationship weight of the intention module to obtain a first text vector, and calculate the product of the second intermediate vector and the second relationship weight of the filling module to obtain a second text vector;

[0174] The first text vector is converted through the output layer of the filling module to obtain the target entity, and the second text vector is converted through the output layer of the intention module to obtain the target intention.

[0175] Specifically, after the text to be recognized is encoded by the language module in the target recognition model, a text feature vector corresponding to the text to be recognized is obtained, and then the text feature vector is input into the filling module for processing to obtain a first intermediate vector, and the text feature vector is input into the intention module for processing to obtain a second intermediate vector; the first intermediate vector specifically refers to the relationship weight that has not yet been introduced into the intention module, and the entity labeling result output by the filling module; the second intermediate vector specifically refers to the relationship weight that has not yet been introduced into the filling module, and the intention understanding result output by the intention module.

[0176] Based on this, the product of the first intermediate vector and the first relationship weight of the intention module is calculated to obtain the first text vector, and the product of the second intermediate vector and the second relationship weight of the filling module is calculated to obtain the second text vector; the first text vector and the second text vector specifically refer to the results output after the filling module and the intention module assist each other, and are expressed in the form of vectors. Finally, the first text vector is converted through the output layer of the filling module to obtain the target entity, and the second text vector is converted through the output layer of the intention module to obtain the target intention.

[0177] In practical applications, the target entity specifically refers to the content with a relatively high degree of importance in the text to be recognized, and the target intent specifically refers to the meaning to be expressed by the text to be recognized.

[0178] See also Figure 4 As shown in the figure, for example, if the user says "I want to go to City A to watch the flag-raising ceremony", the target recognition model can be used to identify the text and determine that the intention of "I want to go to City A to watch the flag-raising ceremony" is "information type". The named entities in this text include {City A, National Flag}, which can achieve faster determination of the user's expression intention and improve the structuring effect.

[0179] The specific implementation method includes: after obtaining the text to be recognized "I want to go to City A to watch the flag raising", the text to be recognized is input into the target recognition model for entity extraction and intent understanding, firstly, the text to be recognized "I want to go to City A to watch the flag raising" is segmented to obtain a word unit set consisting of "I", "I want to", "go to", "City A", "watch", "raise", and "national flag", and then input into the target recognition model, and the BERT model as a language module in the target recognition model encodes the word unit set consisting of "I", "I want to", "go to", "City A", "watch", "raise", and "national flag", and obtains the text feature vector H corresponding to "I want to go to City A to watch the flag raising";

[0180] Secondly, the text feature vector H is simultaneously input into the filling module and the intention module in the target recognition model. The filling module and the intention module introduce relationship weights during the training process to achieve joint learning. Therefore, when performing entity recognition and intention recognition, the output results can be mutually influenced. Specifically, the filling module uses formula (6) The entity annotation result vector [n] = 20*5 affected by the intention module can be obtained. According to the entity annotation result vector [n] = 20*5, the entity recognition result corresponding to the text to be recognized "I want to go to City A to watch the flag raising" can be obtained. The entity recognition result is shown in Table 1:

[0181] I think go A city look Lift National flag H1 H2 H3 H4 H5 H6 H7 O O O B O O I

[0182] Table 1

[0183] Among them, B represents the starting position of a named entity, I represents the non-starting position of a named entity, and O represents a non-named entity. At the same time, the intention module uses formula (7) The intention recognition result vector [d]=1*10 affected by the filling module can be obtained. According to the intention recognition result vector [d]=1*10, the intention understanding result corresponding to the text to be recognized "I want to go to City A to watch the flag-raising ceremony" can be obtained: information class, thereby understanding the intention of the text to be recognized and extracting the entities and entity relationships in the text to be recognized.

[0184] In practical applications, since the functions of the filling module and the intention module are different, after the text is encoded by the language module, the obtained text feature vector can be converted according to the input of the filling module and the intention module. For example, the output of the BERT model is a 10*10 vector, while the input of the filling module is a 20*5 vector, and the input of the intention module is a 1*10 vector. At this time, the 10*10 vector can be converted to obtain the input that meets the filling module input and the intention module input. And since the output of the target recognition model is the output of the integrated intention module and the filling module, it can be an integrated vector of the 20*5 output of the filling module and the 1*10 output of the intention module.

[0185] In addition, after obtaining the target intent and the target entity, in order to be more convenient for application in actual scenarios, a target template can be selected according to the target intent to generate a target text that meets the scenario requirements. In this embodiment, the specific implementation method is as follows:

[0186] A target template is selected from preset filling templates according to the target intent.

[0187] The target entity is structurally processed according to the filling rule of the target template and the target template to obtain the target text.

[0188] Specifically, the preset filling template refers to a template set according to the actual application scenario. Different application scenarios may require the use of different templates. For example, in a navigation scenario, the selected template needs to include an image to better express the meaning of the text to be recognized, or in an information reading scenario, the selected template needs to include a table to better express the meaning of the text to be recognized, or in a data analysis scenario, the selected template needs to include a statistical chart to better express the meaning of the text to be recognized. In actual applications, the preset filling template can be set according to the actual application scenario, and this embodiment does not make too many restrictions here.

[0189] For example, when a user visits a hospital for treatment, he or she may describe his or her symptoms and medication in detail to medical staff. The hospital can then use the target recognition model to identify that the user's intention is to treat the disease, and the entities include {heart disease, Class A drug, hospitalization, Class A hospital}. Then, based on the user's intention to treat the disease, the hospital selects a medical record template from the preset filling templates, and fills in the entities {heart disease, Class A drug, hospitalization, Class A hospital} into the medical record template according to the filling rules of the medical record template, so that hospital staff can directly obtain the medical record text, saving medical staff from excessive operating steps.

[0190] In summary, by selecting a target template to generate a target text according to the target intent, users can obtain the required text without manual filling, which greatly improves the user experience. For texts with more content, the target recognition model can be used to extract key entities and understand the meaning of the text in a shorter time, effectively saving the time spent on reading the text.

[0191] The text recognition method provided in the present application can effectively improve the recognition accuracy by adopting a target recognition model to perform entity extraction and intent understanding on the text to be recognized, combined with the checks and balances mechanism between the filling module and the intent module in the target recognition model. The target recognition model uses a language module combined with a joint learning method of intent understanding and slot filling to extract target entities and relationships, further improving the natural language structuring effect.

[0192] The following combination Figure 5 Taking the application of the recognition model training method and text recognition method provided by the present application in a text recognition scenario as an example, the recognition model training method and text recognition method are further described. Figure 5 A processing flow chart for a text recognition scenario provided by an embodiment of the present application is shown, which specifically includes the following steps:

[0193] Step S502, obtaining training text.

[0194] Specifically, the training text is the text that needs to train the recognition model to be trained, such as a sentence, a paragraph, or an article. In this embodiment, "Frist class fares from Boston to Denver" is used as the training text to describe the model training process of this solution in detail.

[0195] Step S504: perform word segmentation processing on the training text to obtain a word unit set.

[0196] Step S506: input the word unit set into the language module in the recognition model to be trained for encoding processing to obtain the text feature vector of the training text.

[0197] Specifically, the recognition model to be trained is a model that can achieve intent understanding and slot filling. The recognition model to be trained integrates a language model (language module, used for encoding text), a filling module (for slot filling) and an intent module (for intent understanding).

[0198] Based on this, the training text "Frist class fares from Boston to Denver" is segmented to obtain a word unit set consisting of "Frist", "class", "fares", "from", "Boston", "to" and "Denver". The BERT model with better results is then used as a language model to encode the word unit set to obtain the text feature vector Ai=Ai=[A 11 , A 12 , A 13 , A 14 , A 15 , A 16 , A 17 ].

[0199] Step S508: input the text feature vector into a filling module in the recognition model to be trained for processing to obtain a first intermediate vector.

[0200] Step S510: input the text feature vector into the intention module in the recognition model to be trained for processing to obtain a second intermediate vector.

[0201] In actual applications, there is no order in which the above steps S508 and S510 are executed, and the two can be executed simultaneously to train the recognition model to be trained.

[0202] Specifically, after obtaining the text feature vector Ai=[A 11 , A 12 , A 13 , A 14 , A 15 , A 16 , A 17 ], the above vector can be input into the filling module and intent module in the recognition model to be trained, and entity extraction and intent understanding can be performed at the same time.

[0203] Based on this, on the one hand, in the process of obtaining the slot filling label, the filling module actually learns the weight of the filling module to the current hidden layer unit for each hidden layer unit in the filling module, and obtains the first intermediate vector corresponding to the slot context at each moment through weighted summation, where the first intermediate vector is the annotation result of the filling module, as shown in Table 2:

[0204] Frist class fares from Boston to Denver B-Cabin Category I-Cabin Category O O B-Departure Place O Destination B

[0205] Table 2

[0206] On the other hand, in the process of understanding the intention, the second intermediate vector specifically refers to the intention understanding result, which can determine the intention type of the text to be recognized, and it can be concluded that Ai = [A 11 , A 12 , A 13 , A 14 , A 15 , A 16 , A 17 ]The corresponding training text intent is travel intent.

[0207] Step S512: Calculate the relationship weight between the filling module and the intention module based on the first intermediate vector and the second intermediate vector.

[0208] Specifically, in order to improve the recognition accuracy of the recognition model to be trained, it will be achieved through the joint learning of the filling module and the intention module. Specifically, the prediction ability of the filling module will be improved through the context vector of the intention (the first intermediate vector). In order to achieve such an effect, it will be achieved through the relationship weight g. At the same time, the prediction ability of the intent module can also be improved through the context vector of the slot filling (the second intermediate vector). In order to achieve such an effect, it will be achieved through the relationship weight l.

[0209] Step S514, obtaining a first text vector obtained after the filling module processes the text feature vector, and a second text vector obtained after the intention module processes the text feature vector.

[0210] Step S516, calculating the product of the relationship weight and the first text vector to obtain a first target text vector, and calculating the product of the relationship weight and the second text vector to obtain a second target text vector.

[0211] Step S518, extracting the entity vector and intention vector corresponding to the training text.

[0212] Step S520, determining a first loss value of the filling module according to the first target text vector and the entity vector, and determining a second loss value of the intent module according to the second target text vector and the intent vector.

[0213] Step S522, determining a target loss value of the recognition model to be trained according to the first loss value and the second loss value, and iteratively training the recognition model to be trained according to the target loss value.

[0214] Step S524, when the target loss value is less than or equal to the preset loss value threshold, a target recognition model that meets the recognition requirements is obtained.

[0215] Specifically, based on the above-mentioned determination of the relationship weight between the filling module and the intention module, it is necessary to adjust the filling module and the intention module according to the relationship weight to achieve the purpose of joint learning. The specific implementation process of the filling module is: At this time, the relationship weight g is determined, and Ai=[A 11 , A 12 , A 13 , A 14 , A 15 , A 16 , A 17 ] The label of each word in the corresponding training texts "Frist", "class", "fares", "from", "Boston", "to" and "Denver" (the first text vector), the influence of the intention module on the filling module (relationship weight) can be introduced into the label determination result. Correspondingly, the implementation process of the intention module is similar to it, and the relationship weight l is introduced into the intention understanding (second text vector) result, so that the filling module combines the influence of the intention module on it (relationship weight).

[0216] Based on this, after the filling module and the intention module are adjusted based on the relationship weight, Ai=[A 11 , A 12 , A 13 , A 14 , A 15 , A 16 , A 17 ] corresponding to the training text "Frist", "class", "fares", "from", "Boston", "to" and "Denver", and the intent module is used to fill in the labels of each word in Ai = [A 11 , A 12 , A 13 , A 14 , A 15 , A 16 , A 17 ]The corresponding training texts "Frist", "class", "fares", "from", "Boston", "to", and "Denver" are used for intent understanding to obtain the label of each word (the first target text vector) and the intent of the training text (the second target text vector).

[0217] Furthermore, in order to train a recognition model that meets recognition requirements, the first loss value (loss1) of the filling module can be obtained based on the first target text vector output by the filling module and the entity vector (real entity vector) corresponding to the training text, and the second loss value (loss2) of the intent module can be obtained based on the second target text vector output by the intent module and the intent vector (real intent vector) corresponding to the training text, for subsequent iterative training of the model.

[0218] After obtaining the first loss value (loss1) and the second loss value (loss2), in order to enable the recognition model to be trained to achieve a high level of slot filling and intent understanding accuracy, iterative training can be performed by calculating the target loss value (total loss) of the first loss value (loss1) and the second loss value (loss2). That is, when the total loss reaches the minimum value, the training can be stopped, and the recognition model to be trained can be used as the target recognition model.

[0219] Based on this, if "I bought a car" is input into the target recognition model, the intent is understood as a purchasing intent, and the recognition results with slots marked as I(O), bought(O), a(O), and car(I) can be obtained.

[0220] In summary, the accuracy of the recognition model to be trained can be improved by joint learning of the filling module and the intent module. By combining the relationship weights, the filling module and the intent module assist each other in semantic understanding, further improving the analysis accuracy of intent understanding and slot filling. Training the model in parallel can effectively improve the model training speed. In the process of use, by using the target recognition model to extract entities and understand intent of the text to be recognized, combined with the checks and balances mechanism of the filling module and the intent module in the target recognition model, the recognition accuracy can be effectively improved. The target recognition model uses the language module combined with the joint learning method of intent understanding and slot filling to extract target entities and relationships, further improving the natural language structuring effect.

[0221] Corresponding to the above method embodiment, the present application also provides an embodiment of a training device for a recognition model, Figure 6 FIG. 1 is a schematic diagram showing a structure of a training device for a recognition model provided by an embodiment of the present application. Figure 6 As shown, the device comprises:

[0222] An acquisition unit 602 is configured to acquire a training text;

[0223] The encoding unit 604 is configured to input the training text into the recognition model to be trained, and perform encoding processing on the training text through the language module in the recognition model to be trained to obtain a feature vector;

[0224] A determination unit 606 is configured to input the feature vector into a filling module and an intention module in the recognition model to be trained for processing, and determine a relationship weight between the filling module and the intention module according to the processing result;

[0225] An adjusting unit 608 is configured to adjust the filling module and the intention module based on the relationship weight, and obtain a first loss value of the filling module and a second loss value of the intention module according to the adjustment result;

[0226] The training unit 610 is configured to iteratively train the recognition model to be trained based on the first loss value and the second loss value until a training stop condition is reached to obtain a target recognition model.

[0227] In an optional embodiment, the determining unit 606 includes:

[0228] a processing subunit configured to input the feature vector into the filling module in the recognition model to be trained for processing to obtain a first intermediate vector, and to input the feature vector into the intention module in the recognition model to be trained for processing to obtain a second intermediate vector;

[0229] The relationship weight calculation subunit is configured to calculate the relationship weight between the filling module and the intention module based on the first intermediate vector and the second intermediate vector.

[0230] In an optional embodiment, the adjusting unit 608 includes:

[0231] A text vector acquisition subunit is configured to acquire a first text vector obtained by the filling module processing the feature vector, and a second text vector obtained by the intention module processing the feature vector;

[0232] a product calculation subunit configured to calculate the product of the relationship weight and the first text vector to obtain a first target text vector, and to calculate the product of the relationship weight and the second text vector to obtain a second target text vector;

[0233] A vector extraction subunit is configured to extract an entity vector and an intention vector corresponding to the training text in a training set to which the training text belongs;

[0234] A loss value determination subunit is configured to determine the first loss value of the filling module based on the first target text vector and the entity vector, and to determine the second loss value of the intent module based on the second target text vector and the intent vector.

[0235] In an optional embodiment, the training unit 610 includes:

[0236] A target loss value calculation subunit is configured to calculate a target loss value of the recognition model to be trained according to the first loss value and the second loss value;

[0237] The iterative training subunit is configured to iteratively train the recognition model to be trained according to the target loss value.

[0238] In an optional embodiment, the training stop condition is that the target loss value is less than or equal to a loss value threshold;

[0239] Correspondingly, if the target loss value is less than or equal to the loss value threshold, the target recognition model is obtained;

[0240] The loss value threshold is determined based on each target loss value generated by the recognition model to be trained during the training process.

[0241] In an optional embodiment, the target loss value calculation subunit includes:

[0242] A weight value determination submodule, configured to determine a weight value of the first loss value and a weight value of the second loss value;

[0243] The weighted sum submodule is configured to perform weighted sum processing on the weight value of the first loss value and the weight value of the second loss value to obtain the target loss value.

[0244] In an optional embodiment, the training device for the recognition model further includes:

[0245] A word segmentation processing unit is configured to perform word segmentation processing on the training text to obtain a word unit set;

[0246] Accordingly, the encoding unit 604 is further configured as follows:

[0247] The word unit set is input into the recognition model to be trained, and the word unit set is encoded by the language module to obtain the feature vector.

[0248] The training device of the recognition model provided in the present application, in the process of training the recognition model to be trained, uses a language module to encode the training text to obtain a feature vector, then processes the feature vector according to the filling module and the intention module in the recognition model to be trained, and determines the relationship weight between the two according to the processing result, then adjusts the filling module and the intention module according to the relationship weight, and obtains the first loss value of the filling module and the second loss value of the intention module, and finally iteratively trains the recognition model to be trained according to the first loss value and the second loss value, so as to improve the accuracy of the recognition model to be trained by joint learning of the filling module and the intention module, and by combining the relationship weight, the filling module and the intention module assist each other in semantic understanding, thereby further improving the analysis accuracy of intent understanding and slot filling, and training the model in parallel can effectively improve the model training speed.

[0249] The above is a schematic scheme of a training device for a recognition model of this embodiment. It should be noted that the technical scheme of the training device for the recognition model and the technical scheme of the training method for the recognition model described above belong to the same concept, and the details of the technical scheme of the training device for the recognition model that are not described in detail can be found in the description of the technical scheme of the training method for the recognition model described above.

[0250] In addition, the components in the device claim should be understood as the functional modules that must be established to implement the steps of the program flow or the steps of the method, and the functional modules are not actually functionally divided or separated. The device claim defined by such a group of functional modules should be understood as the functional module architecture of the solution mainly implemented by the computer program recorded in the specification, and should not be understood as the physical device that mainly implements the solution by hardware.

[0251] Corresponding to the above method embodiment, the present application also provides a text recognition device embodiment, Figure 7 FIG. 1 is a schematic diagram showing the structure of a text recognition device provided by an embodiment of the present application. Figure 7 As shown, the device comprises:

[0252] A text acquisition unit 702 is configured to acquire text to be recognized;

[0253] The model processing unit 704 is configured to input the text to be recognized into the target recognition model for entity extraction and intent understanding, and obtain the target entity and target intent corresponding to the text to be recognized.

[0254] Wherein, the target recognition model is obtained by training through the above-mentioned recognition model training method.

[0255] In an optional embodiment, the text recognition device further includes:

[0256] A template selection unit is configured to select a target template from preset filling templates according to the target intention;

[0257] The structural processing unit is configured to perform structural processing on the target entity according to the filling rule of the target template and the target template to obtain the target text.

[0258] In an optional embodiment, the model processing unit 704 includes:

[0259] The encoding subunit is configured to input the text to be recognized into the target recognition model, and perform encoding processing on the text to be recognized through the language module in the target recognition model to obtain a text feature vector;

[0260] The processing subunit is configured to input the text feature vector into the filling module and the intention module in the target recognition model for processing, and obtain the target entity and the target intention according to the processing results.

[0261] In an optional embodiment, the text recognition device further includes:

[0262] A word segmentation unit is configured to perform word segmentation processing on the text to be recognized to obtain a word unit set;

[0263] Accordingly, the model processing unit 704 is further configured to:

[0264] The word unit set is input into the language module for encoding processing to obtain the text feature vector.

[0265] In an optional embodiment, the processing subunit includes:

[0266] a processing submodule, configured to input the text feature vector into the filling module for processing to obtain a first intermediate vector, and input the text feature vector into the intention module for processing to obtain a second intermediate vector;

[0267] a calculation submodule configured to calculate the product of the first intermediate vector and the first relationship weight of the intention module to obtain a first text vector, and to calculate the product of the second intermediate vector and the second relationship weight of the filling module to obtain a second text vector;

[0268] The conversion submodule is configured to convert the first text vector through the output layer of the filling module to obtain the target entity, and to convert the second text vector through the output layer of the intention module to obtain the target intention.

[0269] The text recognition device and the text recognition method provided in the present application can effectively improve the recognition accuracy by adopting a target recognition model to perform entity extraction and intent understanding on the text to be recognized, combined with the checks and balances mechanism between the filling module and the intent module in the target recognition model. The target recognition model uses a language module combined with a joint learning method of intent understanding and slot filling to extract target entities and relationships, further improving the natural language structuring effect.

[0270] The above is a schematic scheme of a text recognition device of this embodiment. It should be noted that the technical scheme of the text recognition device and the technical scheme of the above text recognition method belong to the same concept, and the details not described in detail in the technical scheme of the text recognition device can be referred to the description of the technical scheme of the above text recognition method.

[0271] In addition, the components in the device claim should be understood as the functional modules that must be established to implement the steps of the program flow or the steps of the method, and the functional modules are not actually functionally divided or separated. The device claim defined by such a group of functional modules should be understood as the functional module architecture of the solution mainly implemented by the computer program recorded in the specification, and should not be understood as the physical device that mainly implements the solution by hardware.

[0272] Figure 8 The structure block diagram of a computing device 800 provided according to an embodiment of the present application is shown. The components of the computing device 800 include but are not limited to a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and the database 850 is used to store data.

[0273] The computing device 800 also includes an access device 840 that enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of network interface (e.g., a network interface card (NIC)) that is wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a World Wide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0274] In one embodiment of the present application, the above components of the computing device 800 and Figure 8 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 8The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of the present application. Those skilled in the art may add or replace other components as needed.

[0275] The computing device 800 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or PC. The computing device 800 may also be a mobile or stationary server.

[0276] The processor 820 is used to execute computer executable instructions to implement the steps of the recognition model training method and the text recognition method.

[0277] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above-mentioned recognition model training method and text recognition method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above-mentioned recognition model training method and text recognition method.

[0278] An embodiment of the present application also provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, are used for a training method for a recognition model and a text recognition method.

[0279] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned recognition model training method and text recognition method belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be found in the description of the technical scheme of the above-mentioned recognition model training method and text recognition method.

[0280] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0281] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0282] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0283] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0284] The preferred embodiments of the present application disclosed above are only used to help explain the present application. The optional embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the present application. The present application selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present application, so that those skilled in the art can understand and use the present application well. The present application is only limited by the claims and their full scope and equivalents.

Claims

1. A method for training a recognition model, characterized in that: include: Get training text; Inputting the training text into the recognition model to be trained, encoding the training text through the language module in the recognition model to be trained to obtain a feature vector; Input the feature vector into a filling module in the recognition model to be trained for processing to obtain a first intermediate vector, and input the feature vector into an intention module in the recognition model to be trained for processing to obtain a second intermediate vector, and calculate a relationship weight between the filling module and the intention module based on the first intermediate vector and the second intermediate vector; Obtaining a first text vector obtained by the filling module processing the feature vector, and a second text vector obtained by the intention module processing the feature vector; Calculating the product of the relationship weight and the first text vector to obtain a first target text vector, and calculating the product of the relationship weight and the second text vector to obtain a second target text vector; Determine a first loss value of the filling module according to the first target text vector and the entity vector corresponding to the training text, and determine a second loss value of the intent module according to the second target text vector and the intent vector corresponding to the training text; The recognition model to be trained is iteratively trained based on the first loss value and the second loss value until a training stop condition is reached to obtain a target recognition model.

2. The method for training a recognition model according to claim 1, characterized in that: The iterative training of the to-be-trained recognition model based on the first loss value and the second loss value comprises: Calculate the target loss value of the recognition model to be trained according to the first loss value and the second loss value; The recognition model to be trained is iteratively trained according to the target loss value.

3. The method for training a recognition model according to claim 2, characterized in that: The training stop condition is that the target loss value is less than or equal to the loss value threshold; Correspondingly, if the target loss value is less than or equal to the loss value threshold, the target recognition model is obtained; The loss value threshold is determined based on each target loss value generated by the recognition model to be trained during the training process.

4. The method for training a recognition model according to claim 2, characterized in that: The calculating the target loss value of the to-be-trained recognition model according to the first loss value and the second loss value includes: Determining a weight value of the first loss value and a weight value of the second loss value; A weighted summation process is performed on the weight value of the first loss value and the weight value of the second loss value to obtain the target loss value.

5. The method for training a recognition model according to claim 1, characterized in that: Before the step of inputting the training text into the recognition model to be trained, encoding the training text through the language module in the recognition model to be trained, and obtaining the feature vector, the step further includes: Performing word segmentation processing on the training text to obtain a word unit set; Correspondingly, the step of inputting the training text into the recognition model to be trained, encoding the training text through the language module in the recognition model to be trained, and obtaining a feature vector includes: The word unit set is input into the recognition model to be trained, and the word unit set is encoded by the language module to obtain the feature vector.

6. A text recognition method, characterized in that: include: Get the text to be recognized; Inputting the text to be recognized into a target recognition model for entity extraction and intent understanding, and obtaining a target entity and a target intent corresponding to the text to be recognized; Wherein, the target recognition model is obtained by training using the recognition model training method described in any one of claims 1 to 5.

7. The text recognition method according to claim 6, characterized in that: Also includes: Selecting a target template from the preset filling templates according to the target intention; The target entity is structurally processed according to the filling rule of the target template and the target template to obtain the target text.

8. The text recognition method according to claim 6, characterized in that: The step of inputting the text to be recognized into a target recognition model for entity extraction and intent understanding to obtain a target entity and a target intent corresponding to the text to be recognized includes: Inputting the text to be recognized into the target recognition model, encoding the text to be recognized through the language module in the target recognition model to obtain a text feature vector; The text feature vector is input into the filling module and the intention module in the target recognition model for processing, and the target entity and the target intention are obtained according to the processing results.

9. The text recognition method according to claim 8, characterized in that: Before the step of inputting the text to be recognized into the target recognition model, encoding the text to be recognized through the language module in the target recognition model, and obtaining the text feature vector, the method further includes: Performing word segmentation processing on the text to be recognized to obtain a word unit set; Correspondingly, the text to be recognized is input into the target recognition model, and the language module in the target recognition model encodes the text to be recognized to obtain a text feature vector, including: The word unit set is input into the language module for encoding processing to obtain the text feature vector.

10. The text recognition method according to claim 8, characterized in that: The step of inputting the text feature vector into a filling module and an intention module in the target recognition model for processing, and obtaining the target entity and the target intention according to the processing result, comprises: Inputting the text feature vector into the filling module for processing to obtain a first intermediate vector, and inputting the text feature vector into the intention module for processing to obtain a second intermediate vector; Calculate the product of the first intermediate vector and the first relationship weight of the intention module to obtain a first text vector, and calculate the product of the second intermediate vector and the second relationship weight of the filling module to obtain a second text vector; The first text vector is converted through the output layer of the filling module to obtain the target entity, and the second text vector is converted through the output layer of the intention module to obtain the target intention.

11. A training device for a recognition model, characterized in that: include: An acquisition unit, configured to acquire training text; An encoding unit is configured to input the training text into the recognition model to be trained, and perform encoding processing on the training text through the language module in the recognition model to be trained to obtain a feature vector; a determining unit configured to input the feature vector into a filling module in the recognition model to be trained for processing to obtain a first intermediate vector, and input the feature vector into an intention module in the recognition model to be trained for processing to obtain a second intermediate vector, and calculate a relationship weight between the filling module and the intention module based on the first intermediate vector and the second intermediate vector; The adjustment unit is configured to obtain a first text vector obtained by the filling module processing the feature vector, and a second text vector obtained by the intention module processing the feature vector; calculate the product of the relationship weight and the first text vector to obtain a first target text vector, and calculate the product of the relationship weight and the second text vector to obtain a second target text vector; Determine a first loss value of the filling module according to the first target text vector and the entity vector corresponding to the training text, and determine a second loss value of the intent module according to the second target text vector and the intent vector corresponding to the training text; The training unit is configured to iteratively train the recognition model to be trained based on the first loss value and the second loss value until a training stop condition is reached to obtain a target recognition model.

12. A text recognition device, characterized in that: include: A text acquisition unit is configured to acquire text to be recognized; A model processing unit is configured to input the text to be recognized into a target recognition model for entity extraction and intent understanding, and obtain a target entity and a target intent corresponding to the text to be recognized; Wherein, the target recognition model is obtained by training using the recognition model training method described in any one of claims 1 to 5.

13. A computing device, characterized in that: include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the method described in any one of claims 1 to 5 or 6-10.

14. A computer-readable storage medium storing computer instructions, characterized in that: When the instruction is executed by a processor, the steps of the method described in any one of claims 1 to 5 or 6 to 10 are implemented.

Citation Information

Patent Citations

  • Semantic comprehension method in task type dialogue system

    CN111104498A