Text label determination method, device, computer device, and storage medium
By interacting with single-word encoding and attention mechanism in the text label determination method, the problem of low accuracy of text label determination in traditional methods is solved, and a higher accuracy of label matching is achieved.
Patent Information
- Application Number
- CN202110412379.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-04-16
Smart Images

Figure CN113761188B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, apparatus, computer device, and storage medium for determining text labels. Background Art
[0002] With the development of artificial intelligence technology, text label determination technology has emerged. Text label determination refers to matching corresponding concept labels to an article, and a concept label is a phrase-level description used to reflect the content or central theme of the article.
[0003] In traditional technologies, text label determination is often performed through unsupervised extraction or a two-tower model in supervised training. Unsupervised extraction can specifically be methods such as unsupervised similarity calculation and graph-based ranking. Performing text label determination based on a two-tower model mainly means establishing a two-tower model composed of an article modeling module and a phrase modeling module. After obtaining the representations of the article and the phrase through these two modules respectively, different loss functions are selected to train a classification or ranking model.
[0004] However, in the traditional method of unsupervised extraction, due to the lack of supervision information, the extracted results usually have the problem of deviating from the central theme of the original article, resulting in a low accuracy of text label determination; when performing text label determination based on a two-tower model, due to problems such as asymmetric semantic representations, lack of interaction between the article and the phrase, and lack of interaction between candidate phrases, there is also a problem of low accuracy of text label determination. Summary of the Invention
[0005] Based on this, it is necessary to provide a method, apparatus, computer device, and storage medium for determining text labels that can improve the accuracy of label matching for the above technical problems.
[0006] A method for determining text labels, the method comprising:
[0007] Obtain a spliced text, where the spliced text includes the spliced candidate labels and the target text;
[0008] Encode each single character in the spliced text to obtain a word vector corresponding to each single character;
[0009] Use the attention mechanism to interact with each single character according to the word vector to obtain a feature vector corresponding to each single character in the candidate label;
[0010] Perform sequence labeling classification on each single character in the candidate label according to the feature vector to obtain a sequence labeling result corresponding to each single character in the candidate label;
[0011] Determine the target label corresponding to the target text according to the sequence labeling result.
[0012] A text label determination device, the device comprising:
[0013] An acquisition module, configured to acquire a spliced text, the spliced text including the spliced candidate labels and the target text;
[0014] An encoding module, configured to encode each single character in the spliced text to obtain a word vector corresponding to each single character;
[0015] An interaction module, configured to use an attention mechanism to interact with each single character according to the word vectors to obtain a feature vector corresponding to each single character in the candidate labels;
[0016] A classification module, configured to perform sequence labeling classification on each single character in the candidate labels according to the feature vectors to obtain a sequence labeling result corresponding to each single character in the candidate labels;
[0017] A processing module, configured to determine a target label corresponding to the target text according to the sequence labeling result.
[0018] A computer device, including a memory and a processor, the memory storing a computer program, and when the processor executes the computer program, the following steps are implemented:
[0019] Acquire a spliced text, the spliced text including the spliced candidate labels and the target text;
[0020] Encode each single character in the spliced text to obtain a word vector corresponding to each single character;
[0021] Use an attention mechanism to interact with each single character according to the word vectors to obtain a feature vector corresponding to each single character in the candidate labels;
[0022] Perform sequence labeling classification on each single character in the candidate labels according to the feature vectors to obtain a sequence labeling result corresponding to each single character in the candidate labels;
[0023] Determine a target label corresponding to the target text according to the sequence labeling result.
[0024] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0025] Acquire a spliced text, the spliced text including the spliced candidate labels and the target text;
[0026] Encode each single character in the spliced text to obtain a word vector corresponding to each single character;
[0027] Use an attention mechanism to interact with each single character according to the word vectors to obtain a feature vector corresponding to each single character in the candidate labels;
[0028] Perform sequence labeling classification on each single character in the candidate labels according to the feature vectors, and obtain the sequence labeling results corresponding to each single character in the candidate labels;
[0029] Determine the target label corresponding to the target text according to the sequence labeling results.
[0030] The above text label determination method, device, computer device and storage medium, after obtaining the text to be spliced, first encode each single character in the spliced text to obtain the word vectors corresponding to each single character, and then use the attention mechanism to interact with each single character according to the word vectors to obtain the feature vectors corresponding to each single character in the candidate labels, and then perform sequence labeling classification on each single character in the candidate labels according to the feature vectors. It can obtain accurate sequence labeling results under the condition of fully interacting with each label in the candidate labels and the semantics of each label and the target text. Therefore, the sequence labeling results can be used to match all target labels corresponding to the target text from the candidate labels at one time, which can improve the label matching accuracy. Description of the Drawings
[0031] Figure 1 It is a schematic flowchart of the text label determination method in an embodiment;
[0032] Figure 2 It is a schematic diagram of the text label determination method in an embodiment;
[0033] Figure 3 It is a schematic diagram of the text label determination method in an embodiment;
[0034] Figure 4 It is a schematic flowchart of the text label determination method in another embodiment;
[0035] Figure 5 It is a structural block diagram of the text label determination device in an embodiment;
[0036] Figure 6 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments
[0037] This application relates to the field of artificial intelligence technology. Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results in theory, methods, technologies and application systems. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.
[0038] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. What this application mainly involves is natural language processing technology. Natural Language Processing (NLP) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.
[0039] In order to make the purpose, technical solution and advantages of this application clearer, the following further details this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0040] In one embodiment, as Figure 1 shown, a method for determining text tags is provided. In this embodiment, this method is exemplified by being applied to a terminal. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms, or a node in a blockchain. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted device, etc., but is not limited thereto. When the method for determining text tags provided in this embodiment is implemented through the interaction between the terminal and the server, the terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this here. In this embodiment, the method includes the following steps:
[0041] Step 102, obtain the spliced text, where the spliced text includes the spliced candidate tags and the target text.
[0042] Among them, the spliced text refers to the spliced text that needs to perform tag matching. Here, splicing means splicing the candidate tags and the target text into a spliced text. A tag refers to a phrase-level description that reflects the content or central theme of an article. For example, a tag can specifically refer to a concept tag. A candidate tag refers to a set of tags available for matching. Determining the text tag is to determine the target tag corresponding to the target text from the candidate tags. The target text refers to the text that needs to match the tag, and the target tag corresponding to the target text can provide a general description of the target text.
[0043] For example, for the target text "A small-sized apartment of X square meters inside the unit, the decoration cost Y ten thousand yuan, and the effect after decoration is really good! Nowadays, decorating a house may easily cost more than one hundred thousand yuan. However, the money spent should be worth it, after all, everyone's money doesn't come easily", the candidate tags can be small-sized apartment decoration, villa decoration, etc. By using text tag determination, the target tag corresponding to the target text can be obtained as small-sized apartment decoration.
[0044] Specifically, when determining the text tag, the terminal will first obtain the candidate tags and the target text, splice the tags in the candidate tags to obtain the spliced candidate tags, and then splice the spliced candidate tags and the target text to obtain the spliced text. It should be noted that there are at least two tags in the candidate tags. Splicing the tags in the candidate tags means arranging the tags in sequence. For example, when the candidate tags include tag A and tag B, the spliced candidate tags obtained by splicing the tags in the candidate tags can specifically be in the form of tag A + tag B or tag B + tag A. Here, splicing does not affect or change individual tags.
[0045] Step 104: Encode each single character in the spliced text to obtain a word vector corresponding to each single character.
[0046] Among them, encoding means converting each single character in the spliced text into a word vector of a preset fixed length through embedding encoding, that is, representing each single character with a vector for subsequent processing. The preset fixed length can be set according to needs.
[0047] Specifically, after obtaining the spliced text, the terminal will encode each single character in the spliced text, convert each single character in the spliced text into a word vector of a preset fixed length, and different single characters are represented by different word vectors.
[0048] Step 106: Use the attention mechanism to interact with each single character according to the word vectors to obtain a feature vector corresponding to each single character in the candidate tags.
[0049] Wherein, interaction means that after encoding each word, the word vector is first processed by the feature extraction network to extract more context information, and then the interaction information between the candidate label and the target text is obtained by using the unidirectional or bidirectional attention mechanism and the context information. Through the interaction information, it can be inferred which part of the candidate label is more important to the target text. Wherein, the feature extraction network can specifically be a recurrent neural network, a convolutional neural network, etc., and this embodiment is not specifically limited here. Wherein, the interaction information can specifically refer to a similarity matrix obtained by the similarity between each word in the spliced text, and the similarity between each word in the spliced text can be obtained by the word vector corresponding to each word. A feature vector refers to a vector used to represent a word after interaction. For example, a feature vector can specifically refer to a vector with the same length as a word vector. In the feature vector of each word, the word vectors of other words in the candidate label are integrated.
[0050] Specifically, after obtaining the word vectors corresponding to each word, the terminal will use the attention mechanism to calculate the similarity between each word according to the word vectors, and obtain the similarity matrix corresponding to the spliced text according to the similarity between each word. According to the similarity in the similarity matrix, the relative weight coefficient between each word can be calculated, and the relative weight coefficient can be used to represent the importance of each word in the spliced text. After obtaining the relative weight coefficient, the terminal can obtain the feature vector corresponding to each word in the candidate label according to the relative weight coefficient and the word vector of each word. It should be noted that the interactive process is to mine the relationship between the candidate label and the target text by calculating the similarity between each word in the spliced text, so as to realize the use of the word vector of each word in the spliced text to re-represent each word and obtain the feature vector of each word. In order to deeply mine the relationship between the candidate label and the target text, the interactive process between the two may sometimes be performed multiple times to simulate the behavior of humans repeating reading when performing reading comprehension. Therefore, after obtaining a feature vector once, the terminal will use the feature vector as a new word vector of the single word, and interact again according to the feature vector to continuously update the feature vector until the number of interactions reaches a preset interaction number threshold. The preset interaction number threshold here can be set as needed.
[0051] Step 108 , performing sequence labeling classification on each word in the candidate label according to the feature vector, and obtaining a sequence labeling result corresponding to each word in the candidate label.
[0052] Among them, sequence labeling classification refers to predicting the sequence labels corresponding to each single character in the candidate labels. For example, the sequence labels can specifically refer to BIO (Begin, Intermediate, Other) labels. The sequence labeling result refers to the predicted sequence labels corresponding to each single character in the candidate labels. This sequence label is used to distinguish each single character. Through the sequence label, the corresponding relationship between each single character and the target text can be determined, and then the corresponding relationship between the label in the candidate label corresponding to the single character and the target text can be obtained.
[0053] Specifically, after obtaining the feature vector, the terminal will obtain the trained label vector conversion matrix, and convert the feature vector into a sequence label vector according to the trained label vector conversion matrix, obtaining the sequence label vectors corresponding to each single character. The element values of each element in the sequence label vector correspond to the category probabilities of the single character belonging to each preset sequence label. After the terminal obtains the sequence label vector, it can perform sequence labeling classification on each single character in the candidate label according to the sequence label vector and the corresponding relationship between the elements in the preset sequence label vector and the preset sequence labels, obtaining the category probabilities of each single character belonging to each preset sequence label. By sorting the category probabilities, the sequence labeling result corresponding to each single character can be obtained.
[0054] Among them, the sequence label vector is used to represent the sequence label information corresponding to each single character. For example, the sequence label vector can specifically be a vector obtained with the category probabilities of each single character belonging to each preset sequence label as elements. For example, when the sequence label is a BIO label, the form of the sequence label vector can specifically be [x y z], where x, y, and z are the probabilities of the single character belonging to the B, I, and O labels respectively. Among them, in the BIO label, B represents the start position, I represents the middle position, and O represents the irrelevant position.
[0055] Among them, the corresponding relationship between the elements in the sequence label vector and the preset sequence labels is set in advance. For example, when the sequence label is a BIO label, it can be preset that the first element in the feature vector corresponds to the B label, the second element corresponds to I, and the third label corresponds to O. In this way, after obtaining the sequence label vector, the category probabilities of each single character belonging to each preset sequence label can be obtained based on the sequence label vector.
[0056] Step 110, determine the target label corresponding to the target text according to the sequence labeling result.
[0057] Among them, the target label refers to the label that can correctly match the target text and can be screened from the candidate labels.
[0058] Specifically, after obtaining the sequence annotation result, the terminal will determine the valid sequence tags from the sequence annotation result according to the preset valid tags. Based on the valid sequence tags, the valid single characters can be determined, and then the target tags corresponding to the target text can be determined according to the valid single characters. Among them, the preset valid tags refer to the tags with practical meanings set in advance. For example, when the sequence tag is a BIO tag, since the tag B can be used to represent the start position, the tag I can be used to represent the middle position, and the tag O can be used to represent other positions, then the tags B and I can be understood as tags with practical meanings, and the tag O can be understood as a tag without practical meaning. Therefore, when setting the valid tags, the tags B and I need to be set as valid tags.
[0059] The above text tag determination method, after obtaining the text to be spliced, first encodes each single character in the spliced text to obtain the word vectors corresponding to each single character, then uses the attention mechanism to interact with each single character according to the word vectors to obtain the feature vectors corresponding to each single character in the candidate tags, and then performs sequence annotation classification on each single character in the candidate tags according to the feature vectors. It can obtain accurate sequence annotation results when fully interacting with the semantics of each tag in the candidate tags and each tag and the target text, so that the target tags corresponding to all the target texts can be matched from the candidate tags at one time by using the sequence annotation results, which can improve the accuracy of tag matching.
[0060] In one embodiment, using the attention mechanism to interact with each single character according to the word vectors to obtain the feature vectors corresponding to each single character in the candidate tags includes:
[0061] Using the attention mechanism to calculate the similarity between each single character according to the word vectors to obtain a similarity matrix corresponding to the spliced text;
[0062] Normalize the similarity matrix to determine the relative weight coefficients between each single character;
[0063] Perform vector weighting according to the relative weight coefficients and the word vectors to obtain the feature vectors corresponding to each single character in the candidate tags.
[0064] Among them, the similarity is used to characterize the similarity degree between each single character. For example, specifically, the similarity can refer to the vector similarity between the word vectors of each single character. The similarity matrix refers to the matrix composed of the similarities between each single character. The relative weight coefficient is used to represent the importance degree between each single character in the spliced text. Here, the importance degree between each single character refers to the importance degree of other single characters to any single character in the spliced text.
[0065] Specifically, the terminal will use the attention mechanism to calculate the similarity between each single word and all single words in the candidate label based on the word vectors. According to the calculated similarity, a similarity matrix corresponding to the concatenated text is obtained. Based on the similarity values in the similarity matrix, the similarity matrix is normalized to determine the relative weight coefficients between each single word. The relative weight coefficient is the normalized similarity value. After obtaining the relative weight coefficients, for each single word in the candidate label, the terminal will perform vector weighting based on the word vectors of other single words in the candidate label and its relative weight coefficient to obtain the feature vector corresponding to each single word. To deeply explore the relationship between the candidate label and the target text, the interaction process between the two may be executed multiple times to simulate the behavior of humans repeating reading when performing reading comprehension. Therefore, after obtaining the feature vectors once, the terminal will use the feature vectors as the new word vectors of the single words and perform interactions again based on the feature vectors to continuously update the feature vectors until the number of interaction times reaches the preset interaction times threshold. Here, the preset interaction times threshold can be set as needed.
[0066] In this embodiment, by using the attention mechanism to interact with each single word according to the word vectors, the relationship between the candidate label and the target text can be fully explored, and the feature vectors corresponding to each single word in the candidate label can be obtained.
[0067] In one embodiment, performing sequence labeling classification on each single word in the candidate label according to the feature vectors, and the obtained sequence labeling results corresponding to each single word in the candidate label include:
[0068] Obtain the trained label vector conversion matrix;
[0069] According to the feature vectors and the label vector conversion matrix, obtain the sequence label vectors corresponding to each single word in the candidate label;
[0070] Perform sequence labeling classification on each single word in the candidate label according to the sequence label vectors, and obtain the sequence labeling results corresponding to each single word in the candidate label.
[0071] Among them, the trained label vector conversion matrix refers to the matrix used to convert the feature vectors into sequence label vectors, which can be obtained through pre-training. The size of the label vector conversion matrix is related to the feature vectors and the preset number of sequence labels. For example, when the size of the feature vectors is 1*N and the preset number of sequence labels is 3, the size of the label vector conversion matrix can be obtained as N*3. Using this label vector conversion matrix, the feature vectors can be converted into 1*3 sequence label vectors.
[0072] Specifically, after obtaining the feature vector, the terminal will acquire the trained label vector conversion matrix, use the label vector conversion matrix to convert the feature vector, convert the feature vector into a corresponding sequence label vector, and then perform sequence annotation classification on each single character in the candidate label according to the element values in the sequence label vector, to obtain the sequence annotation result corresponding to each single character in the candidate label.
[0073] In this embodiment, by obtaining the label vector conversion matrix, it is possible to use the label vector conversion matrix to achieve the conversion of the feature vector, obtain the sequence label vector corresponding to each single character, and further be able to perform sequence annotation classification on each single character in the candidate label according to the sequence label vector, to obtain the sequence annotation result corresponding to each single character in the candidate label.
[0074] In one embodiment, performing sequence annotation classification on each single character in the candidate label according to the sequence label vector to obtain the sequence annotation result corresponding to each single character in the candidate label includes:
[0075] Performing sequence annotation classification on each single character in the candidate label according to the sequence label vector to determine the category probability of each single character belonging to each preset sequence label;
[0076] According to the category probability, obtain the sequence annotation result corresponding to each single character in the candidate label.
[0077] Wherein, the preset sequence label refers to a pre-set label used to distinguish the position of each single character in the label. For example, the preset sequence label can specifically refer to the BIO label.
[0078] Specifically, the element values of each element in the sequence label vector correspond to the category probability of the single character belonging to each preset sequence label. After the terminal obtains the sequence label vector, it can perform sequence annotation classification on each single character in the candidate label according to the sequence label vector and the corresponding relationship between the elements in the preset sequence label vector and the preset sequence label, to obtain the category probability of each single character belonging to each preset sequence label. By sorting the category probabilities, the sequence annotation result corresponding to each single character can be obtained. Among them, the sequence annotation result refers to the preset sequence label corresponding to the maximum category probability. For example, for the single character A, when its category probabilities of belonging to label B, label I, and label O are 0.6, 0.3, and 0.1 respectively, the sequence annotation result of the single character A can be obtained as label B, that is, the start position of the target label correctly matched by this single character A.
[0079] In this embodiment, by performing sequence annotation classification on each single character in the candidate label according to the sequence label vector to obtain the category probability of each single character belonging to each preset sequence label, it is possible to obtain the sequence annotation result corresponding to each single character in the candidate label according to the category probability.
[0080] In one embodiment, determining a target label corresponding to a target text according to the sequence annotation result includes:
[0081] Determining valid sequence labels according to the sequence annotation result;
[0082] Filtering out target labels corresponding to the target text from the candidate labels according to the valid sequence labels.
[0083] Among them, a valid sequence label refers to a label with practical significance in the sequence annotation result. For example, when the sequence label is a BIO label, since label B and label I will respectively point to the start position and the middle position of the matching label, it is a label with practical significance. Since label O will point to a non-matching label, it is a label without practical significance. A target label refers to a label that matches the target text.
[0084] Specifically, after obtaining the sequence annotation result, the terminal will filter out valid sequence labels from the sequence annotation result according to the preset valid labels, and filter out labels corresponding to the valid sequence labels from the candidate labels, and use the labels corresponding to the valid sequence labels as the target labels corresponding to the target text. Among them, the preset valid labels refer to the labels with practical significance set in advance.
[0085] In this embodiment, by determining valid sequence labels according to the sequence annotation result, and filtering out target labels corresponding to the target text from the candidate labels according to the valid sequence labels, it is possible to determine the target labels by using the valid sequence labels.
[0086] In one embodiment, the sequence annotation results corresponding to each single character in the candidate labels in the above embodiment are obtained through a text label matching model;
[0087] The construction process of the text label matching model includes:
[0088] Obtaining an initial text matching model and classification matching training data. The classification matching training data includes the spliced classification labels and the classification matching texts corresponding to the classification labels, and the classification labels carry classification sequence labels;
[0089] Training the initial text matching model according to the classification matching training data to obtain an initial text label matching model;
[0090] Obtaining label matching training data. The label matching training data includes the spliced training labels and the label matching texts matching the training labels, and the training labels carry training sequence labels;
[0091] Training the initial text label matching model according to the label matching training data to obtain a trained text label matching model.
[0092] Among them, the pre-trained text label matching model refers to a pre-trained model for text label matching. The pre-trained text label matching model includes an encoding layer and an output layer. The encoding layer is used to encode and interact with the input data to obtain a feature vector corresponding to the input data. The output layer is mainly used to perform sequence label classification on each single character in the candidate labels according to the feature vector. For example, the output layer can specifically be a fully connected layer for sequence annotation classification. For example, the pre-trained text label matching model can specifically refer to a model constructed based on the machine reading comprehension task. In the traditional model constructed based on the machine reading comprehension task, usually the context and the question are given as the input, and the model needs to find the answer in the context according to the question. In this embodiment, in the scenario of text label determination, the text label matching model takes the target text as the question and the candidate labels as the context, and the obtained answer is the correct label.
[0093] For example, the model constructed based on the machine reading comprehension task can specifically refer to a reading comprehension model based on BERT (Bidirectional Encoder Representations from Transformers), such as Figure 2 shown. In the reading comprehension model based on BERT, the candidate labels are concatenated as the context, and the title and the text body (i.e., the target text) are used as the question. The obtained answer is the sequence annotation result corresponding to each single character in the candidate labels. According to this sequence annotation, the target label can be determined from the candidate labels. Further, as Figure 2 shown, in addition to the title and the text body in the question, there is also prior knowledge. Here, the prior knowledge refers to the description related to the target text. For example, the prior knowledge can specifically refer to the classification information of the target text, which can help the model better learn the semantic representation of the target text. When there is prior knowledge, the prior knowledge will participate in the text label matching as a part of the target text.
[0094] Among them, the classification matching training data refers to the training set used for classification training. The classification matching training data includes the spliced classification labels and the classification matching texts corresponding to the classification labels. It should be noted that the classification labels here include the first classification labels that match the classification matching texts and the second classification labels that do not match the classification matching texts. The first classification labels and the second classification labels can be distinguished by the classification sequence labels carried by the classification labels. The classification sequence labels refer to the labels pre-labeled for the first classification labels and the second classification labels, and are used to label the positions of each single character in the classification labels of the first classification labels and the second classification labels. For example, the labeled labels can specifically refer to the BIO labels. For the first classification labels, B represents its start position and I represents the middle position. For the second classification labels, O represents its entire position. For example, for the first classification label AAAA, its classification sequence label can specifically be BIII, and for the second classification label BBBBB, its classification sequence label can specifically be OOOOO.
[0095] For example, the classification labels can specifically correspond to search queries. Among them, there are multiple first classification labels and second classification labels that represent search query phrases. The search query phrases and the corresponding classification matching texts can be obtained from the database storing search logs. By mining the search logs, the search queries and the corresponding clicked articles can be obtained. Thus, the search query can be used as the first classification label of the corresponding clicked article, and other non-corresponding search queries can be used as the second classification labels. The search queries contain the supervision signals from user clicks and have a large data scale, which are good classification data. However, there is still a certain gap between the search queries and the real labels. Therefore, preferably, a closer classification data source can be used.
[0096] For example, a closer classification data source can specifically refer to the second- and third-level classification data that is semantically close to the labels. Among them, the classification labels can specifically correspond to the second-level classification. There are multiple first classification labels and second classification labels that represent the third-level classification. Here, the second-level classification and the third-level classification refer to the classifications divided according to the classification scope. The scope of the second-level classification is larger than that of the third-level classification. Among the second-level classifications, there are multiple third-level classifications. For example, under the first-level classification of sports, there are second-level classifications such as basketball and football. Under the second-level classification of basketball, there are third-level classifications such as NBA (National Basketball Association) and CBA (China Basketball Association).
[0097] Among them, the initial text matching model refers to the model used to implement text matching, including an encoding layer and an output layer. For example, the initial text matching model can specifically refer to a model constructed based on the machine reading comprehension task. In the traditional model constructed based on the machine reading comprehension task, usually the context and the question are given as inputs, and the model needs to find the answer in the context according to the question. In this embodiment, in the scenario of classification matching, the classification matching text corresponding to the classification label is used as the question, and the classification matching training data is used as the context, and the classification result is output, that is, the classification task is converted into a label matching task, and the initial text matching model is trained using the classification matching training data. Since the label is close to the second- and third-level classification semantics, therefore, by introducing the training of classification matching, the initial text matching model can learn fine-grained semantic discrimination, thereby strengthening the fine-grained semantic discrimination ability of the initial text matching model.
[0098] Among them, the label matching training data refers to the training set used to train the initial text label matching model. The label matching training data includes the spliced training labels and the label matching text that matches the training labels. The sequence label refers to the label pre-labeled for each single character in the training label, which is used to label the position of each single character in the training label in the training label. For example, the sequence label can specifically refer to the BIO label, where B represents the start position of the training label, and I represents the middle position of the training label. The initial text label matching model refers to the text label matching model to be trained. The initial text label matching model includes an encoding layer and an output layer. The encoding layer therein is used to encode and interact with the label matching training data to obtain the feature vector corresponding to the label matching training data, and the output layer is used to perform sequence labeling classification on the training labels according to the feature vector to obtain the sequence labeling results corresponding to each single character in the training label.
[0099] It should be noted that the number of training labels here can be more than one, that is, it includes all the labels that match the label matching text, and different training labels are distinguished by sequence labels. For example, when the sequence label is the BIO label, B represents the start position of the training label, and I represents the middle position of the training label. The sequence label corresponding to each training label is from the start position to the previous middle position before the next start position. For example, when the spliced training labels are two training labels AAAA and BBBBB, the corresponding sequence label is BIIIBIIII.
[0100] Specifically, the terminal will obtain the classification matching training data and the initial text matching model, input the classification matching training data into the initial text matching model, first encode and interact with the classification matching training data through the encoding layer therein to obtain the feature vector corresponding to the classification matching training data, and then perform sequence labeling classification on the classification labels according to the feature vector through the output layer therein to obtain the classification sequence labeling results corresponding to each single character in the classification labels. After obtaining the classification sequence labeling results, the terminal will calculate the model loss function by comparing the classification sequence labeling results of the same single character with the sequence labels corresponding to the single character in the classification sequence labels, adjust the parameters of the initial text matching model according to the model loss function, and retrain the initial text matching model with adjusted parameters according to the classification matching training data until the model loss function meets the preset training end condition, thereby obtaining the initial text label matching model. Among them, the preset training end condition can be set as needed, including but not limited to conditions such as the model loss function being less than the preset loss function threshold and the model loss function converging.
[0101] Specifically, after obtaining the initial text label matching model, the terminal will obtain the label matching training data, input the label matching training data into the initial text label matching model, first encode and interact with the label matching training data through the encoding layer therein to obtain the feature vector corresponding to the label matching training data, and then perform sequence labeling classification on the training labels according to the feature vector through the output layer therein to obtain the sequence labeling results corresponding to each single character in the training labels. After obtaining the sequence labeling results, the terminal will calculate the model loss function by comparing the sequence labeling results of the same single character with the sequence labels corresponding to the single character in the sequence labels, adjust the parameters of the initial text label matching model according to the model loss function, and retrain the initial text label matching model with adjusted parameters according to the label matching training data until the model loss function meets the preset training end condition, thereby obtaining the trained text label matching model. Among them, the preset training end condition can be set as needed, including but not limited to conditions such as the model loss function being less than the preset loss function threshold and the model loss function converging.
[0102] Further, when the output layer performs sequence labeling classification on the training labels based on the feature vectors, it mainly realizes the conversion of the feature vectors through the label vector conversion matrix in the output layer, converts the feature vectors into corresponding sequence label vectors, and then realizes the sequence labeling classification of each single character in the training labels based on the sequence label vectors, so as to obtain the sequence labeling results corresponding to each single character in the training labels. Among them, the size of the label vector conversion matrix is related to the feature vectors and the preset number of sequence labels. For example, when the size of the feature vectors is 1*N and the preset number of sequence labels is 3, the size of the label vector conversion matrix can be obtained as N*3, and the feature vectors can be converted into 1*3 sequence label vectors by using this label vector conversion matrix.
[0103] In this embodiment, by first introducing classification matching training data to pre-train the initial text matching model, the initial text matching model can learn fine-grained semantic distinctions, thereby strengthening the fine-grained semantic distinction ability of the initial text matching model, obtaining a more accurate initial text label matching model, and then the accurate trained text label matching model can be obtained by training the initial text label matching model using the label matching training data.
[0104] In one embodiment, training the initial text label matching model according to the label matching training data to obtain the trained text label matching model includes:
[0105] Training the initial text label matching model according to the label matching training data to obtain a preliminary trained text label matching model;
[0106] Generating task joint training data according to the classification matching training data and the label matching training data;
[0107] Training the preliminary trained text label matching model according to the task joint training data to obtain the trained text label matching model.
[0108] Specifically, when training the initial text label matching model based on the training data matched by tags to obtain the trained text label matching model, in addition to directly training to obtain the trained text label matching model, it is also possible to first obtain the preliminary trained text label matching model, then generate task joint training data, and train the preliminary trained text label matching model according to the task joint training data to obtain the trained text label matching model. Among them, the task joint training data includes classification matching training data and label matching training data. Training the preliminary trained text label matching model according to the task joint training data means simultaneously using the classification matching training data and the label matching training data to train the preliminary trained text label matching model, that is, alternately using the classification matching training data or the label matching training data to train the preliminary trained text label matching model.
[0109] Furthermore, during the training process, a task identifier will also be concatenated in front of the label matching text and the classification matching text to distinguish the text label matching task and the classification task. Since the classification task is a task related to the text label matching task, by introducing the classification task as an auxiliary task for multi-task learning, the performance of the model can be improved using task joint training. It should be noted that there may be the same matching text in the label matching text and the classification matching text, that is, there is target matching text that can be used for both the classification task and the text label matching task. Preferably, when the terminal performs task joint training, it can screen out the same matching text from the classification matching training data and the label matching training data as the target matching text, and obtain the task joint training data according to the target matching text.
[0110] For example, as Figure 3 shown, the target matching text "Snail rice noodles are nothing! Are you brave enough to try the strong-flavored food in Guangxi?" is included in both the classification matching training data and the label matching training data. When using this target matching text for task joint training, for the classification task, its classification labels include the first classification label "Must-eat in the city", and the second classification labels "Internet-famous food", "Food and drink guide", etc. For the text label matching task, its training labels include "Guangxi cuisine", "Snail rice noodle recipe", "Guangxi tourism", etc. Among them, the sequence label of "Guangxi cuisine" is BIII, the sequence label of "Snail rice noodle recipe" is OOOOO, and the sequence label of "Guangxi tourism" is OOOOO.
[0111] In this embodiment, by first generating task joint training data and then training according to the task joint training data, the performance of the preliminary trained text label matching model can be improved through task joint training, and a trained text label matching model that can accurately determine text labels can be obtained.
[0112] The present application also provides an application scenario, which applies the above-mentioned text label determination method. Specifically, the application of the text label determination method in this application scenario is as follows:
[0113] Among them, the target text is "A small-sized apartment with X square meters inside, and the decoration cost Y ten thousand yuan. The decoration effect is really good! Nowadays, decorating a house may cost more than one hundred thousand yuan casually. However, the money spent should be worth it, after all, everyone's money doesn't come easily.", the candidate labels are "Small-sized apartment decoration, Villa decoration, Rural housing", and the target label is "Small-sized apartment decoration";
[0114] The terminal first obtains the spliced text, where the spliced text includes the spliced candidate labels and the target text. Then it encodes each single character in the spliced text to obtain the word vectors corresponding to each single character, uses the attention mechanism to interact each single character according to the word vectors to obtain the feature vectors corresponding to each single character in the candidate labels, classifies each single character in the candidate labels by sequence labeling according to the feature vectors to obtain the sequence labeling results corresponding to each single character in the candidate labels, and determines the target label corresponding to the target text according to the sequence labeling results. Among them, when the spliced candidate labels are "Small-sized apartment decoration + Villa decoration + Rural housing" and the sequence labels are BIO labels, the corresponding sequence labeling results can be obtained as "BIIII + OOOO + OOOO", and the target label can be obtained as "Small-sized apartment decoration" according to this sequence labeling result.
[0115] As Figure 4 shown, the present application also provides a process schematic diagram to illustrate the text label determination method of the present application. The text label determination method specifically includes the following steps:
[0116] Step 402, obtain the spliced text, where the spliced text includes the spliced candidate labels and the target text;
[0117] Step 404, encode each single character in the spliced text to obtain the word vectors corresponding to each single character;
[0118] Step 406, use the attention mechanism to calculate the similarity between each single character according to the word vectors to obtain the similarity matrix corresponding to the spliced text;
[0119] Step 408, normalize the similarity matrix to determine the relative weight coefficients between each single character;
[0120] Step 410, perform vector weighting according to the relative weight coefficients and the word vectors to obtain the feature vectors corresponding to each single character in the candidate labels;
[0121] Step 412, obtain the trained label vector conversion matrix;
[0122] Step 414: Obtain sequence label vectors corresponding to each single character in the candidate labels according to the feature vector and the label vector transformation matrix;
[0123] Step 416: Perform sequence annotation classification on each single character in the candidate labels according to the sequence label vectors, and determine the class probabilities of each single character belonging to each preset sequence label;
[0124] Step 418: Obtain the sequence annotation results corresponding to each single character in the candidate labels according to the class probabilities;
[0125] Step 420: Determine the valid sequence labels according to the sequence annotation results;
[0126] Step 422: Screen out the target labels corresponding to the target text from the candidate labels according to the valid sequence labels.
[0127] Compared with the traditional method, the text label determination method proposed in this application can fully realize the interaction between the target text and the candidate labels, and can complete the matching process of all labels in the candidate labels at one time, with stronger matching ability and lower cost. In addition, in the above embodiments, training optimization of multi-domain data migration from far to near is also proposed, including introducing training for classification matching first, strengthening the fine-grained semantic discrimination ability of the model, and improving the performance of the model using task joint training, etc., which can further improve the effect of text label determination. By comparing with existing unsupervised strategy methods, classification methods based on GBDT (Gradient Boosting Decision Tree), etc. on the same data set, it can be found that the text label determination method proposed in this application has an obvious improvement in effect.
[0128] Specifically, the comparison data can be as shown in Table 1. Among them, classification data pre-training refers to the way of introducing training for classification matching, classification auxiliary task refers to task joint training, query data pre-training refers to pre-training optimization using previously collected query data, and the MRC (Machine Reading Comprehension) matching model refers to the text label matching model without training optimization:
[0129] Table 1
[0130]
[0131]
[0132] It should be understood that although the steps in the various flowcharts involved in the above embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the various flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0133] In one embodiment, as Figure 5 shown, a text label determination device is provided. This device can be a software module, a hardware module, or a combination of both to become a part of a computer device. Specifically, the device includes: an acquisition module 502, an encoding module 504, an interaction module 506, a classification module 508, and a processing module 510, where:
[0134] The acquisition module 502 is used to acquire the spliced text, and the spliced text includes the spliced candidate labels and the target text;
[0135] The encoding module 504 is used to encode each single character in the spliced text to obtain a word vector corresponding to each single character;
[0136] The interaction module 506 is used to interact with each single character according to the word vector by using the attention mechanism to obtain a feature vector corresponding to each single character in the candidate label;
[0137] The classification module 508 is used to perform sequence annotation classification on each single character in the candidate label according to the feature vector to obtain a sequence annotation result corresponding to each single character in the candidate label;
[0138] The processing module 510 is used to determine the target label corresponding to the target text according to the sequence annotation result.
[0139] The above text label determination device, after acquiring the text to be spliced, first encodes each single character in the spliced text to obtain a word vector corresponding to each single character, then uses the attention mechanism to interact with each single character according to the word vector to obtain a feature vector corresponding to each single character in the candidate label, and then performs sequence annotation classification on each single character in the candidate label according to the feature vector. It can obtain accurate sequence annotation results under the condition of fully interacting with the semantics of each label in the candidate label and each label and the target text, so that the sequence annotation results can be used to match all the target labels corresponding to the target text from the candidate labels at one time, which can improve the label matching accuracy.
[0140] In one embodiment, the interaction module is further configured to use the attention mechanism to calculate the similarity between each single word according to the word vectors, obtain a similarity matrix corresponding to the concatenated text, normalize the similarity matrix, determine the relative weight coefficients between each single word, and perform vector weighting according to the relative weight coefficients and the word vectors to obtain the feature vectors corresponding to each single word in the candidate label.
[0141] In one embodiment, the classification module is further configured to obtain a trained label vector conversion matrix, obtain the sequence label vectors corresponding to each single word in the candidate label according to the feature vectors and the label vector conversion matrix, and perform sequence annotation classification on each single word in the candidate label according to the sequence label vectors to obtain the sequence annotation results corresponding to each single word in the candidate label.
[0142] In one embodiment, the classification module is further configured to perform sequence annotation classification on each single word in the candidate label according to the sequence label vectors, determine the category probabilities of each single word belonging to each preset sequence label, and obtain the sequence annotation results corresponding to each single word in the candidate label according to the category probabilities.
[0143] In one embodiment, the processing module is further configured to determine valid sequence labels according to the sequence annotation results, and screen out the target labels corresponding to the target text from the candidate labels according to the valid sequence labels.
[0144] In one embodiment, the sequence annotation results corresponding to each single word in the candidate label in the above embodiment can be obtained through a text label matching model. The text label determination device further includes a model processing module, and the model processing module includes a text label matching model. The model processing module is further configured to obtain an initial text matching model and classification matching training data. The classification matching training data includes concatenated classification labels and classification matching texts corresponding to the classification labels. The classification labels carry classification sequence labels. Train the initial text matching model according to the classification matching training data to obtain an initial text label matching model. Obtain label matching training data. The label matching training data includes concatenated training labels and label matching texts matching the training labels. The training labels carry training sequence labels. Train the initial text label matching model according to the label matching training data to obtain a trained text label matching model.
[0145] In one embodiment, the model processing module is further configured to train the initial text label matching model according to the label matching training data to obtain a preliminary trained text label matching model, generate task joint training data according to the classification matching training data and the label matching training data, and train the preliminary trained text label matching model according to the task joint training data to obtain a trained text label matching model.
[0146] For the specific limitations of the text label determination device, reference can be made to the limitations of the text label determination method in the above text, which will not be elaborated here. Each module in the above text label determination device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0147] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structural diagram can be as Figure 6 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a text label determination method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0148] Those skilled in the art can understand that Figure 6 the structure shown in
[0149] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0150] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, it implements the steps in the above method embodiments.
[0151] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the steps in the above method embodiments.
[0152] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0153] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0154] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for determining text tags, characterized in that, The method comprises: Acquire a concatenated text, wherein the concatenated text includes concatenated candidate tags and target text; the candidate tags refer to a set of tags to be matched; the target text refers to the text that needs to match the tags; Encoding each single word in the concatenated text to obtain a word vector corresponding to each single word; Using an attention mechanism to interact with each word according to the word vector, so as to mine the relationship between the candidate tag and the target text, and obtain a feature vector corresponding to each word in the candidate tag; Perform sequence labeling classification on each word in the candidate label according to the feature vector to obtain a sequence labeling result corresponding to each word in the candidate label; Determining valid sequence tags according to the sequence annotation results; According to the valid sequence labels, a target label corresponding to the target text is screened out from the candidate labels; the target label is used to give a general description of the target text.
2. The method according to claim 1, wherein The method of using the attention mechanism to interact with each word according to the word vector to mine the relationship between the candidate tag and the target text, and obtaining the feature vector corresponding to each word in the candidate tag includes: Calculating the similarity between the individual words according to the word vectors using an attention mechanism to obtain a similarity matrix corresponding to the concatenated text; Normalizing the similarity matrix to determine the relative weight coefficients between the individual words; Vector weighting is performed according to the relative weight coefficient and the word vector to obtain a feature vector corresponding to each word in the candidate tag.
3. The method according to claim 1, wherein The step of performing sequence labeling classification on each word in the candidate label according to the feature vector to obtain a sequence labeling result corresponding to each word in the candidate label comprises: Get the trained label vector transformation matrix; Obtaining a sequence label vector corresponding to each word in the candidate label according to the feature vector and the label vector conversion matrix; The individual words in the candidate tags are sequence labeled and classified according to the sequence label vector to obtain the sequence labeling results corresponding to the individual words in the candidate tags.
4. The method according to claim 3, wherein The step of performing sequence labeling classification on each word in the candidate label according to the sequence label vector to obtain a sequence labeling result corresponding to each word in the candidate label includes: Perform sequence labeling classification on each word in the candidate label according to the sequence label vector, and determine the probability that each word belongs to each preset sequence label; According to the category probability, the sequence labeling result corresponding to each word in the candidate tag is obtained.
5. The method according to any one of claims 1 to 4, characterized in that, The sequence labeling results corresponding to each word in the candidate label are obtained by a text label matching model; The process of constructing the text label matching model includes: Acquire an initial text matching model and classification matching training data, wherein the classification matching training data includes concatenated classification labels and classification matching texts corresponding to the classification labels, and the classification labels carry classification sequence labels; Training the initial text matching model according to the classification matching training data to obtain an initial text label matching model; Obtain label matching training data, where the label matching training data includes spliced training labels and label matching texts that match the training labels, and the training labels carry training sequence labels; Train the initial text label matching model according to the label matching training data to obtain a trained text label matching model.
6. The method according to claim 5, characterized in that, The training of the initial text label matching model according to the label matching training data to obtain a trained text label matching model includes: Train the initial text label matching model according to the label matching training data to obtain a preliminary trained text label matching model; Generate task joint training data according to the classification matching training data and the label matching training data; Train the preliminary trained text label matching model according to the task joint training data to obtain a trained text label matching model.
7. A text label determination device, characterized in that The device includes: An acquisition module, configured to acquire a spliced text, where the spliced text includes a spliced candidate label and a target text; the candidate label refers to a set of labels to be matched; the target text refers to the text for which a label needs to be matched; An encoding module, configured to encode each single character in the spliced text to obtain a word vector corresponding to each single character; An interaction module, configured to utilize an attention mechanism to interact with each single character according to the word vector to mine the relationship between the candidate label and the target text, and obtain a feature vector corresponding to each single character in the candidate label; A classification module, configured to perform sequence annotation classification on each single character in the candidate label according to the feature vector to obtain a sequence annotation result corresponding to each single character in the candidate label; A processing module, configured to determine a valid sequence label according to the sequence annotation result, and filter out a target label corresponding to the target text from the candidate labels according to the valid sequence label; the target label is used to give a general description of the target text.
8. The device according to claim 7, characterized in that, The interaction module is further configured to calculate the similarity between each single character according to the word vector by using the attention mechanism to obtain a similarity matrix corresponding to the spliced text, normalize the similarity matrix, determine the relative weight coefficient between each single character, and perform vector weighting according to the relative weight coefficient and the word vector to obtain a feature vector corresponding to each single character in the candidate label.
9. The device according to claim 7, characterized in that, The classification module is further configured to obtain a trained label vector conversion matrix, obtain a sequence label vector corresponding to each single character in the candidate label according to the feature vector and the label vector conversion matrix, and perform sequence annotation classification on each single character in the candidate label according to the sequence label vector to obtain a sequence annotation result corresponding to each single character in the candidate label.
10. The device according to claim 9, wherein The classification module is further configured to perform sequence annotation classification on each single character in the candidate label according to the sequence label vector, determine the category probability that each single character belongs to each preset sequence label, and obtain a sequence annotation result corresponding to each single character in the candidate label according to the category probability.
11. The device according to any one of claims 7-10, characterized in that, The sequence annotation results corresponding to each single character in the candidate tags are obtained through a text tag matching model; the device further includes a model processing module, and the module processing module is used to obtain an initial text matching model and classification matching training data, and the classification matching training data includes spliced classification tags and classification matching texts corresponding to the classification tags, and the classification tags carry classification sequence tags, and the initial text matching model is trained according to the classification matching training data to obtain an initial text tag matching model, and tag matching training data is obtained, and the tag matching training data includes spliced training tags and tag matching texts matching the training tags, and the training tags carry training sequence tags, and the initial text tag matching model is trained according to the tag matching training data to obtain a trained text tag matching model.
12. The device according to claim 11, characterized in that, The module processing module is further used to train the initial text tag matching model according to the tag matching training data to obtain a preliminary trained text tag matching model, generate task joint training data according to the classification matching training data and the tag matching training data, and train the preliminary trained text tag matching model according to the task joint training data to obtain a trained text tag matching model.
13. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Sequence labeling method and device and computer equipment
CN111985229A