Sequence labeling method, device, computer equipment, and storage medium

By performing word segmentation and conversion of text sequences, generating identification sequences and inputting sequence annotation model, analyzing and combining tags to obtain multiple second tags, and annotating fields, the problem that traditional methods cannot label multiple tags to fields is solved, and more diverse tag detection is achieved.

CN114328837BActive Publication Date: 2025-05-23SUZHOU LANGDONG NET TEC CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111654465.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-05-23
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

Traditional sequence annotation methods cannot place multiple tags on a field, resulting in a relatively single tag detection method and cannot meet the needs of complex text information extraction.

Method used

By obtaining the text sequence, performing word segmentation processing and conversion, generating an identification sequence, and inputting it into the sequence annotation model, obtaining the first label of the field. When the first tag contains a combined tag, the combined tag is parsed to obtain a plurality of second tags, and the fields are marked according to these tags to ensure compliance with the tag logical relationship.

Benefits of technology

It realizes the generation of combined labels for fields in text sequences and uses multiple labels to annotate fields, which enhances the diversity of label detection methods of the sequence labeling model and can handle complex text information more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114328837B_ABST
    Figure CN114328837B_ABST
Patent Text Reader

Abstract

The present application relates to a sequence labeling method, apparatus, computer equipment, and storage medium. The method comprises: obtaining a text sequence, converting the text sequence, and obtaining an identification sequence corresponding to the text sequence; inputting the identification sequence into a sequence labeling model to obtain a first label corresponding to a field in the text sequence; when the first label includes a combined label, parsing the combined label to obtain a plurality of second labels corresponding to the combined label; and labeling the field according to the plurality of second labels. Compared with the traditional sequence labeling method in which only one label can be used to label a field, the present method can generate a combined label for the field in the text sequence, and use the plurality of second labels obtained by parsing the combined label to label the field in the text sequence, thereby making the label detection method of the sequence labeling model more diverse.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer application technology, and in particular to a sequence labeling method, apparatus, computer equipment, storage medium and computer program product. Background Art

[0002] Information extraction technology is a technology that extracts some field information (such as entities, key fact descriptions, etc.) from natural language text. To extract information from natural language text, the natural language text must first be sequence labeled.

[0003] In traditional technology, text can be input into a sequence labeling model, and the text can be processed by the sequence labeling model to generate a label sequence corresponding to the text. The label sequence is then decoded to obtain a label result corresponding to each field in the text, and the label result is used to perform sequence labeling on the text. However, using the sequence labeling method in traditional technology, the label output by the sequence labeling model for each field is a single label, which has the problem of a relatively simple label detection method and inability to label a field with multiple labels. Summary of the invention

[0004] Based on this, it is necessary to provide a sequence labeling method, apparatus, computer equipment, storage medium and computer program product that can label fields with multiple labels in order to address the above technical problems.

[0005] In a first aspect, the present application provides a sequence annotation method. The method comprises:

[0006] Acquire a text sequence, perform word segmentation processing on the text sequence to obtain a plurality of word segmentation characters, and convert each of the word segmentation characters to obtain an identification sequence corresponding to the text sequence;

[0007] Inputting the identification sequence into a sequence labeling model to obtain a first label corresponding to a field in the text sequence;

[0008] When the first tag includes a combination tag, the combination tag is parsed to obtain multiple second tags corresponding to the combination tag, the field is labeled according to the multiple second tags, and the multiple second tags corresponding to the combination tag conform to a preset tag logical relationship.

[0009] In one of the embodiments, when the target field appears multiple times in the text sequence, the method further includes:

[0010] Determining whether the relationship between the multiple first tags corresponding to the target field conforms to the tag logical relationship;

[0011] When the relationship between the plurality of first tags conforms to the tag logical relationship, the plurality of first tags corresponding to the target field in the text sequence are accepted.

[0012] In one embodiment, the method further comprises:

[0013] When the relationship between the plurality of first tags does not conform to the tag logical relationship, the plurality of first tags corresponding to the target field are deleted.

[0014] In one embodiment, inputting the identification sequence into a sequence labeling model to obtain a first label corresponding to a field in the text sequence includes:

[0015] Inputting the identification sequence into the sequence labeling model to generate a label sequence corresponding to the identification sequence, wherein the labels in the label sequence carry label identifications;

[0016] The tag sequence is decoded according to the tag identifier to obtain the first tag corresponding to the field in the text sequence.

[0017] In one embodiment, the tag identifier includes a start identifier and a non-start identifier; and decoding the tag sequence according to the tag identifier to obtain the first tag corresponding to the field in the text sequence includes:

[0018] Starting from the first start identifier of the tag sequence, searching for a group of adjacent start identifiers and non-start identifiers in sequence to obtain multiple identifier groups;

[0019] A field is generated according to a partial text sequence corresponding to the identification group, and the first label corresponding to the field is generated according to the label corresponding to the identification group.

[0020] In one embodiment, obtaining the text sequence includes:

[0021] Get the original text sequence;

[0022] When the text length of the original text sequence is greater than a threshold, the original text sequence is segmented into sentences to obtain a plurality of text sentences;

[0023] According to the text sentence length of each of the text sentences, the multiple text sentences are divided to obtain multiple text sequences, wherein the text length of each of the text sequences is less than the threshold, and there are overlapping text sentences between two adjacent text sequences.

[0024] In a second aspect, the present application also provides a sequence labeling device. The device comprises:

[0025] The identification sequence generation module is used to obtain a text sequence, perform word segmentation processing on the text sequence to obtain a plurality of word segmentation characters, and convert each of the word segmentation characters to obtain an identification sequence corresponding to the text sequence;

[0026] A first label acquisition module, used to input the identification sequence into a sequence labeling model to obtain a first label corresponding to a field in the text sequence;

[0027] A field labeling module is used to parse the combined tag to obtain multiple second tags corresponding to the combined tag when the first tag includes the combined tag, and to label the field according to the multiple second tags, so that the multiple second tags corresponding to the combined tag conform to a preset tag logical relationship.

[0028] In a third aspect, the present application further provides a computer device, wherein the computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the sequence labeling method described in any one of the embodiments of the first aspect is implemented.

[0029] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the sequence labeling method described in any one of the embodiments of the first aspect is implemented.

[0030] In a fifth aspect, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the sequence labeling method described in any one of the embodiments of the first aspect is implemented.

[0031] The above-mentioned sequence labeling method, apparatus, computer equipment, storage medium and computer program product convert the text sequence to obtain an identification sequence corresponding to the text sequence, input the identification sequence into the sequence labeling model, obtain the first label corresponding to the field in the text sequence, and when the first label includes a combined label, parse the combined label to obtain multiple second labels corresponding to the combined label, label the field according to the multiple second labels, and can generate a combined label for the field in the text sequence, and use the multiple second labels obtained by parsing the combined label to label the field in the text sequence. Therefore, compared with the traditional sequence labeling method that can only use one label to label the field, the sequence labeling method provided by the present application can obtain multiple labels corresponding to the field, and use multiple labels to label the field, so that the label detection method of the sequence labeling model is more diverse. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 A schematic diagram of a sequence labeling method in one embodiment;

[0033] Figure 2 A schematic diagram of a flow chart of a text sequence acquisition step in an embodiment;

[0034] Figure 3 is a schematic flow chart of a sequence labeling method in another embodiment;

[0035] Figure 4 is a structural block diagram of a sequence labeling device in one embodiment;

[0036] Figure 5 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0038] In one embodiment, Figure 1 As shown, a sequence labeling method is provided. This embodiment uses the method applied to a server as an example. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart TVs, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented as an independent server or a server cluster consisting of multiple servers.

[0039] In this embodiment, the method is applied to a case where there is a tag logical relationship between multiple tags, and includes the following steps:

[0040] Step S102, obtaining a text sequence, performing word segmentation processing on the text sequence to obtain a plurality of word segmentation characters, and converting each of the word segmentation characters to obtain an identification sequence corresponding to the text sequence.

[0041] Among them, the text sequence can be used to represent an unlabeled text fragment composed of multiple characters. The identification sequence can be obtained by converting the text sequence according to the preset mapping relationship between characters and identifications. For example, the character "A" corresponds to the identification "xxy", and the character "B" corresponds to the identification "xxz", then the identification sequence corresponding to the text sequence "AAB" can be "xxy-xxy-xxz". In one example, the identification sequence can be an id sequence obtained by converting the text sequence according to the segmentation vocabulary.

[0042] Specifically, the server pre-stores a plurality of mapping relationships between characters and identifiers (e.g., a mapping relationship between characters and identifiers in a word segmentation vocabulary). The server responds to the sequence labeling request and obtains a text sequence. The text sequence is segmented to obtain a plurality of segmentation characters. According to the mapping relationship between the character and the identifier, each segmentation character is converted to obtain an identifier sequence corresponding to the text sequence. The sequence labeling request may be manually triggered by a user, for example, a user clicks a corresponding sequence labeling button on a page to trigger a sequence labeling request for a text sequence; or it may be automatically triggered by a server, for example, when the server detects the existence of a text sequence, it automatically triggers a sequence labeling request for a text sequence to obtain the text sequence.

[0043] Step S104: input the identification sequence into a sequence labeling model to obtain a first label corresponding to a field in the text sequence.

[0044] The sequence labeling model may be a language representation model, such as a BERT model (Bidirectional Encoder Representation from Transformers, a self-encoding language model), a bidirectional LSTM model (Bi-directional LSTM, a bidirectional long short-term memory network model), an XLNet model (an autoregressive language model), an ERNIE model (Enhanced Representation from kNowledge IntEgration, a knowledge enhanced semantic representation model), etc. In one embodiment, by using a text sequence sample carrying a label to train a preliminarily trained language representation model, and using the trained language representation model as a trained sequence labeling model, the cost and time of sequence labeling model training can be reduced, and the sequence labeling model can achieve a higher sequence labeling accuracy.

[0045] Specifically, a trained sequence labeling model is pre-deployed in the server. The identifier sequence is input into the sequence labeling model, and the correlation between each identifier in the identifier sequence and multiple tags is determined by the sequence labeling model, and the tag with the highest correlation is used as the tag corresponding to the identifier. According to the tag corresponding to each identifier, a first tag corresponding to each field in the text sequence is generated.

[0046] Step S106: when the first tag includes a combined tag, the combined tag is parsed to obtain a plurality of second tags corresponding to the combined tag, and the field is labeled according to the plurality of second tags.

[0047] Among them, in the embodiment of the present application, the label may include a combination label and a single label. The combination label can be used to represent a label formed by combining multiple second labels. For example, the combination label "recipient-defendant" can be generated by combining the second label "recipient" and the second label "defendant". The second label can be used to represent the single label that constitutes the combination label, and the multiple second labels corresponding to the combination label conform to the preset label logical relationship. Among them, the label logical relationship can be set by the user according to the semantic logic between multiple labels. Due to the limitation of the label logical relationship between multiple second labels, the number of combination labels composed of multiple second labels is limited, so that the sequence labeling model provided in the embodiment of the present application and the traditional sequence labeling model have a small difference in cost when used. In an example, when there are a "defendant" label and a "plaintiff" label, since the two labels do not conform to the label logical relationship, the two labels cannot be combined to generate a corresponding combination label.

[0048] Specifically, the server detects the first tag corresponding to each field in the text sequence. When the server determines that there is a combined tag in the first tag corresponding to the field, the server uses the second tag to parse the combined tag and determines the multiple second tags in the combined tag. The field is annotated with multiple second tags as the annotation result corresponding to the field. When the server determines that there is no combined tag in the first tag corresponding to the field, the field is annotated with the first tag as the annotation result corresponding to the field. The annotation result corresponding to each field in the text sequence is used as the sequence annotation result of the text sequence. In one example, further, the server can extract information from the text sequence through the sequence annotation result of the text sequence to obtain field information corresponding to each annotation result.

[0049] In one example, when the server generates a first label corresponding to a field through a sequence labeling model and contains a combined label "recipient-defendant", the combined label is parsed to obtain the second label "recipient" and "defendant", and the field is labeled with "recipient" and "defendant" as the labeling result corresponding to the field.

[0050] In one embodiment, when the first label corresponding to the field generated by the server through the sequence labeling model is a single label, the first label is directly used to label the field.

[0051] In the above sequence labeling method, the text sequence is converted to obtain an identification sequence corresponding to the text sequence, and the identification sequence is input into the sequence labeling model to obtain a first label corresponding to the field in the text sequence. When the first label includes a combined label, the combined label is parsed to obtain multiple second labels corresponding to the combined label, and the field is labeled according to the multiple second labels. It is possible to generate a combined label for the field in the text sequence, and the multiple second labels obtained by parsing the combined label are used to label the fields in the text sequence. Therefore, compared with the traditional sequence labeling method that can only use one label to label the field, the sequence labeling method provided in the present application can obtain multiple labels corresponding to the field, and use multiple labels to label the field, so that the label detection method of the sequence labeling model is more diverse.

[0052] In one embodiment, when a target field appears multiple times in a text sequence, the sequence labeling method further includes: determining whether the relationship between multiple first labels corresponding to the target field conforms to a preset label logical relationship; when the relationship between multiple first labels conforms to the label logical relationship, accepting multiple first labels corresponding to the target field in the text sequence.

[0053] Specifically, the server pre-stores a label logical relationship between multiple labels. When the server determines that the target field appears multiple times in the text sequence, the server obtains the first label corresponding to the target field at each position in the text sequence, and obtains multiple first labels corresponding to the target field. The label logical relationship is used to judge the multiple first labels corresponding to the target field. When the server determines that the relationship between the multiple first labels conforms to the label logical relationship, the server accepts the multiple first labels corresponding to the target field in the text sequence, and annotates the target field according to the multiple first labels.

[0054] In one example, when the first label corresponding to the target field is "defendant" or "recipient", the two first labels conform to the label logical relationship, and the "defendant" and "recipient" labels corresponding to the target field are received, and the "defendant" and "recipient" labels are used as the annotation results of the target field.

[0055] In this embodiment, when the target field appears multiple times in the text sequence, the label logical relationship is used to judge multiple first labels corresponding to the target field, and multiple first labels that meet the label logical relationship are received, so as to improve the accuracy of sequence labeling of the text sequence.

[0056] In one embodiment, the sequence labeling method further includes: when the relationship between the multiple first labels does not conform to the label logical relationship, deleting the multiple first labels corresponding to the target field.

[0057] Specifically, when the server determines that the target field appears multiple times in the text sequence, the first tag corresponding to the target field at each position in the text sequence is obtained to obtain multiple first tags corresponding to the target field. The multiple first tags corresponding to the target field are judged using the tag logical relationship. When the server determines that the multiple first tags corresponding to the target field do not conform to the tag logical relationship, the multiple first tags corresponding to the target field in the text sequence are rejected, such as deleting each first tag corresponding to the target field.

[0058] In one example, when the first labels corresponding to the target field are "defendant" and "plaintiff", the two first labels do not conform to the label logical relationship, and the "defendant" and "plaintiff" labels corresponding to the target field are rejected, and the "defendant" and "plaintiff" labels corresponding to the target field are deleted.

[0059] In one example, when the server determines that the multiple first tags corresponding to the target field do not conform to the tag logical relationship, the multiple first tags corresponding to the target field in the text sequence are rejected, and prompt information is generated and displayed (such as the multiple first tags corresponding to the target field do not conform to the tag logical relationship).

[0060] In this embodiment, by deleting a plurality of first labels that do not conform to a logical relationship, it is possible to avoid logically conflicting annotation results in the sequence annotation results of the text sequence.

[0061] In one embodiment, step S104, inputting the identification sequence into the sequence labeling model to obtain the first label corresponding to the field in the text sequence, includes: inputting the identification sequence into the sequence labeling model, generating a label sequence corresponding to the identification sequence, decoding the label sequence according to the label identification, and obtaining the first label corresponding to the field in the text sequence.

[0062] The tags in the tag sequence carry tag identifiers. The tag identifiers may include, but are not limited to, non-entity identifiers. Tags carrying non-entity identifiers may be used to represent non-noun tags, and there is no need to perform sequence annotation on the fields corresponding to the tags.

[0063] Specifically, the server inputs the identification sequence into the sequence labeling model, determines the correlation between each identification and multiple tags through the sequence labeling model, and establishes a correlation matrix between identification and tags. The tag with the highest correlation with each identification is used as the tag corresponding to the identification, and the correlation matrix is ​​converted into a tag sequence. The tag identification carried by each tag in the tag sequence is detected, and the tags carrying non-entity identification are deleted. According to the deleted tag sequence, the first tag corresponding to the field in the text sequence is obtained.

[0064] In an example, the training process of the sequence labeling model is explained by taking the ERNIE model as an example:

[0065] First, the server obtains multiple text sequence samples with labels, performs word segmentation on the text sequence samples, and uses the vocabulary of the preliminarily trained ERNIE model (the vocabulary stores the mapping relationship between multiple characters and identifiers) to convert the multiple characters obtained after the word segmentation to obtain the identifier corresponding to each character in the text sequence sample, generate the identifier sequence corresponding to the text sequence sample, and use the label corresponding to the field in the text sequence sample as the label corresponding to the identifier sequence. The identifier sequence and the label corresponding to the identifier sequence are input as training data into the preliminarily trained ERNIE model. A layer of fully connected neural network is added to the preliminarily trained ERNIE model, so that the size of the hidden state dimension in the ERNIE model after adding the fully connected neural network is the same as the number of labels (for example, when the number of identifiers in the identifier sequence is 512 and the number of labels is 21, the size of the correlation matrix generated by the sequence labeling model is 512×21). The correlation between the identifier and each label is obtained by adding the ERNIE model after the fully connected neural network is added, and the correlation matrix between the identifier and the label is established. The label with the highest correlation with the identifier in the correlation matrix is ​​used as the predicted label corresponding to the identifier. The difference between the label corresponding to the identification sequence and the predicted label is determined, and the weight parameters of the ERNIE model after adding the fully connected neural network are adjusted according to the difference until the difference meets the preset conditions to obtain the trained ERNIE model.

[0066] In this embodiment, by generating a label sequence corresponding to the identification sequence, deleting the label carrying the non-entity identifier, and obtaining the first label corresponding to the field in the text sequence according to the deleted label sequence, the amount of data processed by the server can be reduced, thereby improving the efficiency of sequence labeling.

[0067] In one embodiment, the tag identifier includes a start identifier and a non-start identifier. Decoding the tag sequence according to the tag identifier to obtain a first tag corresponding to a field in the text sequence includes: starting from the first start identifier of the tag sequence, searching for a group of adjacent start identifiers and non-start identifiers in sequence to obtain multiple identifier groups; generating a field in the text sequence according to a partial identifier sequence corresponding to the identifier group; and generating a first tag corresponding to the field according to the tag corresponding to the identifier group.

[0068] Specifically, the server starts from the first tag carrying the start identifier in the tag sequence, and searches for several non-start identifiers adjacent to the current start identifier and after the current start identifier in turn, and uses the current start identifier and the corresponding non-start identifier as an adjacent group of start identifiers and non-start identifiers to obtain multiple identifier groups, and the content of the tags in each identifier group is consistent. According to the partial identifier sequence corresponding to each identifier group, multiple characters in the text sequence are combined to obtain a field corresponding to each identifier group. The tag content corresponding to the identifier group is used as the first tag corresponding to the field.

[0069] In one example, the server may use multiple non-starting identifiers between the first starting identifier and the second starting identifier as non-starting identifiers adjacent to the first starting identifier to obtain an identifier group corresponding to the first starting identifier.

[0070] In this embodiment, the label sequence is decoded by the start identifier and the non-start identifier to determine multiple identifier groups, and multiple characters are combined using the identifier groups to obtain fields corresponding to each identifier group. The label content corresponding to the identifier group is used as the first label corresponding to the field, which can improve the accuracy of sequence labeling of the fields.

[0071] In one embodiment, Figure 2 As shown, step S102, obtaining a text sequence, converting the text sequence, and obtaining an identification sequence corresponding to the text sequence, includes:

[0072] Step S202, obtaining an original text sequence.

[0073] Step S204: when the text length of the original text sequence is greater than a threshold, the original text sequence is segmented into sentences to obtain a plurality of text sentences.

[0074] Step S206, dividing the multiple text sentences according to the text sentence length of each text sentence to obtain multiple text sequences.

[0075] The text length of each text sequence is less than a threshold, and there are overlapping text sentences between two adjacent text sequences.

[0076] Specifically, a threshold value of text length is pre-stored in the server. The server obtains an original text sequence and determines the text length of the original text sequence. The text length of the original text sequence is compared with the threshold value. When the server determines that the text length of the original text sequence is greater than the threshold value, the original text sequence is segmented to obtain a plurality of text sentences. Starting from the first text sentence in the original text sequence, the text lengths of the text sentences are sequentially superimposed in the order of the text sentences until the first text length of the superimposed multiple text sentences is less than the threshold value, and the sum of the first text length and the text length of the next text sentence is greater than the threshold value, and the superimposed multiple text sentences are regarded as a text sequence. Starting from the text sentence at the end of the text sequence, the above operation is repeated to generate a new text sequence until the last text sentence in the original text sequence is processed to obtain a plurality of text sequences.

[0077] In this embodiment, the original text sequence is segmented into sentences and the text sentences are divided according to their sentence lengths to obtain multiple text sequences. In addition, there are overlapping text sentences between two adjacent text sequences. This can increase the overlapping parts between adjacent text sequences, so that the sequence labeling model can learn contextual information, thereby improving the accuracy of sequence labeling.

[0078] In one embodiment, Figure 3 As shown, a sequence labeling method is provided, including:

[0079] Step S302, obtaining an original text sequence, and when the text length of the original text sequence is greater than a threshold, dividing the original text sequence into sentences to obtain a plurality of text sentences.

[0080] Step S304, dividing the multiple text sentences according to the text sentence lengths to obtain multiple text sequences, converting each text sequence to obtain an identification sequence corresponding to each text sequence.

[0081] Specifically, the server obtains an original text sequence, determines the text length of the original text sequence, and when the text length of the original text sequence is greater than a threshold, the original text sequence is segmented to obtain multiple text sentences. According to the text sentence lengths of the text sentences, multiple text sentences are divided, and the divided multiple text sentences are used as a text sequence to obtain multiple text sequences. Each text sequence is segmented, and the multiple segmentation characters obtained after the segmentation are converted according to the mapping relationship between the characters and the identifiers to obtain an identifier sequence corresponding to each text sequence. The specific text sequence generation operation and the identifier sequence generation operation can be implemented with reference to the text sequence generation method and the identifier sequence generation method provided in the above embodiments, which will not be specifically elaborated here.

[0082] Step S306: Input the identification sequence into the sequence labeling model to generate a label sequence corresponding to the identification sequence, and determine multiple identification groups in the label sequence according to the label identification.

[0083] Step S308: Generate fields in the text sequence according to the partial identification sequence corresponding to the identification group, and generate a first label corresponding to the field according to the label corresponding to the identification group.

[0084] Specifically, the server sequentially inputs each identification sequence into the sequence labeling model to obtain the correlation between the identification and the label, determines the label corresponding to the identification according to the correlation, and generates a label sequence corresponding to the identification sequence. Decode the label sequence according to the label identification carried by the labels in the label sequence to determine multiple identification groups in the label sequence. Determine multiple characters corresponding to the identification group according to the partial identification sequence corresponding to the identification group, and combine the multiple characters to obtain the field corresponding to the identification group. Use the label corresponding to the identification group to generate a first label corresponding to the field.

[0085] In one example, the text sequence obtained by the server is "Company A is in Place B". The text sequence is segmented to obtain multiple segmented characters "A", "com", "pany", "is", "in", "Place", "B". Each segmented character is converted using the mapping relationship between the character and the identification to obtain an identification sequence corresponding to the text sequence. The identification sequence is input into the sequence labeling model to generate a label sequence corresponding to the identification sequence as "Start identification - recipient, Non - start identification - recipient, Non - start identification - recipient, Non - entity identification, Start identification - location, Non - start identification - location". Delete the labels carrying non - entity identification in the label sequence. Take "Start identification - recipient, Non - start identification - recipient, Non - start identification - recipient" as the first identification group, and take "Start identification - location, Non - start identification - location" as the second identification group. Combine the multiple characters corresponding to the first identification group to obtain the field "Company A" corresponding to the first identification group, and take the label content "recipient" corresponding to the first identification group as the first label of the field "Company A". Combine the multiple characters corresponding to the second identification group to obtain the character "Place B" corresponding to the second identification group, and take the label content "location" corresponding to the second identification group as the first label of the character "Place B".

[0086] Step S310: When the number of times a field appears in the text sequence is multiple, determine whether the relationship between the multiple first labels corresponding to the field conforms to the label logical relationship, and receive the multiple first labels that conform to the label logical relationship.

[0087] Step S312: When the first label includes a combined label, parse the combined label to obtain multiple second labels corresponding to the combined label, and label the field according to the multiple second labels.

[0088] Specifically, when the server determines that a field appears multiple times in a text sequence, multiple first tags corresponding to the field are obtained, and it is determined whether the relationship between the multiple first tags corresponding to the field conforms to the tag logical relationship. When the server determines that the relationship between the multiple first tags conforms to the tag logical relationship, each first tag corresponding to the field is received. When the first tag includes a combined tag, the combined tag is parsed to obtain multiple second tags corresponding to the combined tag, and the field is annotated with the first tag of the non-combination tag and the multiple second tags. When the server determines that the relationship between the multiple first tags does not conform to the tag logical relationship, each first tag corresponding to the field is deleted.

[0089] In this embodiment, by sentence segmenting the original text sequence, a text sequence in which there are overlapping text sentences between multiple adjacent text sequences is obtained, which enables the sequence labeling model to learn the context information in the original text sequence, thereby improving the accuracy of sequence labeling; by converting the text sequence, the identification sequence obtained after the conversion is input into the sequence labeling model, a label sequence corresponding to the identification sequence is generated, the label sequence is decoded, and a first label corresponding to the field in the text sequence is generated. When the field appears multiple times, the multiple first labels corresponding to the field are verified using the label logical relationship, and multiple first labels that conform to the label logical relationship are received, which can improve the accuracy of sequence labeling of the field; compared with the traditional sequence labeling method in which only one label can be used to label the field, the sequence labeling method provided in the present application parses the combined label and uses the multiple second labels obtained after the parsing to label the field, which can make the label detection method of the sequence labeling model more diverse.

[0090] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0091] Based on the same inventive concept, the embodiment of the present application also provides a sequence labeling device for implementing the sequence labeling method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more sequence labeling device embodiments provided below can refer to the limitations on the sequence labeling method above, and will not be repeated here.

[0092] In one embodiment, Figure 4 As shown, a sequence labeling device 400 is provided, comprising: an identification sequence generating module 402, a first label obtaining module 404 and a field labeling module 406, wherein:

[0093] The identification sequence generation module 402 is used to obtain a text sequence, perform word segmentation processing on the text sequence to obtain a plurality of word segmentation characters, and convert each word segmentation character to obtain an identification sequence corresponding to the text sequence.

[0094] The first label acquisition module 404 is used to input the identification sequence into the sequence labeling model to obtain the first label corresponding to the field in the text sequence.

[0095] The field labeling module 406 is used to parse the combined tag to obtain multiple second tags corresponding to the combined tag when the first tag includes the combined tag, and label the field according to the multiple second tags, so that the multiple second tags corresponding to the combined tag conform to the preset tag logical relationship.

[0096] In one embodiment, when a target field appears multiple times in a text sequence, the sequence labeling device 400 includes: a first label verification module, used to determine whether the relationship between multiple first labels corresponding to the target field conforms to a preset label logical relationship; when the relationship between the multiple first labels conforms to the label logical relationship, accept the multiple first labels corresponding to the target field in the text sequence.

[0097] In one embodiment, the first tag checking module is further used to: when the relationship between the multiple first tags does not conform to the tag logical relationship, delete the multiple first tags corresponding to the target field.

[0098] In one embodiment, the first label acquisition module 404 includes: a label sequence generation unit, which is used to input the identification sequence into the sequence annotation model to generate a label sequence corresponding to the identification sequence, and the label in the label sequence carries the label identification; a label sequence decoding unit, which is used to decode the label sequence according to the label identification to obtain the first label corresponding to the field in the text sequence.

[0099] In one embodiment, the label identifier includes a start identifier and a non-start identifier, and the label sequence decoding unit includes: an identifier group generation subunit, which is used to start from the first start identifier of the label sequence and sequentially search for a group of adjacent start identifiers and non-start identifiers to obtain multiple identifier groups; a field generation subunit, which is used to generate a field in a text sequence according to a partial identifier sequence corresponding to the identifier group; and a first label generation subunit, which is used to generate a first label corresponding to the field according to the label corresponding to the identifier group.

[0100] In one embodiment, the identification sequence generation module 402 includes: an original text sequence processing unit, which is used to obtain an original text sequence, and when the text length of the original text sequence is greater than a threshold, the original text sequence is divided into sentences to obtain multiple text sentences; a text sequence generation unit, which is used to divide the multiple text sentences according to the text sentence length of each text sentence to obtain multiple text sequences, wherein the text length of each text sequence is less than the threshold, and there are overlapping text sentences between two adjacent text sequences.

[0101] Each module in the above sequence labeling device can be implemented in whole or in part by software, hardware and a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each module above.

[0102] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store a threshold value of text length. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a sequence labeling method is implemented.

[0103] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0104] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0105] In one embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0106] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0107] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0108] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0109] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0110] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A sequence labeling method, It is characterized in that The method comprises: Acquire a text sequence, perform word segmentation processing on the text sequence to obtain a plurality of word segmentation characters, and convert each of the word segmentation characters to obtain an identification sequence corresponding to the text sequence; Inputting the identification sequence into a sequence labeling model to obtain a first label corresponding to a field in the text sequence includes: Inputting the identification sequence into the sequence labeling model to generate a label sequence corresponding to the identification sequence, wherein the labels in the label sequence carry label identifiers, and the label identifiers include a start identifier and a non-start identifier; Starting from the first start identifier of the tag sequence, searching for a group of adjacent start identifiers and non-start identifiers in sequence to obtain multiple identifier groups; Generate a field in the text sequence according to a partial identification sequence corresponding to the identification group; generating the first label corresponding to the field according to the label corresponding to the identification group; When the first tag includes a combination tag, the combination tag is parsed to obtain multiple second tags corresponding to the combination tag, the field is labeled according to the multiple second tags, and the multiple second tags corresponding to the combination tag conform to a preset tag logical relationship.

2. The method according to claim 1, It is characterized in that When the target field appears multiple times in the text sequence, the method further includes: Determining whether the relationship between the multiple first tags corresponding to the target field conforms to the tag logical relationship; When the relationship between the plurality of first tags conforms to the tag logical relationship, the plurality of first tags corresponding to the target field in the text sequence are accepted.

3. The method according to claim 2, It is characterized in that The method further comprises: When the relationship between the plurality of first tags does not conform to the tag logical relationship, the plurality of first tags corresponding to the target field are deleted.

4. The method according to any one of claims 1 to 3, It is characterized in that The obtaining of the text sequence comprises: Get the original text sequence; When the text length of the original text sequence is greater than a threshold, the original text sequence is segmented into sentences to obtain a plurality of text sentences; According to the text sentence length of each of the text sentences, the multiple text sentences are divided to obtain multiple text sequences, wherein the text length of each of the text sequences is less than the threshold, and there are overlapping text sentences between two adjacent text sequences.

5. A sequence labeling device, It is characterized in that The device comprises: The identification sequence generation module is used to obtain a text sequence, perform word segmentation processing on the text sequence to obtain a plurality of word segmentation characters, and convert each of the word segmentation characters to obtain an identification sequence corresponding to the text sequence; A first label acquisition module, used to input the identification sequence into a sequence labeling model to obtain a first label corresponding to a field in the text sequence; A field labeling module, configured to, when the first label includes a combination label, parse the combination label to obtain a plurality of second labels corresponding to the combination label; label the field according to the plurality of second labels, wherein the plurality of second labels corresponding to the combination label conform to a preset label logical relationship; Wherein, the first label acquisition module includes: A label sequence generating unit, configured to input the identification sequence into the sequence labeling model, and generate a label sequence corresponding to the identification sequence, wherein the labels in the label sequence carry label identifiers, and the label identifiers include a start identifier and a non-start identifier; The tag sequence decoding unit comprises: An identification group generating subunit is used to search for a group of adjacent starting identifications and non-starting identifications in sequence starting from the first starting identification of the tag sequence to obtain a plurality of identification groups; A field generation subunit, used to generate fields in the text sequence according to a partial identification sequence corresponding to the identification group; The first label generating subunit is used to generate the first label corresponding to the field according to the label corresponding to the identification group.

6. The device according to claim 5, It is characterized in that When the target field appears multiple times in the text sequence, the device further includes: The first label verification module is used to determine whether the relationship between the multiple first labels corresponding to the target field conforms to the label logical relationship; when the relationship between the multiple first labels conforms to the label logical relationship, accept the multiple first labels corresponding to the target field in the text sequence.

7. The device according to claim 6, It is characterized in that The first tag checking module is further configured to delete the first tags corresponding to the target field when the relationship between the first tags does not conform to the tag logical relationship.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program. It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

10. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Sequence labeling method and system and computer equipment

    CN111222317A

  • Resume data information analysis and matching method and device, electronic equipment and medium

    CN111428488A