Tag identification method, device, electronic device and storage medium
By processing character encoding results in text content in parallel multi-layer perceptrons, the problem of inaccurate label prediction in the prior art is solved, and more accurate label type prediction is achieved.
Patent Information
- Application Number
- CN202111520187.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2041-12-13
AI Technical Summary
In the prior art, when extracting labels corresponding to proper nouns in text content, problems often occur in the type of predicted labels are inaccurate.
By obtaining the input text and splitting it into characters, after encoding processing, the encoding result is input into the parallel multi-layer perceptron. Each multi-layer perceptron has its own corresponding label, and the dimensions of the output vector correspond one by one to the characters, thereby obtaining the dimensions where the score value exceeds the preset value to determine the label.
The accuracy of label type prediction of nouns represented by characters is improved, making the generation process of output vectors more targeted.
Smart Images

Figure CN114254075B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to a tag identification method, device, electronic device and storage medium. Background Art
[0002] In the prior art, when faced with text content in a specific technical field, in order to speed up the acquisition of information, technical means are usually used to extract the classification of proper nouns in the text content, and the classification can also be called a label.
[0003] However, when extracting labels corresponding to proper nouns in text content, the prior art often encounters the problem of inaccurate predicted label types. Summary of the invention
[0004] The embodiments of the present application provide a tag recognition method, device, electronic device and storage medium, which can improve the problem of inaccurate tag prediction of proper nouns in text content in the prior art.
[0005] The present invention provides a tag identification method, which includes:
[0006] Obtaining input text, the input text comprising a first number of characters;
[0007] Performing encoding processing on each of the characters to obtain a corresponding encoding result;
[0008] Inputting the first number of encoding results into each of the second number of multilayer perceptrons to obtain an output vector output by each of the multilayer perceptrons, wherein the second number of multilayer perceptrons are connected in parallel, each of the multilayer perceptrons has a corresponding label, and the dimension of each of the output vectors is the first number, and the first number of dimensions corresponds to the first number of characters one by one;
[0009] For each output vector of the second number of output vectors, at least one dimension whose score value exceeds the preset score value is obtained from the first number of dimensions, wherein the at least one dimension has at least one corresponding character, and the label of the at least one character is the label corresponding to the output vector to which the at least one dimension belongs.
[0010] The present application also provides a tag identification device, the device comprising:
[0011] An input text acquisition unit, used to acquire input text, wherein the input text includes a first number of characters;
[0012] A coding processing unit, used for performing coding processing on each of the characters to obtain a corresponding coding result;
[0013] an output vector acquisition unit, configured to input the first number of encoding results into each of the second number of multilayer perceptrons to obtain an output vector output by each of the multilayer perceptrons, wherein the second number of multilayer perceptrons are connected in parallel, each of the multilayer perceptrons has a corresponding label, and the dimension of each of the output vectors is the first number, and the first number of dimensions corresponds to the first number of characters one by one;
[0014] A label determination unit is used to obtain, for each output vector of the second number of output vectors, at least one dimension whose score value exceeds a preset score value from the first number of dimensions, wherein the at least one dimension has at least one corresponding character, and the label of the at least one character is the label corresponding to the output vector to which the at least one dimension belongs.
[0015] In some embodiments, the encoding processing unit includes:
[0016] A model acquisition subunit, used to acquire a label recognition model, wherein the label recognition model includes a CNN and a second number of multi-layer perceptrons connected in parallel;
[0017] The result acquisition subunit is used to encode each of the characters through the CNN to obtain a corresponding encoding result.
[0018] In some embodiments, the apparatus further comprises:
[0019] A training text acquisition unit, used to acquire a training text, wherein the training text includes a third number of training characters;
[0020] A training code acquisition unit, used to perform encoding processing on each of the training characters to obtain a corresponding training code result;
[0021] A training output acquisition unit, used for inputting the third number of training encoding results into each of the second number of multilayer perceptrons to obtain a training output vector output by each of the multilayer perceptrons, each of which has a corresponding label;
[0022] A label declaration information unit, used to obtain label declaration information, wherein the label declaration information is used to declare a label of a target type with label information in the training text;
[0023] A perceptron acquisition unit, used to acquire a target multilayer perceptron corresponding to the label of the target type;
[0024] A loss summing calculation unit, used to calculate the loss according to the output vector output by the target multilayer perceptron, and obtain the sum of the losses;
[0025] The summing and convergence unit is used to obtain a trained label recognition model when the sum of the losses converges.
[0026] In some embodiments, the apparatus further comprises:
[0027] The summing non-convergence unit is used to update the parameters of the CNN and the target multi-layer perceptron according to the sum of the losses when the sum of the losses does not converge, until the sum of the losses converges, thereby obtaining a trained label recognition model.
[0028] In some embodiments, the apparatus further comprises:
[0029] The training text screening unit is used to screen selected text from multiple texts to be screened, wherein the selected text includes nouns with labels belonging to the target type, and the selected text is a training text that is not labeled.
[0030] In some embodiments, the training text screening unit includes:
[0031] A first encoding subunit is used to perform a first encoding process on a fourth number of characters included in the text to be screened to obtain a first encoding result;
[0032] a second encoding subunit, configured to perform a second encoding process on a fifth number of characters included in the target type dictionary to obtain a second encoding result, wherein the target type dictionary includes a plurality of target nouns, the plurality of target nouns all belong to the label of the target type, and the plurality of target nouns include the fifth number of characters in total;
[0033] An interaction construction subunit, used for constructing the interaction between the first encoding result and the second encoding result according to the attention mechanism to obtain an interaction result;
[0034] A two-dimensional vector subunit, used for performing a full connection transformation on the interaction result to obtain a two-dimensional vector result;
[0035] A score value acquisition subunit is used to acquire a score value of a first-dimensional vector in the two-dimensional vector result, wherein the score value of the first-dimensional vector reflects the probability that the text to be filtered includes a noun with a label belonging to the target type;
[0036] The text selection subunit is used to select a preset number of texts to be screened with the highest scores of the first dimensional vector from the multiple texts to be screened, and the preset number of texts to be screened are the selected texts.
[0037] In some embodiments, the apparatus further comprises:
[0038] A model acquisition unit, used to acquire a probability prediction model, wherein the probability prediction model includes a first RoBERTa model, a second RoBERTa model, an attention layer, and a fully connected layer;
[0039] A first encoding subunit, specifically configured to perform a first encoding process on the fourth number of characters using the first RoBERTa model;
[0040] A second encoding subunit, specifically configured to perform a second encoding process on the fifth number of characters using the second RoBERTa model;
[0041] An interaction construction subunit, specifically used to construct the interaction between the first encoding result and the second encoding result by using the attention layer;
[0042] The two-dimensional vector quantum unit is specifically used to use the fully connected layer to perform a fully connected transformation on the interaction result.
[0043] In some embodiments, the apparatus further comprises:
[0044] A result judgment unit, used to judge whether the result of the score value of the first dimensional vector is accurate by using the trained binary classification model;
[0045] A parameter updating unit is used to update the parameters of the first RoBERTa model, the second RoBERTa model, the attention layer and the fully connected layer when the result of the score value of the first dimensional vector is inaccurate.
[0046] In some embodiments, the apparatus further comprises:
[0047] The label screening unit is used to screen labels of the target type from a plurality of labels to be screened, wherein the plurality of labels to be screened correspond to the same target technical field.
[0048] In some embodiments, the tag screening unit includes:
[0049] A test set recognition subunit, used to input a test sample set belonging to the target technical field into an original label recognition model, and obtain label recognition results corresponding to characters in the test sample set;
[0050] A test set parameter determination subunit, configured to determine a first parameter and a second parameter of each of the multiple labels to be filtered according to the label recognition results corresponding to the characters in the test sample set and the labels actually corresponding to the characters in the test sample set, wherein the multiple labels to be filtered are all labels actually corresponding to the characters in the test sample set;
[0051] A training set parameter determination subunit, used to determine the third parameter of each to-be-screened label according to the number of occurrences of each label in all labels actually corresponding to the characters in the training sample set;
[0052] The tag screening subunit is used for determining, for each of the tags to be screened, that the tag to be screened is a tag of the target type if the first parameter, the second parameter and the third parameter of the tag to be screened all meet preset conditions.
[0053] In the label recognition method provided in the embodiment of the present application, an input text can be obtained, and the input text can be split into a first number of characters, and then each character in the first number of characters is encoded to obtain a corresponding encoding result, a total of the first number of encoding results. Subsequently, the first number of encoding results are input as a whole into each of the second number of multilayer perceptrons, so that the output vector output by each multilayer perceptron can be obtained, a total of the second number of output vectors, wherein each multilayer perceptron has its own corresponding different labels; the number of dimensions of each output vector is the first number, and the first number of dimensions corresponds to the first number of characters one by one. For each output vector, the score values of the first number of dimensions can be compared with the preset score values, and at least one dimension whose score value exceeds the preset score value can be obtained. The above-mentioned at least one dimension corresponds to at least one character, and the at least one character is a noun in the input text, and the label to which the noun belongs is the label corresponding to the output vector to which the above-mentioned at least one dimension belongs.
[0054] In the present application, multiple multi-time perceptrons connected in parallel and corresponding to multiple labels can generate an output vector, and the score values of the dimensions in the output vector can be compared with the preset score values, so that the characters corresponding to the dimensions whose score values exceed the preset score values can be obtained, and further the labels to which these characters belong. The present application utilizes multiple multi-time perceptrons connected in parallel and corresponding to multiple labels, so that the generation process of the output vector is more targeted, so that the type of label of the noun represented by the character can be predicted more accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0056] Figure 1 It is a scene schematic diagram of the tag recognition method provided in the embodiment of the present application;
[0057] Figure 2 It is a flowchart of a tag identification method provided by an embodiment of the present application;
[0058] Figure 3 A schematic diagram of a model showing a specific implementation of the tag recognition model in the application stage;
[0059] Figure 4 A model schematic diagram showing a specific implementation of the tag recognition model in the training phase;
[0060] Figure 5 A model schematic diagram showing a specific implementation of a probability prediction model;
[0061] Figure 6 A schematic diagram showing a comparison between the marking method in the prior art and the marking method corresponding to the present application is shown;
[0062] Figure 7 A model schematic diagram showing a specific implementation of the original tag recognition model;
[0063] Figure 8 This is a structural schematic diagram of a tag identification device provided by an embodiment of the present application;
[0064] Fig. 9 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0065] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0066] Embodiments of the present application provide a tag identification method, device, electronic device, and storage medium.
[0067] The tag identification device can be integrated into an electronic device, which can be a terminal, a server, or other device. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop, or a personal computer (PC). The server can be a single server or a server cluster consisting of multiple servers.
[0068] In some embodiments, the tag identification device may also be integrated into multiple electronic devices. For example, the tag identification device may be integrated into multiple servers, and the tag identification method of the present application may be implemented by multiple servers.
[0069] In some embodiments, the server may also be implemented in the form of a terminal.
[0070] For example, the electronic device described above may execute the following method: obtaining input text, the input text including a first number of characters; encoding each of the characters to obtain a corresponding encoding result; inputting the first number of encoding results into each of the second number of multilayer perceptrons to obtain an output vector output by each of the multilayer perceptrons, wherein the second number of multilayer perceptrons are connected in parallel, each of the multilayer perceptrons has its own corresponding label, and the dimension of each of the output vectors is a first number, and the first number of dimensions corresponds one-to-one to the first number of characters; for each of the second number of output vectors, obtaining at least one dimension whose score value exceeds a preset score value from the first number of dimensions, wherein the at least one dimension has a corresponding at least one character, and the label of the at least one character is the label corresponding to the output vector to which the at least one dimension belongs.
[0071] For more information, please see Figure 1 , let's take the input text "I had a headache for three days, and took some medicine for pain relief but it didn't work" as an example. The input text includes 13 characters, namely "head, pain, three, days,,, took, used, powder, benefit, pain, not, seen, effect". Each of the above 13 characters is encoded, and the processing results corresponding to each character can be obtained; for example, "head, pain, three, days,,, took, used, powder, benefit, pain, not, seen, effect" corresponds to "result 1, result 2, result 3, result 4, result 5, result 6, result 7, result 8, result 9, result 10, result 11, result 12, result 13" respectively.
[0072] After obtaining the above 13 processing results, the above 13 processing results can be taken as a whole and input into multilayer perceptron 1, multilayer perceptron 2... multilayer perceptron n. Each multilayer perceptron has its own corresponding label: multilayer perceptron 1 corresponds to label 1, multilayer perceptron 2 corresponds to label 2... multilayer perceptron n corresponds to label n. Each of the n multilayer perceptrons will generate an output vector after outputting the 13 processing results, for a total of n output vectors. Each output vector is a 13-dimensional vector, and the 13 dimensions correspond to the 13 characters one by one.
[0073] For each output vector, the score values of the 13 dimensions of the output vector can be compared with the preset score values respectively, so as to select at least one dimension whose score value exceeds the preset score value. At least one dimension has its own corresponding character, and the label of the character is the label corresponding to the output vector to which the above-mentioned at least one dimension belongs.
[0074] For more information, please see Figure 1 , let's assume that the 1st dimension of the output vector corresponding to label 1 is the dimension whose score value exceeds the preset score value, then the character corresponding to 1st dimension is the "head" in the input text "Headache for three days, taking Sanlitong has not worked", and the label of "head" is label 1. Let's also assume that the 1st and 2nd dimensions of the output vector corresponding to label 2 are both dimensions whose score values exceed the preset score value, then the characters corresponding to 1st and 2nd dimensions are the "head" and "pain" in the input text "Headache for three days, taking Sanlitong has not worked", so it can be regarded that the label of the noun "headache" composed of the characters corresponding to 1st and 2nd dimensions is label 2, and similarly, the label corresponding to the noun "Sanlitong" is label n. Label 1 can be a human body part label, that is, all nouns belonging to the type of human body parts belong to label 1; label 2 can be a symptom label, and label n can be a drug label.
[0075] It should be noted that the serial numbers of the following embodiments are not intended to limit the preferred order of the embodiments.
[0076] In this embodiment, a tag identification method is provided, such as Figure 2 As shown, the tag identification method is applied in a terminal, and the specific process of the method may include the following steps 110 to 140:
[0077] 110. Obtain input text, where the input text includes a first number of characters.
[0078] The input text is the text whose label is to be predicted, and the input text may be text information describing information in a certain professional field. Professional fields are fields that are not easily understood or familiar to ordinary people and have professional technical thresholds, such as the medical field, the legal field, etc.
[0079] The input text may include a first number of characters. The first number is a positive integer, and the specific value of the first number should not be understood as a limitation of the present application. Optionally, after obtaining the input text, the input text may be segmented to obtain the first number of characters.
[0080] 120. Perform encoding processing on each of the characters to obtain a corresponding encoding result.
[0081] For each character in the first number of characters, encoding processing may be performed, so that an encoding result corresponding to each character may be obtained.
[0082] Optionally, in a specific implementation, step 120 may specifically include the following steps 121 to 122:
[0083] 121. Obtain a label recognition model, wherein the label recognition model includes a CNN and a second number of multi-layer perceptrons connected in parallel.
[0084] The label recognition model is a model for label recognition of nouns included in the input text, wherein the nouns included in the input text may be composed of at least one character. The label recognition model may include a convolutional neural network (CNN) and a second number of multilayer perceptrons (MLP), wherein the second number of multilayer perceptrons are connected in parallel. The second number is a positive integer, and the specific value of the second number should not be understood as a limitation on the present application.
[0085] 122. Encode each of the characters through the CNN to obtain a corresponding encoding result.
[0086] CNN can be a gated CNN or a bidirectional gated CNN. The specific type of CNN should not be understood as a limitation to the present application.
[0087] For more information, please see Figure 3 , Figure 3 A specific implementation method of the label recognition model in the application stage is shown. Next, let's take the professional field of medicine, the input text is "Headache for three days, taking Sanlitong has no effect", and the CNN is a bidirectional gated CNN as an example to illustrate the embodiment of this application.
[0088] The encoding process of each character through CNN can specifically include the following process:
[0089] For each of the 13 characters "headache, three, days, ...
[0090] In the above process, splitting each character into an m-dimensional vector can be achieved by the following methods: 1. Initializing each character randomly into an m-dimensional vector; 2. Initializing each character using a preset m-dimensional vector. It should be understood that splitting characters into m-dimensional vectors can also be achieved by other methods, and the specific implementation method should not be understood as a limitation of the present application.
[0091] Set a convolution kernel of n*m, where n is a positive integer less than 13. Subsequently, use this n*m convolution kernel to perform convolution on the above 13*m matrix by moving it from left to right. The step size for each time step is 1, and after moving 13 times, a first convolution result of 13*1 can be obtained; then use this n*m convolution kernel to perform convolution on the above 13*m matrix from right to left. The step size for each time step is 1, and after moving 13 times, a second convolution result of 13*1 can be obtained. Then vertically splice the first convolution result and the second convolution result to obtain a splicing result of 13*2. Each column of this 13*2 is the encoding result corresponding to the corresponding character among 13 characters. For example, the encoding result corresponding to the first character "head" is the first column of the splicing result of 13*2, and the encoding result corresponding to the second character "pain" is the second column of the splicing result of 13*2... and so on.
[0092] 130. Input the first quantity of the encoding results into each of the second quantity of multi-layer perceptrons to obtain an output vector output by each multi-layer perceptron; wherein, the second quantity of multi-layer perceptrons are connected in parallel, each multi-layer perceptron has its own corresponding label, and the dimension of each output vector is the first quantity, and the first quantity of dimensions correspond one by one to the first quantity of characters.
[0093] The first quantity of encoding results as a whole can be respectively input into each multi-layer perceptron, so that the output vectors output by each multi-layer perceptron can be obtained. Since the number of multi-layer perceptrons is the second quantity, the number of output vectors is also the second quantity.
[0094] Continue to illustrate with the above example. For details, please refer to Figure 3 , the encoding result is input into multi-layer perceptron 1, and multi-layer perceptron 1 will output an output vector corresponding to label 1; the encoding result is input into multi-layer perceptron 2, and multi-layer perceptron 2 will output an output vector corresponding to label 2... the encoding result is input into multi-layer perceptron n, and multi-layer perceptron n will output an output vector corresponding to label n.
[0095] It should be understood that the corresponding relationship between the output vector and the label can be realized by the multi-layer perceptron that outputs the output vector. Since each multi-layer perceptron has its own corresponding label, the output vector output by the multi-layer perceptron also has a corresponding label.
[0096] Each output vector has the first quantity of dimensions, and each dimension has its own corresponding score value. The first quantity of dimensions correspond one by one to the first quantity of characters. For details, please refer to Figure 3The output vector corresponding to label 1 has 13 dimensions, and the output vector corresponding to label 2 also has 13 dimensions. The i-th dimension among the 13 dimensions of the output vector corresponds to the i-th character among the 13 characters, where i is a positive integer not exceeding a first quantity. In this example, i is a positive integer not exceeding 13.
[0097] 140. For each of the second quantity of output vectors, obtain at least one dimension from the first quantity of dimensions whose score value exceeds a preset score value. Among them, the at least one dimension has at least one corresponding character, and the label of the at least one character is the label corresponding to the output vector to which the at least one dimension belongs.
[0098] For each output vector, the score values of the first quantity of dimensions in the output vector are compared with the preset score value. If the score value exceeds the preset score value, it indicates that the character represented by the corresponding dimension has a higher possibility of belonging to the label corresponding to the output vector. Therefore, it can be determined that the label of the character represented by the dimension whose score value exceeds the preset score value is the label corresponding to the output vector.
[0099] Continuing with the above example, for the output vector corresponding to label n, for example, if the score values of the 8th, 9th, and 10th dimensions in the dimension exceed the preset score value, then the characters corresponding to the 8th, 9th, and 10th dimensions can be obtained, namely "san", "li", "tong", and then it is determined that the label corresponding to the noun "Sanlitong" composed of the above three characters is label n.
[0100] In the above embodiment, multiple multi-perceptrons connected in parallel and corresponding to multiple labels respectively can generate output vectors, and the score values of the dimensions in the output vectors can be compared with the preset score value, so that the characters corresponding to the dimensions whose score values exceed the preset score value can be obtained, and further the labels to which these characters belong can be obtained. This application utilizes multiple multi-perceptrons connected in parallel and corresponding to multiple labels respectively, making the generation process of the output vector more targeted, and thus making the prediction of the type of the label of the noun represented by the character more accurate.
[0101] Optionally, in a specific embodiment, before step 120, the embodiments of the present application may further include the following steps A1 to A8:
[0102] A1. Obtain a training text, where the training text includes a third quantity of training characters.
[0103] The training text here is the text that has been specifically labeled with the target type label. Among them, the target type label is a label that meets one or several characteristics. For example, the target type label can be a label that appears infrequently in the training set; the target type label can also be a label that is still difficult to accurately identify after multiple trainings; the target type label can also be a label that appears infrequently and is difficult to accurately identify. The value of the third number is a positive integer, and the specific numerical value of the third number should not be understood as a limitation on this application.
[0104] For more information, please see Figure 4 , Figure 4 A specific implementation of the label recognition model in the training phase is shown. It can be seen that Figure 4 The training phase shown is similar to Figure 3 Compared with the application stage shown in FIG. 1 , the input volume includes not only the training text corresponding to the input text but also annotation declaration information. The annotation declaration information will be described in detail when introducing step A4.
[0105] A2. Encode each of the training characters to obtain a corresponding training encoding result.
[0106] The specific implementation method of step A2 is the same as that of step 120, and will not be repeated here.
[0107] A3. Input the third number of training encoding results into each of the second number of multilayer perceptrons to obtain a training output vector output by each of the multilayer perceptrons, wherein each of the multilayer perceptrons has a corresponding label.
[0108] The specific implementation method of step A3 is the same as that of step 130, and will not be described here.
[0109] A4. Obtain annotation declaration information, where the annotation declaration information is used to declare a label of a target type with annotation information in the training text.
[0110] The annotation declaration information is used to declare the type of label that the corresponding training text is specifically annotated with. Specifically, the annotation declaration information is a data string with a data volume of the third number of bits, and each data bit of the data string uniquely corresponds to a label. If the training text is specifically annotated with one or several labels, then one or several data bits of the data string can be set with annotation information, and other data bits of the data string may not be set with the above-mentioned annotation information. Setting the annotation information can be set to one, and not setting the annotation information can be set to zero. For example, let's assume that the training text is specifically annotated with label 1, and the training text may also include characters belonging to other labels, but the other labels have not been annotated by the staff. In this case, the annotation declaration information will be at the data position corresponding to label 1, and the other data bits will be set to zero. For details, please refer to Figure 4 .
[0111] A5. Obtain a target multilayer perceptron corresponding to the label of the target type.
[0112] Since there is a one-to-one correspondence between labels and multiple perceptrons, the corresponding target multiple perceptron can be obtained according to the label of the target type.
[0113] A6. Calculate the loss according to the output vector output by the target multilayer perceptron, and obtain the sum of the losses.
[0114] In the above implementation, when calculating the loss, the calculation can be performed only based on the output vector of the target multilayer perceptron; and for the non-target multilayer perceptron, the corresponding loss can be set to zero. When calculating the sum of the losses, the loss is calculated only based on the output vector of the target multilayer perceptron. If there is only one target multilayer perceptron, the corresponding loss is taken as the sum of the losses.
[0115] The training text is only specifically labeled with target type labels, so there may be training characters in the training text that belong to non-target type labels but are not labeled. If the loss corresponding to the non-target multilayer perceptron is not set to zero, after the label recognition model recognizes the training characters belonging to the non-target type labels, since the non-target type labels are not labeled, errors will be introduced when calculating the loss; this error is generated in the following way: although the non-target type labels are accurately recognized, the loss value is calculated incorrectly because the non-target type labels are not labeled.
[0116] In the above implementation, in the process of training the label recognition model, the label recognition model can be trained to increase the target type of labels, thereby improving the recognition accuracy of the label recognition model for the target type of labels; at the same time, it can effectively avoid the error in the recognition of non-target type labels caused by increasing the training of target type labels. Compared with the prior art of increasing the labeling of target type labels and other types of labels, it saves the staff's working time for labeling other types of labels, and improves the label recognition model in a targeted manner.
[0117] A7. If the sum of the losses converges, a trained label recognition model is obtained.
[0118] A8. If the sum of the losses does not converge, update the parameters of the CNN and the target multilayer perceptron according to the sum of the losses until the sum of the losses converges, thereby obtaining a trained label recognition model.
[0119] If the sum of the losses does not converge, the parameters of the multilayer perceptron corresponding to the labeled label are updated according to the sum of the losses, and the new loss is calculated according to the output vector of the updated target multilayer perceptron output, and the new sum of the losses is obtained. If the new sum of the losses does not converge, it is iterated until the sum of the losses converges to obtain a trained label recognition model.
[0120] In the above implementation, the loss is calculated according to the output vector output by the target multilayer perceptron, which can be specifically implemented in the following manner:
[0121] Let's assume that the label of the target type corresponding to the target multilayer perceptron is label c, then
[0122] According to the formula Calculate the loss value L of label c c Among them, Ω neg represents all unlabeled positions in the output vector corresponding to label c, Ω pos Indicates the position of all labeled labels in the output vector corresponding to label c, s j Represents the score value corresponding to the j-th dimension.
[0123] Optionally, in a specific implementation manner, before step A1, the embodiment of the present application may further include the following steps:
[0124] A selected text is screened from a plurality of texts to be screened, wherein the selected text includes nouns having a label belonging to the target type, and the selected text is a training text that is not labeled.
[0125] The text to be screened is text information to be screened that describes information in a certain professional field. The text to be screened may include nouns with labels belonging to the target type, or may not include nouns with labels belonging to the target type. It should be understood that the selected text screened out from the multiple texts to be screened at this time is a training text that has not been specifically labeled with labels of the target type by the staff. After the selected text is screened out, the next step is for the staff to label the target type.
[0126] In the above-mentioned implementation, the selected text is filtered out from multiple texts to be filtered, and the work of identifying whether the text includes nouns with labels belonging to the target type is performed by a machine, rather than by staff. The staff can directly mark the nouns with labels belonging to the target type in the selected text, which can further reduce the workload of the staff.
[0127] Optionally, the step of “screening selected text from a plurality of texts to be screened, wherein the selected text includes a noun having a label belonging to the target type” specifically includes the following steps B1 to B6:
[0128] For each of the multiple texts to be filtered, the following steps B1 to B5 are performed:
[0129] B1. Perform a first encoding process on a fourth number of characters included in the text to be screened to obtain a first encoding result.
[0130] The value range of the fourth number is a positive integer, and the specific value of the fourth number should not be understood as a limitation on the present application. The first encoding result includes a fourth number of n-dimensional vectors, and the fourth number of n-dimensional vectors correspond one-to-one to the fourth number of characters, and the first encoding result is a matrix of the fourth number × n. Among them, the value of n is a positive integer, for example, n can take values such as 80, 100, 200, etc., and its specific value should not be understood as a limitation on the present application.
[0131] Optionally, step B1 is performed in a probability prediction model, which includes a first RoBERTa model, a second RoBERTa model, an attention layer, and a fully connected layer. For details, see Figure 5 .
[0132] Step B1 may specifically utilize the first RoBERTa model to perform a first encoding process on the fourth number of characters.
[0133] B2. Perform a second encoding process on the fifth number of characters included in the target type dictionary to obtain a second encoding result.
[0134] The target type dictionary includes multiple target nouns, all of which belong to the target type label, and the multiple target nouns include the fifth number of characters. For example, let the target type label be label c, the target type dictionary be dictionary c, there are p nouns in dictionary c, the labels of the p nouns are all label c, and the total number of characters of the p nouns is the fifth number.
[0135] The second encoding result includes a fifth number of n-dimensional vectors, the fifth number of n-dimensional vectors correspond one-to-one to the fifth number of characters, and the second encoding result is a matrix of the fifth number × n. Step B2 can specifically use the second RoBERTa model to perform a second encoding process on the fifth number of characters. For details, see Figure 5 .
[0136] B3. According to the attention mechanism, construct the interaction between the first encoding result and the second encoding result to obtain an interaction result.
[0137] Step B3 can specifically use the attention layer to construct the interaction between the first encoding result and the second encoding result. For details, see Figure 5 .
[0138] Assuming that the fourth number is 13 and the fifth number is 100, the first encoding result is a 13×n matrix with 13 n-dimensional vectors; the second encoding result is a 100×n matrix with 100 n-dimensional vectors. The first encoding result and the second encoding result have a total of 113 n-dimensional vectors. Step B3 specifically includes the following process:
[0139] For each of the 113 n-dimensional vectors, the following steps can be performed:
[0140] Let us take the first n-dimensional vector as an example for explanation. The first n-dimensional vector is any n-dimensional vector among the 113 n-dimensional vectors.
[0141] Calculate the inner product of the first n-dimensional vector and 113 vectors, and get a total of 113 products.
[0142] By performing softmax transformation on these 113 products, we can get the weight values corresponding to the first n-dimensional vector and all 113 n-dimensional vectors.
[0143] Then, the weighted sum of 113 n-dimensional vectors is calculated, and the weighted sum is the interaction result of the first n-dimensional vectors.
[0144] Through the above method, the interaction result of each n-dimensional vector in all 113 n-dimensional vectors can be obtained.
[0145] B4. Perform a full connection transformation on the interaction result to obtain a two-dimensional vector result.
[0146] In step B4, the fully connected layer may be used to perform a fully connected transformation on the interaction result to obtain a two-dimensional vector result. The two-dimensional vector result includes a first-dimensional vector and a second-dimensional vector, wherein the first-dimensional vector is used to describe the possibility that the text to be screened includes nouns with labels belonging to the target type, and the second-dimensional vector is used to describe the possibility that the text to be screened does not include nouns with labels belonging to the target type.
[0147] B5. Obtain the score value of the first dimension vector in the two-dimensional vector result.
[0148] The score value of the first dimensional vector reflects the probability that the text to be filtered includes a noun with a label belonging to the target type. The higher the score value, the greater the probability that the text to be filtered includes a noun with a label belonging to the target type.
[0149] By performing the above steps B1 to B5 for each of the multiple texts to be screened, the score value of the first dimension vector corresponding to each of the multiple texts to be screened can be obtained. Optionally, after obtaining the score value of the first dimension vector corresponding to each of the multiple texts to be screened, step B6 is performed:
[0150] B6. Select a preset number of texts to be screened with the highest scores of the first dimensional vectors from the multiple texts to be screened, and the preset number of texts to be screened are the selected texts.
[0151] In step B6, after obtaining the score value of the first dimension vector corresponding to each text to be filtered, the multiple texts to be filtered can be sorted in descending order according to the size of the score value, and then a preset number of texts to be filtered are intercepted and used as the selected texts.
[0152] In the above implementation, a preset number of to-be-screened texts with the highest scores of the first-dimensional vectors can be used as selected texts. A high score of the first-dimensional vector means that the to-be-screened texts have a high probability of including nouns with labels belonging to the target type, so that the selected texts obtained by the staff have a high probability of including nouns with labels belonging to the target type, so that the staff can focus on the labeling work of the target type labels, reducing the screening work of the labeled texts and improving work efficiency.
[0153] For more information, please see Figure 6 , Figure 6The figure shows the comparison between the traditional annotation method (i.e., the annotation method in the prior art) and the proposed annotation method (i.e., the annotation method corresponding to the present application). Let "dizziness" be a noun belonging to the label of the target type, and the label of the target type is "adverse reaction". Figure 6 It can be seen that in the labeling method of the prior art, not only the noun "dizziness" of the target type label "adverse reaction" needs to be labeled, but also other types of labels need to be labeled, for example, "head" is labeled as a "part" label, "headache" is labeled as a "symptom" label, "three days" is labeled as a "duration" label, "Sanlitong" is labeled as a "drug" label, etc.
[0154] As for the labeling method corresponding to this application, the staff only needs to label the noun "dizziness" of the target type label "adverse reaction", and the nouns belonging to other types of labels do not need to be labeled, which greatly improves the work efficiency of the staff.
[0155] Optionally, in one implementation, after step B5, the embodiment of the present application may further include the following steps:
[0156] The trained binary classification model is used to determine whether the score value of the first dimensional vector is accurate; if inaccurate, the parameters of the first RoBERTa model, the second RoBERTa model, the attention layer, and the fully connected layer are updated.
[0157] In the above-mentioned implementation, after calculating the score value of the first-dimensional vector in step B5, the trained binary classification model can be used to determine whether the score value result of the first-dimensional vector is accurate. If it is inaccurate, the parameters of the probability prediction model can be updated based on the judgment result of the binary classification model, thereby continuously improving the probability prediction model.
[0158] In a specific implementation, the probability prediction model can be trained in the following manner:
[0159] Let's take label c as an example. First, we obtain a batch of text segments labeled with label c from the labeled text, and then use this batch of text segments and dictionary c to construct a batch of positive samples. Then, we obtain a batch of text segments that do not contain label c from the labeled text, and use this batch of text segments and dictionary c to construct a batch of negative samples. Then, we use the constructed positive and negative samples to train the probability prediction model.
[0160] Optionally, in a specific implementation, before the step of “filtering the selected text from the multiple texts to be filtered”, the embodiment of the present application may further include the following steps:
[0161] Filter the tags of the target type from multiple tags to be filtered, where the multiple tags to be filtered correspond to the same target technical field.
[0162] As described above, the tags of the target type can be tags with a low frequency of occurrence in the training set; the tags of the target type can also be tags that are still difficult to be accurately recognized after multiple trainings; the tags of the target type can also be tags that have both a low frequency of occurrence and are difficult to be accurately recognized.
[0163] For example, taking the tags of the target type as tags that have both a low frequency of occurrence and are difficult to be accurately recognized, and the target technical field as the medical field, the description is as follows: Before filtering the training text, it is necessary to filter out the tags with a low frequency of occurrence and difficult to be accurately recognized from multiple tags to be filtered belonging to the same medical field.
[0164] Optionally, the step of "filtering the tags of the target type from multiple tags to be filtered" may specifically include the following steps C1 to C4:
[0165] C1. Input the test sample set belonging to the target technical field into the original tag recognition model, and obtain the tag recognition results corresponding to the characters in the test sample set.
[0166] The original tag recognition model may include a feature extraction network and a fully connected layer. For details, please refer to Figure 7 . Among them, regarding the training process of the original tag recognition model, the loss of the original tag recognition model can be calculated by the CRF method, and when the loss converges, the trained original tag recognition model is obtained. Input any test sample in the test sample set into the original tag recognition model, and the tag recognition result corresponding to the characters of this test sample can be obtained.
[0167] For example, taking the test sample "Toothache, taking aspirin is effective" as an example for illustration:
[0168] Split the test sample "Toothache, taking aspirin is effective" into 13 characters "tooth, ache, pain, ,, take, use, a, spirin, see, effect", and then use the feature extraction network to encode the above 13 characters to obtain the encoding result.
[0169] Optionally, the feature extraction network can be a CNN. The CNN can be a gated CNN or a bidirectional gated CNN. Let's assume the feature extraction network is a bidirectional gated CNN. The process of encoding 13 characters by the bidirectional gated CNN is the same as the process of encoding each character in step 122, so it will not be elaborated here. It should be understood that the feature extraction network can also be a BILSTM, and the specific type of the feature extraction network should not be regarded as a limitation to this application.
[0170] Input the encoding result into the fully connected layer of the original label recognition model, and then we can obtain the (2×a + 1)-dimensional vector corresponding to each of the 13 characters "toothache, take aspirin, effective". Here, a is the number of labels that the original label recognition model can recognize. Since each label is divided into the beginning part of the label and the non-beginning part of the label, so we take 2×a; and there are characters in the test sample that do not belong to any label, such as ",", so there is also an empty representing not belonging to any label. Therefore, the vector corresponding to each character is a (2×a + 1)-dimensional vector, and each (2×a + 1)-dimensional vector corresponding to each character has its own score value.
[0171] For each of the multiple characters, perform a softmax transformation on the (2×a + 1)-dimensional vector of the character, and then we can obtain the probability value of each character as one of the (2×a + 1) label situations. Take the label situation with the largest probability value as the label of the corresponding character, and then we can obtain the label of the character.
[0172] Repeat the above steps, and then we can realize the recognition of the label for each test sample in the test sample set.
[0173] C2. Determine the first parameter and the second parameter of each to-be-screened label among the multiple to-be-screened labels according to the label recognition result corresponding to the character in the test sample set and the actual label corresponding to the character in the test sample set, where the multiple to-be-screened labels are all the labels actually corresponding to the characters in the test sample set.
[0174] After obtaining the label recognition result of each test sample in the test sample set through step C1, determine the first parameter and the second parameter of each to-be-screened label according to the above label recognition result and the actual label corresponding to the character in the test sample set. Among them, the first parameter and the second parameter can be used to reflect whether the to-be-screened label can be accurately recognized; the first parameter can be the F1 score, and the second parameter can be the confidence level.
[0175] The F1 score of each to-be-screened label can be calculated in the following way:
[0176] For each label to be filtered, the precision and recall of each label to be filtered can be calculated first.
[0177] The precision rate indicates how many of the samples predicted as positive are actually positive samples. There are two cases of samples predicted as positive: samples that are actually positive are predicted as positive (TP), and samples that are actually negative are predicted as positive (FP). Therefore, the precision rate precision = TP / (TP+FP).
[0178] The recall rate indicates how many positive examples in the sample are predicted correctly, and the prediction results of positive examples are as follows: predicting the actual positive sample as positive (TP), and predicting the actual positive sample as negative (FN). Therefore, the recall rate recall = TP / (TP+FN).
[0179] After calculating the precision and recall through the above steps, you can use the following formula Calculate the F1 score.
[0180] The confidence of each label to be filtered can be calculated as follows:
[0181] Let label c be any label to be filtered among multiple labels to be filtered.
[0182] According to the formula Calculate the confidence S c ; Among them, P k is the probability value of the character in the test sample set that is predicted to be labeled c and whose actual label is indeed c at label c, ∑ L P k is the sum of the probability values of the x characters in the test sample set that are predicted to have label c and whose actual label is indeed c at label c, and L is the total number of characters in the test sample set that are predicted to have label c and whose actual label is indeed c.
[0183] C3. Determine the third parameter of each to-be-screened label according to the number of occurrences of each label in all labels actually corresponding to the characters in the training sample set.
[0184] The third parameter is a parameter that reflects the frequency of occurrence of the label to be screened. The third parameter can be the label frequency. Let's continue to use label c as an example to explain: Get the number of times label c appears in the training sample set. The number of times it appears is the label frequency N of label c. c .
[0185] C4. For each of the tags to be screened, if the first parameter, the second parameter and the third parameter of the tag to be screened all meet the preset conditions, the tag to be screened is determined to be a tag of the target type.
[0186] For each of the multiple tags to be filtered, the first parameter, the second parameter and the third parameter of the tag to be filtered can be obtained respectively, and when the first parameter, the second parameter and the third parameter all meet the preset conditions, the corresponding tag to be filtered is determined as a tag of the target type.
[0187] Let's continue to use label c as an example. The first parameter corresponding to label c is F c , the second parameter corresponding to label c is S c , the third parameter corresponding to label c is N c .
[0188] If the above three parameters meet the following conditions:
[0189] F c <90%
[0190] N c <250
[0191] S c <70%
[0192] Then it can be determined that label c is a label of the target type.
[0193] In the label recognition method provided in the embodiment of the present application, an input text can be obtained, and the input text can be split into a first number of characters, and then each character in the first number of characters is encoded to obtain a corresponding encoding result, a total of the first number of encoding results. Subsequently, the first number of encoding results are input as a whole into each of the second number of multilayer perceptrons, so that the output vector output by each multilayer perceptron can be obtained, a total of the second number of output vectors, wherein each multilayer perceptron has its own corresponding different labels; the number of dimensions of each output vector is the first number, and the first number of dimensions corresponds to the first number of characters one by one. For each output vector, the score values of the first number of dimensions can be compared with the preset score values, and at least one dimension whose score value exceeds the preset score value can be obtained. The above-mentioned at least one dimension corresponds to at least one character, and the at least one character is a noun in the input text, and the label to which the noun belongs is the label corresponding to the output vector to which the above-mentioned at least one dimension belongs.
[0194] In the present application, multiple multi-time perceptrons connected in parallel and corresponding to multiple labels can generate an output vector, and the score values of the dimensions in the output vector can be compared with the preset score values, so that the characters corresponding to the dimensions whose score values exceed the preset score values can be obtained, and further the labels to which these characters belong. The present application utilizes multiple multi-time perceptrons connected in parallel and corresponding to multiple labels, so that the generation process of the output vector is more targeted, so that the type of label of the noun represented by the character can be predicted more accurately.
[0195] In order to better implement the above method, the embodiment of the present application also provides a tag identification device, which can be integrated in an electronic device, and the electronic device can be a terminal, a server, etc. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop, or a personal computer (PC) and other devices; the server can be a single server or a server cluster composed of multiple servers. For example, Figure 8 As shown, the tag identification device may include:
[0196] An input text acquisition unit 801 is used to acquire an input text, where the input text includes a first number of characters;
[0197] The encoding processing unit 802 is used to perform encoding processing on each of the characters to obtain a corresponding encoding result;
[0198] The output vector acquisition unit 803 is used to input the first number of encoding results into each of the second number of multilayer perceptrons to obtain an output vector output by each of the multilayer perceptrons, wherein the second number of multilayer perceptrons are connected in parallel, each of the multilayer perceptrons has a corresponding label, and the dimension of each of the output vectors is the first number, and the first number of dimensions corresponds to the first number of characters one by one;
[0199] The label determination unit 804 is used to obtain, for each output vector of the second number of output vectors, at least one dimension whose score value exceeds a preset score value from the first number of dimensions, wherein the at least one dimension has at least one corresponding character, and the label of the at least one character is the label corresponding to the output vector to which the at least one dimension belongs.
[0200] In some embodiments, the encoding processing unit 802 includes:
[0201] A model acquisition subunit, used to acquire a label recognition model, wherein the label recognition model includes a CNN and a second number of multi-layer perceptrons connected in parallel;
[0202] The result acquisition subunit is used to encode each of the characters through the CNN to obtain a corresponding encoding result.
[0203] In some embodiments, the apparatus further comprises:
[0204] A training text acquisition unit, used to acquire a training text, wherein the training text includes a third number of training characters;
[0205] A training code acquisition unit, used to perform encoding processing on each of the training characters to obtain a corresponding training code result;
[0206] A training output acquisition unit, used for inputting the third number of training encoding results into each of the second number of multilayer perceptrons to obtain a training output vector output by each of the multilayer perceptrons, each of which has a corresponding label;
[0207] A label declaration information unit, used to obtain label declaration information, wherein the label declaration information is used to declare a label of a target type with label information in the training text;
[0208] A perceptron acquisition unit, used to acquire a target multilayer perceptron corresponding to the label of the target type;
[0209] A loss summing calculation unit, used to calculate the loss according to the output vector output by the target multilayer perceptron, and obtain the sum of the losses;
[0210] The summing and convergence unit is used to obtain a trained label recognition model when the sum of the losses converges.
[0211] In some embodiments, the apparatus further comprises:
[0212] The summing non-convergence unit is used to update the parameters of the CNN and the target multi-layer perceptron according to the sum of the losses when the sum of the losses does not converge, until the sum of the losses converges, thereby obtaining a trained label recognition model.
[0213] In some embodiments, the apparatus further comprises:
[0214] The training text screening unit is used to screen selected text from multiple texts to be screened, wherein the selected text includes nouns with labels belonging to the target type, and the selected text is a training text that is not labeled.
[0215] In some embodiments, the training text screening unit includes:
[0216] A first encoding subunit is used to perform a first encoding process on a fourth number of characters included in the text to be screened to obtain a first encoding result;
[0217] a second encoding subunit, configured to perform a second encoding process on a fifth number of characters included in the target type dictionary to obtain a second encoding result, wherein the target type dictionary includes a plurality of target nouns, the plurality of target nouns all belong to the label of the target type, and the plurality of target nouns include the fifth number of characters in total;
[0218] An interaction construction subunit, used for constructing the interaction between the first encoding result and the second encoding result according to the attention mechanism to obtain an interaction result;
[0219] A two-dimensional vector subunit, used for performing a full connection transformation on the interaction result to obtain a two-dimensional vector result;
[0220] A score value acquisition subunit is used to acquire a score value of a first-dimensional vector in the two-dimensional vector result, wherein the score value of the first-dimensional vector reflects the probability that the text to be filtered includes a noun with a label belonging to the target type;
[0221] The text selection subunit is used to select a preset number of texts to be screened with the highest scores of the first dimensional vector from the multiple texts to be screened, and the preset number of texts to be screened are the selected texts.
[0222] In some embodiments, the apparatus further comprises:
[0223] A model acquisition unit, used to acquire a probability prediction model, wherein the probability prediction model includes a first RoBERTa model, a second RoBERTa model, an attention layer, and a fully connected layer;
[0224] A first encoding subunit, specifically configured to perform a first encoding process on the fourth number of characters using the first RoBERTa model;
[0225] A second encoding subunit, specifically configured to perform a second encoding process on the fifth number of characters using the second RoBERTa model;
[0226] An interaction construction subunit, specifically used to construct the interaction between the first encoding result and the second encoding result by using the attention layer;
[0227] The two-dimensional vector quantum unit is specifically used to use the fully connected layer to perform a fully connected transformation on the interaction result.
[0228] In some embodiments, the apparatus further comprises:
[0229] A result judgment unit, used to judge whether the result of the score value of the first dimensional vector is accurate by using the trained binary classification model;
[0230] A parameter updating unit is used to update the parameters of the first RoBERTa model, the second RoBERTa model, the attention layer and the fully connected layer when the result of the score value of the first dimensional vector is inaccurate.
[0231] In some embodiments, the apparatus further comprises:
[0232] The label screening unit is used to screen labels of the target type from a plurality of labels to be screened, wherein the plurality of labels to be screened correspond to the same target technical field.
[0233] In some embodiments, the tag screening unit includes:
[0234] A test set recognition subunit, used to input a test sample set belonging to the target technical field into an original label recognition model, and obtain label recognition results corresponding to characters in the test sample set;
[0235] A test set parameter determination subunit, configured to determine a first parameter and a second parameter of each of the multiple labels to be filtered according to the label recognition results corresponding to the characters in the test sample set and the labels actually corresponding to the characters in the test sample set, wherein the multiple labels to be filtered are all labels actually corresponding to the characters in the test sample set;
[0236] A training set parameter determination subunit, used to determine the third parameter of each to-be-screened label according to the number of occurrences of each label in all labels actually corresponding to the characters in the training sample set;
[0237] The tag screening subunit is used for determining, for each of the tags to be screened, that the tag to be screened is a tag of the target type if the first parameter, the second parameter and the third parameter of the tag to be screened all meet preset conditions.
[0238] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can refer to the previous method embodiments, which will not be repeated here.
[0239] In the label recognition method provided by the embodiment of the present application, multiple multi-time perceptrons connected in parallel and corresponding to multiple labels can generate an output vector, and the score value of the dimension in the output vector can be compared with the preset score value, so that the characters corresponding to the dimensions whose score value exceeds the preset score value can be obtained, and the labels to which these characters belong can be further obtained. The present application utilizes multiple multi-time perceptrons connected in parallel and corresponding to multiple labels, so that the generation process of the output vector is more targeted, so that the type of label of the noun represented by the character can be predicted more accurately.
[0240] In the present application, the prediction accuracy of the label of the noun represented by the character can be improved.
[0241] The embodiment of the present application also provides an electronic device, which can be a terminal, a server, etc. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop, a personal computer, etc. The server can be a single server or a server cluster composed of multiple servers, etc.
[0242] In some embodiments, the tag identification device may also be integrated into multiple electronic devices. For example, the tag identification device may be integrated into multiple servers, and the tag identification method of the present application may be implemented by multiple servers.
[0243] In this embodiment, the electronic device of this embodiment is an electronic device as an example for detailed description, for example, Fig. 9 As shown, it shows a schematic diagram of the structure of the electronic device involved in the embodiment of the present application, specifically:
[0244] The electronic device may include one or more processors 901 of processing cores, one or more computer-readable storage media memories 902, a power supply 903, an input module 904, and a communication module 905. Those skilled in the art will appreciate that Fig. 9 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0245] The processor 901 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 902 and calling data stored in the memory 902, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. In some embodiments, the processor 901 may include one or more processing cores; in some embodiments, the processor 901 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 901.
[0246] The memory 902 can be used to store software programs and modules. The processor 901 executes various functional applications and data processing by running the software programs and modules stored in the memory 902. The memory 902 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 902 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 902 may also include a memory controller to provide the processor 901 with access to the memory 902.
[0247] The electronic device also includes a power supply 903 for supplying power to various components. In some embodiments, the power supply 903 can be logically connected to the processor 901 through a power management system, so as to manage charging, discharging, and power consumption through the power management system. The power supply 903 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.
[0248] The electronic device may further include an input module 904, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0249] The electronic device may further include a communication module 905. In some embodiments, the communication module 905 may include a wireless module. The electronic device may perform short-range wireless transmission through the wireless module of the communication module 905, thereby providing the user with wireless broadband Internet access. For example, the communication module 905 may be used to help the user send and receive emails, browse web pages, and access streaming media.
[0250] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 901 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 902 according to the following instructions, and the processor 901 will run the application programs stored in the memory 902, thereby realizing various functions, as follows:
[0251] Obtain input text, the input text includes a first number of characters; encode each of the characters to obtain a corresponding encoding result; input the first number of encoding results into each of the second number of multilayer perceptrons to obtain an output vector output by each of the multilayer perceptrons, wherein the second number of multilayer perceptrons are connected in parallel, each of the multilayer perceptrons has its own corresponding label, and each of the output vectors has a first number of dimensions, and the first number of dimensions corresponds to the first number of characters one-to-one; for each of the second number of output vectors, obtain at least one dimension whose score value exceeds a preset score value from the first number of dimensions, wherein the at least one dimension has a corresponding at least one character, and the label of the at least one character is the label corresponding to the output vector to which the at least one dimension belongs.
[0252] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0253] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0254] To this end, an embodiment of the present application provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions can be loaded by a processor to execute the steps in any tag identification method provided in the embodiment of the present application. For example, the instructions can execute the following steps:
[0255] Obtain input text, the input text includes a first number of characters; encode each of the characters to obtain a corresponding encoding result; input the first number of encoding results into each of the second number of multilayer perceptrons to obtain an output vector output by each of the multilayer perceptrons, wherein the second number of multilayer perceptrons are connected in parallel, each of the multilayer perceptrons has its own corresponding label, and each of the output vectors has a first number of dimensions, and the first number of dimensions corresponds to the first number of characters one-to-one; for each of the second number of output vectors, obtain at least one dimension whose score value exceeds a preset score value from the first number of dimensions, wherein the at least one dimension has a corresponding at least one character, and the label of the at least one character is the label corresponding to the output vector to which the at least one dimension belongs.
[0256] The storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0257] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various optional implementations provided in the above embodiments.
[0258] Since the instructions stored in the storage medium can execute the steps in any tag identification method provided in the embodiments of the present application, the beneficial effects that can be achieved by any tag identification method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0259] The above is a detailed introduction to a tag identification method, device, electronic device and computer-readable storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A tag recognition method, characterized in that: The method comprises: Obtaining input text, the input text comprising a first number of characters; Performing encoding processing on each of the characters to obtain a corresponding encoding result; Inputting the first number of encoding results into each of the second number of multilayer perceptrons to obtain an output vector output by each of the multilayer perceptrons, wherein the second number of multilayer perceptrons are connected in parallel, each of the multilayer perceptrons has a corresponding label, and the dimension of each of the output vectors is the first number, and the first number of dimensions corresponds to the first number of characters one by one; For each output vector of the second number of output vectors, obtaining at least one dimension whose score value exceeds the preset score value from the first number of dimensions, wherein the at least one dimension has at least one corresponding character, and the label of the at least one character is the label corresponding to the output vector to which the at least one dimension belongs; Wherein, before encoding each of the characters to obtain the corresponding encoding result, the method further includes the step of obtaining a training sample to train a multilayer perceptron to obtain a trained label recognition model; before obtaining the training sample, the method further includes: Filtering selected text from a plurality of to-be-filtered texts, wherein the selected text includes nouns with labels belonging to a target type, and the selected text is a training text that is not labeled; The step of filtering the selected text from the plurality of texts to be filtered includes: For each of the plurality of texts to be screened, the following steps are performed based on the probability prediction model: Performing a first encoding process on a fourth number of characters included in the text to be filtered to obtain a first encoding result; Performing a second encoding process on a fifth number of characters included in the target type dictionary to obtain a second encoding result, wherein the target type dictionary includes a plurality of target nouns, the plurality of target nouns all belong to the label of the target type, and the plurality of target nouns include the fifth number of characters in total; According to the attention mechanism, construct an interaction between the first encoding result and the second encoding result to obtain an interaction result; Performing a full connection transformation on the interaction result to obtain a two-dimensional vector result; Obtaining a score value of a first-dimensional vector in the two-dimensional vector result, wherein the score value of the first-dimensional vector reflects the probability that the text to be filtered includes a noun with a label belonging to the target type; A preset number of texts to be screened having the highest scores of the first dimensional vectors are selected from the multiple texts to be screened, and the preset number of texts to be screened are the selected texts.
2. The tag recognition method according to claim 1, characterized in that: The encoding process is performed on each of the characters to obtain a corresponding encoding result, including: Acquire a label recognition model, wherein the label recognition model includes a CNN and a second number of multi-layer perceptrons connected in parallel; Each of the characters is encoded by the CNN to obtain a corresponding encoding result.
3. The tag recognition method according to claim 2, characterized in that: The step of obtaining training samples to train a multilayer perceptron to obtain a trained label recognition model includes: Acquire a training text, wherein the training text includes a third number of training characters; Performing encoding processing on each of the training characters to obtain a corresponding training encoding result; Input the third number of training encoding results into each of the second number of multilayer perceptrons to obtain a training output vector output by each of the multilayer perceptrons, each of which has a corresponding label; Acquire annotation declaration information, where the annotation declaration information is used to declare a label of a target type with annotation information in the training text; Obtain a target multilayer perceptron corresponding to the label of the target type; Calculating a loss according to the output vector output by the target multilayer perceptron, and obtaining a sum of the losses; If the sum of the losses converges, a trained label recognition model is obtained.
4. The tag recognition method according to claim 3, characterized in that: After calculating the loss according to the output vector output by the target multilayer perceptron and obtaining the sum of the losses, the method further includes: If the sum of the losses does not converge, the parameters of the CNN and the target multi-layer perceptron are updated according to the sum of the losses until the sum of the losses converges, thereby obtaining a trained label recognition model.
5. The tag recognition method according to claim 4, characterized in that: Before performing the first encoding process on the fourth number of characters included in the text to be filtered, the method further includes: Obtain a probability prediction model, wherein the probability prediction model includes a first RoBERTa model, a second RoBERTa model, an attention layer, and a fully connected layer; The performing first encoding processing on the fourth number of characters included in the text to be filtered includes: Performing a first encoding process on the fourth number of characters using the first RoBERTa model; The performing a second encoding process on the fifth number of characters included in the target type dictionary includes: Performing a second encoding process on the fifth number of characters using the second RoBERTa model; The step of constructing the interaction between the first encoding result and the second encoding result according to the attention mechanism includes: Using the attention layer to construct an interaction between the first encoding result and the second encoding result; The performing a full connection transformation on the interaction result includes: The fully connected layer is used to perform a fully connected transformation on the interaction result.
6. The tag recognition method according to claim 5, characterized in that: After obtaining the score value of the first-dimensional vector in the two-dimensional vector result, the method further includes: Using the trained binary classification model to determine whether the score value of the first dimensional vector is accurate; If inaccurate, update the parameters of the first RoBERTa model, the second RoBERTa model, the attention layer, and the fully connected layer.
7. The tag recognition method according to claim 6, characterized in that: Before filtering the selected text from the plurality of texts to be filtered, the method further includes: The tags of the target type are filtered from a plurality of tags to be filtered, and the plurality of tags to be filtered correspond to the same target technical field.
8. The tag recognition method according to claim 7, characterized in that: The step of filtering the target type of tags from a plurality of tags to be filtered includes: Inputting a test sample set belonging to the target technical field into an original label recognition model to obtain label recognition results corresponding to characters in the test sample set; Determine, according to the label recognition results corresponding to the characters in the test sample set and the labels actually corresponding to the characters in the test sample set, the first parameter and the second parameter of each of the multiple labels to be filtered, wherein the multiple labels to be filtered are all labels actually corresponding to the characters in the test sample set; Determine the third parameter of each to-be-screened label according to the number of occurrences of each label in all labels actually corresponding to the characters in the training sample set; For each of the to-be-screened tags, if the first parameter, the second parameter, and the third parameter of the to-be-screened tag all meet the preset conditions, then the to-be-screened tag is determined to be a tag of the target type; Among them, the first parameter and the second parameter are parameters reflecting whether the tags to be filtered can be accurately identified; the third parameter is a parameter reflecting the frequency of occurrence of the tags to be filtered.
9. A label recognition device, characterized in that: The device comprises: An input text acquisition unit, used to acquire input text, wherein the input text includes a first number of characters; A coding processing unit, used for performing coding processing on each of the characters to obtain a corresponding coding result; an output vector acquisition unit, configured to input the first number of encoding results into each of the second number of multilayer perceptrons to obtain an output vector output by each of the multilayer perceptrons, wherein the second number of multilayer perceptrons are connected in parallel, each of the multilayer perceptrons has a corresponding label, the dimension of each of the output vectors is the first number, the first number of dimensions corresponds to the first number of characters one by one, and each of the multilayer perceptrons is trained for its corresponding label during the training process; a label determination unit, configured to obtain, for each output vector of the second number of output vectors, at least one dimension whose score value exceeds a preset score value from the first number of dimensions, wherein the at least one dimension has at least one corresponding character, and the label of the at least one character is the label corresponding to the output vector to which the at least one dimension belongs; Wherein, before encoding each of the characters to obtain the corresponding encoding result, the method further includes the step of obtaining a training sample to train a multilayer perceptron to obtain a trained label recognition model; before obtaining the training sample, the method further includes: Filtering selected text from a plurality of to-be-filtered texts, wherein the selected text includes nouns with labels belonging to a target type, and the selected text is a training text that is not labeled; The step of filtering the selected text from the plurality of texts to be filtered includes: For each of the plurality of texts to be screened, the following steps are performed based on the probability prediction model: Performing a first encoding process on a fourth number of characters included in the text to be filtered to obtain a first encoding result; Performing a second encoding process on a fifth number of characters included in the target type dictionary to obtain a second encoding result, wherein the target type dictionary includes a plurality of target nouns, the plurality of target nouns all belong to the label of the target type, and the plurality of target nouns include the fifth number of characters in total; According to the attention mechanism, construct an interaction between the first encoding result and the second encoding result to obtain an interaction result; Performing a full connection transformation on the interaction result to obtain a two-dimensional vector result; Obtaining a score value of a first-dimensional vector in the two-dimensional vector result, wherein the score value of the first-dimensional vector reflects the probability that the text to be filtered includes a noun with a label belonging to the target type; A preset number of texts to be screened having the highest scores of the first dimensional vectors are selected from the multiple texts to be screened, and the preset number of texts to be screened are the selected texts.
10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps in the tag identification method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the tag identification method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Sequence labeling method and device and training method of sequence labeling model
CN110210035A
Multi-task language analysis system and method based on shared representation
CN110309511A