Label extraction method, device and equipment

By combining multiple extraction algorithms (semantic matching, title matching and additional information matching) and performing linear summing operations, the problem of insufficient label extraction accuracy and recall in the prior art is solved, and higher label extraction accuracy and recall rate are achieved.

CN114239596BActive Publication Date: 2025-05-30GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111537207.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-15
Publication Date
2025-05-30
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

In the prior art, the accuracy and recall rate of the tag extraction method are insufficient to meet business needs, especially in the case of diversity in text content.

Method used

A label extraction method is adopted to obtain the correlation between the content obtained by multiple extraction algorithms (based on semantic matching, title matching and additional information matching) and perform linear summing of the correlation value through a preset fusion algorithm. When the correlation value is greater than the set threshold, the corresponding label is output.

Benefits of technology

Through the complementary and fusion processing of different extraction algorithms, the accuracy and recall of tag extraction are improved, and accurate tags can be extracted more effectively in the diverse text content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114239596B_ABST
    Figure CN114239596B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, and device for label extraction. The label extraction method includes: respectively obtaining the relevance between content and labels obtained by two or more preset extraction algorithms; performing a linear summation operation on the respectively obtained relevance between content and labels through a preset fusion algorithm to obtain a relevance operation value; and when the relevance operation value is greater than a relevance threshold, outputting the label corresponding to the relevance operation value. The solution provided by this application can improve the accuracy and recall rate of label extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a method, device, and equipment for label extraction. Background Art

[0002] Currently, in order to facilitate users to quickly understand the core points of text content, labels are usually set for the content. Labels are a series of words for classifying and applying content.

[0003] In related technologies, there are mainly two methods for content label extraction: The first is the supervised learning method, which requires a large number of sample texts to be labeled, and the labor cost is quite high. Adding new content labels requires retraining the model, and the retrained model will affect the effect of the original content labels, and most content labels need to be re-evaluated for their effects; The second is the unsupervised learning method. Although this method greatly reduces the manual labeling cost, the accuracy and recall rate of the extracted content labels are not good and cannot meet the business requirements.

[0004] However, most of the methods adopted in related technologies only consider the effects in certain specific scenarios, and in the actual application process, due to the diversity of text content, the accuracy and recall rate of the labels extracted by using the related technology methods need to be improved. Summary of the Invention

[0005] To solve or partially solve the problems existing in related technologies, this application provides a method, device, and equipment for label extraction, which can improve the accuracy and recall rate of label extraction.

[0006] The first aspect of this application provides a method for label extraction, including:

[0007] Respectively obtain the relevance between the content and the label obtained by two or more set extraction algorithms;

[0008] Perform a linear summation operation on the respectively obtained relevance between the content and the label through a preset fusion algorithm to obtain a relevance operation value;

[0009] When the relevance operation value is greater than the relevance threshold, output the label corresponding to the relevance operation value.

[0010] In an embodiment, the step of respectively obtaining the relevance between the content and the label obtained by two or more set extraction algorithms includes:

[0011] Obtain a first relevance between the content and the label obtained by a first extraction algorithm, where the first extraction algorithm is based on semantic matching;

[0012] Obtain the second relevance between the content obtained by the second extraction algorithm and the label, where the second extraction algorithm is based on title matching;

[0013] Obtain the third relevance between the content obtained by the third extraction algorithm and the label, where the third extraction algorithm is based on additional information matching.

[0014] In one embodiment, the linear summation operation of the relevances between the respectively obtained content and the label through a preset fusion algorithm to obtain a relevance operation value includes:

[0015] Obtain a relevance matrix of the relevance between the content and the label according to the number of the labels and the number of the set extraction algorithms;

[0016] Obtain a preset weight matrix, where the weight matrix includes the weights of the set extraction algorithms for the labels;

[0017] Perform a linear summation operation on the relevance matrix and the weight matrix to obtain a relevance operation value.

[0018] In one embodiment, the obtaining the first relevance between the content obtained by the first extraction algorithm and the label, where the first extraction algorithm is based on semantic matching, includes:

[0019] For the text content, after obtaining the keywords of the text content according to the keyword extraction algorithm, input the keywords into the word2vec model to obtain the vectors of the keywords;

[0020] For the text content corresponding to the label, after obtaining the keywords of the text content according to the keyword extraction algorithm, then determine the core words and the related words of the core words, input the core words and the related words into the word2vec model to obtain the word vectors of the core words and the related words, and perform vector summation on the word vectors of the core words and the related words to obtain the vector of the label;

[0021] Perform a first set operation on the vector of the keyword and the vector of the label to obtain the first relevance between the keyword and the label.

[0022] In one embodiment, the related words are divided into highly related words and lowly related words according to the similarity.

[0023] In one embodiment, the obtaining the second relevance between the content obtained by the second extraction algorithm and the label, where the second extraction algorithm is based on title matching, includes:

[0024] For the title of the text content, after obtaining the keywords of the title according to the keyword extraction algorithm, perform a second set operation on the keywords of the title to obtain the second relevance between the title and the label.

[0025] In one embodiment, obtaining a third relevance degree between the content obtained by a third extraction algorithm and a label, where the third extraction algorithm is based on additional information matching, includes:

[0026] For text content, after obtaining the additional information of the text content according to the keyword extraction algorithm, perform a third setting operation on the additional information to obtain a third relevance degree between the additional information and the label.

[0027] The second aspect of the present application provides a label extraction device, including:

[0028] An acquisition module, configured to respectively obtain the relevance degrees between the content obtained by two or more setting extraction algorithms and a label;

[0029] An operation module, configured to perform a linear summation operation on the relevance degrees between the content and the label respectively obtained by the acquisition module through a preset fusion algorithm to obtain a relevance degree operation value;

[0030] An output module, configured to output the label corresponding to the relevance degree operation value when the relevance degree operation value obtained by the operation module is greater than a relevance degree threshold.

[0031] In one embodiment, the acquisition module includes:

[0032] A first acquisition sub-module, configured to obtain a first relevance degree between the content obtained by a first extraction algorithm and a label, where the first extraction algorithm is based on semantic matching;

[0033] A second acquisition sub-module, configured to obtain a second relevance degree between the content obtained by a second extraction algorithm and a label, where the second extraction algorithm is based on title matching;

[0034] A third acquisition sub-module, configured to obtain a third relevance degree between the content obtained by a third extraction algorithm and a label, where the third extraction algorithm is based on additional information matching.

[0035] In one embodiment, the operation module includes:

[0036] A relevance degree sub-module, configured to obtain a relevance degree matrix of the relevance degrees between the content and the label according to the number of labels and the number of setting extraction algorithms;

[0037] A weight sub-module, configured to obtain a preset weight matrix, where the weight matrix includes the weights of the setting extraction algorithms for the labels;

[0038] A calculation sub-module, configured to perform a linear summation operation on the relevance degree matrix and the weight matrix to obtain a relevance degree operation value.

[0039] The third aspect of the present application provides an electronic device, including:

[0040] a processor; and

[0041] a memory storing executable code, which, when executed by the processor, causes the processor to execute the method as described above.

[0042] The fourth aspect of the present application provides a computer-readable storage medium storing executable code, which, when executed by a processor of an electronic device, causes the processor to execute the method as described above.

[0043] The technical solution provided by the present application may include the following beneficial effects:

[0044] In the label extraction method of the present application, two or more preset extraction algorithms are integrated to extract labels. The relevance between the content obtained by different algorithms and the labels is linearly summed through a preset fusion algorithm to obtain a relevance operation value. Finally, when the relevance operation value is greater than the relevance threshold, the label corresponding to the relevance operation value is output. By making up for each other's advantages of different preset extraction algorithms and performing fusion processing on different preset extraction algorithms before finally determining the output label, the accuracy and recall rate of label extraction can be improved.

[0045] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] By describing the exemplary embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. Among them, in the exemplary embodiments of the present application, the same reference numerals generally represent the same components.

[0047] Figure 1 is a schematic flowchart of the label extraction method shown in the embodiments of the present application;

[0048] Figure 2 is another schematic flowchart of the label extraction method shown in the embodiments of the present application;

[0049] Figure 3 is a schematic diagram of the application framework structure of the label extraction method shown in the embodiments of the present application;

[0050] Figure 4 is a schematic diagram of the structure of the label extraction device shown in the embodiments of the present application;

[0051] Figure 5 is another schematic diagram of the structure of the label extraction device shown in the embodiments of the present application;

[0052] Figure 6 It is a schematic structural diagram of an electronic device shown in an embodiment of the present application. Detailed implementation manners

[0053] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application will be more thorough and complete, and can fully convey the scope of the present application to those skilled in the art.

[0054] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0055] It should be understood that although the terms "first", "second", "third", etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.

[0056] Currently in the related art, for most label extraction methods, the accuracy and recall rate of the extracted labels need to be improved. In view of the above problems, the present application provides a label extraction method, which can improve the accuracy and recall rate of label extraction.

[0057] To facilitate the understanding of the solution of the embodiment of the present application, the technical solution of the embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0058] Please refer to Figure 1 , a label extraction method, including:

[0059] S11. Obtain the relevance between the content and the label obtained by two or more set extraction algorithms respectively.

[0060] This step may include: obtaining a first relevance between the content and the label obtained by a first extraction algorithm, where the first extraction algorithm is based on semantic matching; obtaining a second relevance between the content and the label obtained by a second extraction algorithm, where the second extraction algorithm is based on title matching; obtaining a third relevance between the content and the label obtained by a third extraction algorithm, where the third extraction algorithm is based on additional information matching.

[0061] S12. Linearly sum the relevances between the respectively obtained content and the label through a preset fusion algorithm to obtain a relevance operation value.

[0062] This step may include: obtaining a relevance matrix of the relevance between the content and the label according to the number of labels and the number of set extraction algorithms; obtaining a preset weight matrix, where the weight matrix includes the weights of the set extraction algorithms for the labels; performing a linear summation operation on the relevance matrix and the weight matrix to obtain a relevance operation value.

[0063] Among them, obtaining the first relevance between the content and the label obtained by the first extraction algorithm, where the first extraction algorithm is based on semantic matching, includes: for the text content, after obtaining the keywords of the text content according to the keyword extraction algorithm, inputting the keywords into a word2vec model (a related model for generating word vectors) to obtain the vectors of the keywords; for the text content corresponding to the label, after obtaining the keywords of the text content according to the keyword extraction algorithm, determining the core words and related words among them, inputting the core words and related words into the word2vec model to obtain the vectors of the core words and related words, and summing the vectors of the core words and related words to obtain the vector of the label; performing a first set operation on the vector of the keywords and the vector of the label to obtain the first relevance between the keywords and the label.

[0064] Among them, obtaining the second relevance between the content and the label obtained by the second extraction algorithm, where the second extraction algorithm is based on title matching, includes: for the title of the text content, after obtaining the keywords of the title according to the keyword extraction algorithm, performing a second set operation on the keywords of the title to obtain the second relevance between the title and the label.

[0065] Among them, obtaining the third relevance between the content and the label obtained by the third extraction algorithm, where the third extraction algorithm is based on additional information matching, includes: for the text content, after obtaining the additional information of the text content according to the keyword extraction algorithm, performing a third set operation on the additional information to obtain the third relevance between the additional information and the label.

[0066] S13. When the relevance operation value is greater than the relevance threshold, output the label corresponding to the relevance operation value.

[0067] The relevance threshold can be determined empirically or calculated using a set algorithm. When the calculated relevance operation value is greater than the relevance threshold, the label corresponding to the relevance operation value is output, and this label is used as the finally extracted label.

[0068] As can be seen from this embodiment, the label extraction method of the present application combines two or more set extraction algorithms to extract labels. The relevance between the content obtained by different algorithms and the labels is linearly summed through a preset fusion algorithm to obtain a relevance operation value. Finally, when the relevance operation value is greater than the relevance threshold, the label corresponding to the relevance operation value is output. By leveraging the complementary advantages of different set extraction algorithms and fusing different set extraction algorithms before finally determining the output label, the accuracy and recall rate of label extraction can be improved.

[0069] Figure 2 It is another schematic flowchart of the label extraction method shown in the embodiments of the present application.

[0070] The technical solution of the present application can be applied to text content such as articles, posts, pictures and texts, but is not limited thereto. The technical solution of the present application has good application effects for text content such as long texts and short texts. In the technical solution of the present application, multiple label extraction algorithms are used, and an algorithm fusion mechanism based on statistical evaluation is adopted. For example, the present application can use 3 label extraction algorithms, including a semantic matching-based extraction algorithm, a title matching-based extraction algorithm, and an additional information matching-based extraction algorithm, and fuse them. Among them, the semantic matching-based extraction algorithm is the core, and secondly, the title matching-based extraction algorithm and the additional information matching-based extraction algorithm play an auxiliary role. For example, for some professional type labels or activity type labels, applying the title matching-based extraction algorithm and the additional information matching-based extraction algorithm can have an auxiliary effect. For professional type labels such as "wing doors" (the semantics of such labels are relatively professional and concise) and activity type labels such as "Task X", relevant features are often contained in the text content and the content's attached information. It should be noted that this embodiment takes the implementation of extracting labels by fusing 3 label extraction algorithms as an example, but is not limited thereto. According to needs, 2 label extraction algorithms can also be set, and then 2 label extraction algorithms can be fused to implement label extraction, and the principle is similar.

[0071] Please refer to Figure 2 and Figure 3 , a label extraction method, including:

[0072] S21. Obtain the first relevance between the content obtained by the first extraction algorithm and the label, where the first extraction algorithm is based on semantic matching.

[0073] Labels generally refer to a series of words used to classify and apply content. For example, labels such as "autopilot" and "endurance" indicate that the content is related to autopilot or endurance. Labels can be generated according to operational requirements. Generally, there is no correlation or mutual exclusivity between labels, and they will be continuously added and deleted according to operational requirements. It is necessary to extract content periodically and automatically. Under this requirement, the algorithm needs to have relatively strong flexibility.

[0074] Among them, the first extraction algorithm can be a semantic matching-based extraction algorithm, and the semantic matching-based extraction algorithm can include algorithms such as keyword extraction algorithm, text semantic representation generation algorithm, label semantic representation generation algorithm, and content label matching algorithm. The following will be described in detail in combination with Figure 3 respectively.

[0075] In Figure 3 In the schematic diagram shown, different processes are performed on the text content, labels, text title, and text content additional information respectively. For example, for the text content, after preprocessing (such as text cleaning, word segmentation, stop word removal, synonym replacement, etc.), combined with a preset keyword library composed of domain keywords, etc., keywords and word frequencies are obtained according to the keyword extraction algorithm, and the semantic representation vector is obtained by using the text semantic representation generation algorithm and the word2vec model; for the labels, the label-related words and label core words can be obtained by using the label-assisted calibration tool, and the label semantic representation vector is obtained by using the label semantic representation generation algorithm; the results obtained by using the text semantic representation generation algorithm and the label semantic representation generation algorithm are then processed by using the semantic matching-based extraction algorithm to obtain the first label correlation degree; for the text title, the keywords of the title are obtained by using the keyword extraction algorithm and then processed by using the title matching-based extraction algorithm to obtain the second label correlation degree; for the text content additional information, it is processed by using the additional information matching-based extraction algorithm to obtain the third label correlation degree; finally, by using the algorithm fusion mechanism based on statistical evaluation, the first label correlation degree, the second label correlation degree, and the third label correlation degree are fused and calculated to obtain the correlation degree calculation value. When the correlation degree calculation value is greater than the correlation degree threshold, the label corresponding to the correlation degree calculation value is output.

[0076] 1) Using the text semantic representation generation algorithm, obtain the vector of keywords.

[0077] For this step, for the text content, after obtaining the keywords of the text content according to the keyword extraction algorithm, the keywords are input into the word2vec model to obtain the vector of keywords.

[0078] Among them, the processing process of the keyword extraction algorithm in this application includes: obtaining the input text content; obtaining the first score of each candidate word in the input text content according to the first preset algorithm; obtaining the second score of each candidate word in the input text content according to the second preset algorithm; performing a preset operation on the first score of each candidate word and the second score of each candidate word to obtain the third score of each candidate word; determining the keywords of the input text content according to the sorting result of the third scores of each candidate word. Among them, the first score of each candidate word in the input sample can be obtained according to the preset keyword library and the first preset algorithm. The first preset algorithm can be the SeedTextRank algorithm, and the SeedTextRank algorithm includes weighting each candidate word, and the weighting process includes keyword weighting, position weighting, edge weight weighting, etc. Among them, the second score of each candidate word in the input sample can be obtained according to the preset corpus corresponding to the preset keyword library and the second preset algorithm. The second preset algorithm can be the TFIDF (Term Frequency-Inverse Document Frequency) algorithm. The preset corpus can be, for example, the corpus of a certain company or industry. The preset keyword library includes core words and corresponding related words. The core words and corresponding related words in the preset keyword library can be determined in the following way in advance: screening the words with the top scores from the preset corpus; screening the words related to the preset field from the words with the top scores according to the preset field as core words; performing a synonym search on the core words, and performing part-of-speech filtering and length filtering, and screening a set number of words related to the core words as the corresponding related words. Among them, the first score of each candidate word and the second score of each candidate word can be subjected to a weighted summation operation to obtain the third score of each candidate word; the keywords of the input sample are determined according to the sorting result of the third scores of each candidate word and the preset invalid word library.

[0079] Among them, word2vec is an open-source word vector algorithm model. The solution of this application can use the word2vec tool to train a pre-trained model with 5 million articles of Wikipedia as the corpus but not limited to this; then, input a certain community article, train and iterate a set number of times, such as 100 rounds, and finally obtain a word vector model, where the dimension of each vector can be 100. The reason for using the word2vec algorithm to make a word vector model in this application is that the keyword extraction algorithm, text semantic representation generation algorithm, etc. in this application are all a series of bag-of-words models (word vector models), so choosing to use word2vec is a better match, while other semantic deep models such as bert do not have a better effect than word2vec in the solution of this application. The functions of the word2vec model in the solution of this application include: inputting a word can obtain the word vector of this word; inputting a word can output N related words and their word vectors.

[0080] After the keywords and their frequencies of the text are extracted by the keyword extraction algorithm in this application, the words ranked among the top can be selected according to the weights of the keywords, for example, top 10 are selected, and then input into the word2vec word vector model, and a matrix of N×P can be obtained, where N is the number of keywords (N <= 10), and P is the dimension of the word vector (P = 100). This N×P matrix is the semantic representation of the current text.

[0081] 2) Use the label semantic representation generation algorithm to obtain the vector of the label.

[0082] For the text content corresponding to the label, after obtaining the keywords of the text content according to the keyword extraction algorithm, then determine the core words and related words among them, input the core words and related words into the word2vec model, obtain the word vectors of the core words and related words, and perform vector summation on the word vectors of the core words and related words to obtain the vector of the label.

[0083] This step can perform a small amount of annotation on the label, and the processing process is exemplified as follows:

[0084] The first step: First, annotate several core words under a certain label.

[0085] These core words are a series of words most relevant to this label. For example, the core words of the autonomous driving label are NGP (Navigation Guided Pilot), LCC (Lane Centering Control), ACC (Adaptive Cruise Control), autonomous driving, Xpilot (intelligent assistance system), etc. The annotation method is to rank several keywords mentioned by using the keyword extraction algorithm on several pieces of text content under this label according to the word frequency, and then screen out the words with relatively high relevance to this label as core words. It should be noted that about 5 - 10 core words can be annotated.

[0086] The second step: Input these core words into a preset label auxiliary calibration tool, and output a series of related keywords of these core words as related words.

[0087] For example, the above core words can be input into the word2vec model to obtain similar words and their similarities; then, by using the TFIDF algorithm, words with smaller IDF (Inverse Document Frequency) values or smaller TF (Term Frequency) values of the above similar words in the set corpus are calculated, and the words ranked in the top positions in terms of similarity are output, such as the top 100, and then several words related to the label are selected as relevant words. If the similarity is greater than a set threshold, such as 0.3, it is highly relevant; otherwise, it is lowly relevant, so that the relevant words are divided into two levels: highly relevant words and lowly relevant words. By using the TFIDF algorithm, some words with little significance output according to the word2vec model can be deleted.

[0088] Step 3: Input the obtained several core words, highly relevant words, and lowly relevant words into the word2vec model to obtain the word vectors of each word.

[0089] Step 4: For each label, sum the word vectors obtained from its core words, highly relevant words, and lowly relevant words to obtain an N×P vector matrix of the label.

[0090] Where N is the total number of labels, and P = 100 is the vector dimension.

[0091] 3) Use the content label matching algorithm to obtain the first relevance between the keyword and the label.

[0092] According to each keyword c and its corresponding semantic vector Vc output by the text semantic representation generation algorithm for the current text, the relevance w between the keyword c and the label t can be obtained using the following formula:

[0093] w = dist(Vc, Vt) × min(frequence(c), 3) × key_weight(c, t)

[0094] Where:

[0095] c is the text keyword, Vc is the word2vec vector of the text keyword, and Vt is the word2vec vector of the label t. It should be noted that a series of labels can be defined in the business, and this operation is performed on each label in the series of labels. Here, the label t refers to the t-th label in the series of labels;

[0096] dist(Vc, Vt) is the cosine similarity between the vector Vc and the vector Vt;

[0097] The frequence(c) is the number of times the text keyword c appears in the current text content, and min(frequence(c), 3) represents taking the smaller value between frequence(c) and 3;

[0098] The key_weight(c, t) is the weight of the text keyword c. If the text keyword c is the core word in the tag t, the weight can return 4.0. If it is a highly relevant word, the weight can return 2.0. If it is a low relevant word, the weight can return 0.56. It should be noted that the weight values here are only for illustrative purposes and are not limited to this.

[0099] At this time, a total of N × M relevance degrees are obtained, where N is the number of tags and M is the number of keywords in the current text. Summing up the above M relevance degrees can obtain the relevance degree between the current text content and a tag; performing the above operation on each tag can obtain N relevance degrees. Because among the N × M relevance degrees, every M (there are M relevance degrees for the current text keywords) are combined into one, and finally there are N relevance degrees.

[0100] S22. Obtain the second relevance degree between the content obtained by the second extraction algorithm and the tag, where the second extraction algorithm is based on title matching.

[0101] For the title of the text content, after obtaining the keywords of the title according to the keyword extraction algorithm, perform a second setting operation on the keywords of the title to obtain the second relevance degree between the title and the tag.

[0102] Among them, the second extraction algorithm can be a title matching extraction algorithm. This title matching extraction algorithm can be an algorithm that extracts the corresponding tags through the title of the content. The processing process can be as follows:

[0103] The keywords of the title can be extracted according to the keyword extraction algorithm. For each keyword c of the title, the relevance degree w between the keyword c and the tag t can be obtained using the following formula:

[0104] w = key_weight(c, t)

[0105] Among them, the key_weight(c, t) is the weight of the text keyword c. If the text keyword c is the core word in the tag t, the weight can return 1.0, otherwise it returns 0.0. It should be noted that the weight values here are only for illustrative purposes and are not limited to this.

[0106] At this time, a total of M relevance degrees are obtained. Among them, M is the number of keywords. Summing up the M relevance degrees can obtain the relevance degree between the current title and a tag; performing the above operation on each tag can obtain N relevance degrees.

[0107] S23. Obtain the third relevance between the content obtained by the third extraction algorithm and the label, where the third extraction algorithm is based on additional information matching.

[0108] For text content, after obtaining the additional information of the text content according to the keyword extraction algorithm, perform a third setting operation on the additional information to obtain the third relevance between the additional information and the label.

[0109] Among them, the third extraction algorithm can be an extraction algorithm based on additional information matching. This extraction algorithm based on additional information matching is an algorithm that uses the additional information of the content to determine the label of the content through rules, and can include the following rules:

[0110] 1) Highlight information matching

[0111] First, the highlighted content can be extracted from the text content. For example, ##Model ① Promotion##, the extracted highlighted content is "Model ① Promotion".

[0112] For each label, some highlighted content is pre-configured. If the extracted highlighted content is the pre-configured highlighted content under a certain label, output the label relevance w = 1.0, otherwise w = 0.0. It should be noted that the values here are only for illustration but not limited to this.

[0113] Through the above processing, N relevances can be obtained, where N is the number of labels.

[0114] 2) Source matching

[0115] First, the source can be extracted from the content. For example, "Paid Column - Charging".

[0116] For each label, some sources are pre-configured. If the extracted content source is the pre-configured source under a certain label, output the label relevance w = 1.0, otherwise w = 0.0. It should be noted that the values here are only for illustration but not limited to this.

[0117] Through the above processing, N relevances can be obtained, where N is the number of labels.

[0118] 3) Creator matching

[0119] First, the creator can be extracted from the content.

[0120] For each label, some creators are pre-configured. If the extracted creator is the pre-configured source under a certain label, output the label relevance w = 1.0, otherwise w = 0.0. It should be noted that the values here are only for illustration but not limited to this.

[0121] Through the above processing, N relevances can be obtained, where N is the number of labels.

[0122] S24. Linearly sum the relevance degrees of the separately obtained contents and the labels through a preset fusion algorithm to obtain a relevance degree calculation value.

[0123] The above multiple extraction algorithms have all output the relevance degrees of the current text content and each label, that is, N×K relevance degrees are output, where K is the number of algorithms.

[0124] Considering that the accuracy rates of each algorithm under different labels are different, the solution of this application uses an algorithm fusion mechanism based on evaluation, and this algorithm fusion mechanism includes a weight matrix and a threshold vector.

[0125] Among them, the N×K weight matrix is shown as follows, for example:

[0126] w11, w12, w13...w1k

[0127] w21, w22, w23...w2k

[0128] w31, w32, w33...w3k

[0129] wn1, wn2, wn3...wnk

[0130] Among them, the weight matrix wnk represents the weight of the kth algorithm under the label n.

[0131] Among them, the N-dimensional threshold vector is shown as follows, for example:

[0132] S1, S2....SN

[0133] Linearly sum the N×K relevance degrees according to the N×K weight matrix to obtain N relevance degrees representing the current text and N labels.

[0134] Among them, the algorithm weight vector in the weight matrix can be preset in the following way:

[0135] 1) Set the weights of other algorithms to 0;

[0136] 2) Input a certain amount of random evaluation content samples into the current algorithm, and each sample obtains the relevance degree of the current label, and adjust the relevance degree threshold S; each time of adjustment, the accuracy rate and recall rate of the label output can be evaluated, and the maximum accuracy rate with the recall rate greater than the set threshold, such as 70%, can be taken as the weight of the current algorithm under the current label;

[0137] 3) Repeat the above operations for each algorithm to obtain the weights of each algorithm under the current label;

[0138] 4) Repeat the above operations for each tag to obtain the weight of each algorithm for each tag.

[0139] S25. When the relevance operation value is greater than the relevance threshold, output the tag corresponding to the relevance operation value.

[0140] If the relevance of the tag is greater than the relevance threshold S, output this tag.

[0141] Among them, the relevance threshold S can be preset in the following way:

[0142] After obtaining the weights using the above algorithm weight calculation method, set them into the weight matrix, and then input a certain amount of random evaluation content samples into the current algorithm. Adjust the relevance threshold S for each tag to maximize its F1 score (F1 score). Thus, the relevance threshold S for all tags is obtained. The F1 score is a measure for classification problems. In some machine learning competitions for multi-classification problems, the F1 score is often used as the final evaluation method. It is the harmonic mean of precision and recall, with a maximum of 1 and a minimum of 0.

[0143] It can be found from this embodiment that compared with general algorithms, the solution of the present application does not require global retraining and does not require global re-adjustment of parameters when adding or deleting tags. Only the semantic definitions of the added or deleted tags are needed, and the old tags do not require retraining of the model. At the same time, the manual annotation cost is low. In addition, the present application integrates multiple algorithms and uses a fusion mechanism, which also enables the solution of the present application to maintain a high accuracy and recall rate even when the content diversity is very wide.

[0144] Corresponding to the foregoing method embodiment for implementing the application function, the present application also provides a tag extraction device, an electronic device, and corresponding embodiments.

[0145] Figure 4 It is a schematic structural diagram of the tag extraction device shown in the embodiment of the present application.

[0146] Please refer to Figure 4 , a tag extraction device 40, including: an acquisition module 41, an operation module 42, and an output module 43.

[0147] The acquisition module 41 is used to respectively acquire the relevance between the content and the tag obtained by two or more set extraction algorithms. The acquisition module 41 can acquire the first relevance between the content and the tag obtained by the first extraction algorithm, where the first extraction algorithm is based on semantic matching; acquire the second relevance between the content and the tag obtained by the second extraction algorithm, where the second extraction algorithm is based on title matching; acquire the third relevance between the content and the tag obtained by the third extraction algorithm, where the third extraction algorithm is based on additional information matching.

[0148] An operation module 42 is configured to perform a linear summation operation on the relevance between the content respectively obtained by the acquisition module 41 and the label through a preset fusion algorithm to obtain a relevance operation value. The operation module 42 may obtain a relevance matrix of the relevance between the content and the label according to the number of labels and the number of set extraction algorithms; obtain a preset weight matrix, where the weight matrix includes the weights of the set extraction algorithms for the labels; perform a linear summation operation on the relevance matrix and the weight matrix to obtain a relevance operation value.

[0149] An output module 43 is configured to output the label corresponding to the relevance operation value when the relevance operation value obtained by the operation module 42 is greater than a relevance threshold.

[0150] It can be seen from this embodiment that the label extraction device of the present application fuses two or more set extraction algorithms to extract labels, performs a linear summation operation on the relevance between the content respectively obtained by different algorithms and the label through a preset fusion algorithm to obtain a relevance operation value, and finally outputs the label corresponding to the relevance operation value when the relevance operation value is greater than the relevance threshold. By making up for each other's advantages of different set extraction algorithms and performing a fusion process on different set extraction algorithms and then finally determining the output label, the accuracy and recall rate of label extraction can be improved.

[0151] Figure 5 It is another structural schematic diagram of the label extraction device shown in the embodiment of the present application.

[0152] Please refer to Figure 5 , a label extraction device 40, including: an acquisition module 41, an operation module 42, and an output module 43. The functions of the acquisition module 41, the operation module 42, and the output module 43 can be referred to the description in Figure 4 .

[0153] Among them, the acquisition module 41 may include: a first acquisition sub-module 411, a second acquisition sub-module 412, and a third acquisition sub-module 413.

[0154] The first acquisition sub-module 411 is configured to acquire a first relevance degree between content and a label obtained by a first extraction algorithm, where the first extraction algorithm is based on semantic matching. For text content, the first acquisition sub-module 411 can, after obtaining keywords of the text content according to a keyword extraction algorithm, input the keywords into a word2vec model to obtain vectors of the keywords; for the text content corresponding to the label, after obtaining keywords of the text content according to the keyword extraction algorithm, then determine core words and related words thereof, input the core words and related words into the word2vec model to obtain word vectors of the core words and related words, perform vector summation on the word vectors of the core words and related words to obtain a vector of the label; perform a first preset operation according to the vectors of the keywords and the vector of the label to obtain a first relevance degree between the keywords and the label. Among them, related words are divided into highly related words and lowly related words according to similarity.

[0155] The second acquisition sub-module 412 is configured to acquire a second relevance degree between content and a label obtained by a second extraction algorithm, where the second extraction algorithm is based on title matching. For the title of the text content, the second acquisition sub-module 412 can, after obtaining keywords of the title according to a keyword extraction algorithm, perform a second preset operation on the keywords of the title to obtain a second relevance degree between the title and the label.

[0156] The third acquisition sub-module 413 is configured to acquire a third relevance degree between content and a label obtained by a third extraction algorithm, where the third extraction algorithm is based on additional information matching. For text content, the third acquisition sub-module 413 can, after obtaining additional information of the text content according to a keyword extraction algorithm, perform a third preset operation on the additional information to obtain a third relevance degree between the additional information and the label.

[0157] Among them, the operation module 42 may include: a relevance sub-module 421, a weight sub-module 422, and a calculation sub-module 423.

[0158] The relevance sub-module 421 is configured to obtain a relevance matrix of the relevance degree between content and a label according to the number of labels and the number of preset extraction algorithms.

[0159] The weight sub-module 422 is configured to obtain a preset weight matrix, where the weight matrix includes weights of the preset extraction algorithms in the label.

[0160] The calculation sub-module 423 is configured to perform a linear summation operation on the relevance matrix and the weight matrix to obtain a relevance operation value.

[0161] The label extraction device of the present application integrates multiple algorithms and utilizes a fusion mechanism, and also enables the solution of the present application to maintain a high accuracy rate and recall rate even when the diversity of content is very wide.

[0162] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0163] Figure 6 It is a schematic structural diagram of an electronic device shown in an embodiment of the present application.

[0164] See Figure 6 , the electronic device 1000 includes a memory 1010 and a processor 1020.

[0165] The processor 1020 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0166] The memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM can store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage device can be a readable and writable storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory 1010 can include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory 1010 can include a removable storage device that is readable and / or writable, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, super density disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and instantaneous electronic signals transmitted wirelessly or wired.

[0167] Executable code is stored on the memory 1010, and when the executable code is processed by the processor 1020, it can cause the processor 1020 to execute some or all of the methods described above.

[0168] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, and the computer program or the computer program product includes computer program code instructions for executing some or all of the steps of the above method of the present application.

[0169] Alternatively, the present application can also be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium), on which executable code (or a computer program or computer instruction code) is stored. When the executable code (or the computer program or the computer instruction code) is executed by a processor of an electronic device (or a server, etc.), it causes the processor to execute some or all of the steps of the above method according to the present application.

[0170] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for label extraction, characterized in that, it includes: respectively obtaining the relevance between the content and the label obtained by two or more preset extraction algorithms; the relevance is obtained by performing a preset operation on the keywords of the text content through the preset extraction algorithm; obtaining a relevance matrix of the relevance between the content and the label according to the number of the labels and the number of the preset extraction algorithms, the size of the relevance matrix is N×K, N is the number of the labels, K is the number of the preset extraction algorithms, and the relevance matrix includes the relevance between the preset extraction algorithm and the label; obtaining a preset weight matrix, where the weight matrix includes the weights of the preset extraction algorithms for the labels, and wherein, the weight matrix is preset in the following manner: for each preset extraction algorithm, inputting a random evaluation content sample, obtaining the relevance between the random evaluation content sample and the label, adjusting the relevance threshold, and determining the weights of the preset extraction algorithms for the labels according to the recall rate and accuracy rate of the label output evaluated each time; performing a linear summation operation on the relevance matrix and the weight matrix to obtain a relevance operation value; when the relevance operation value is greater than the relevance threshold, outputting the label corresponding to the relevance operation value.

2. The method according to claim 1, characterized in that, the step of respectively obtaining the relevance between the content and the label obtained by two or more preset extraction algorithms includes: obtaining a first relevance between the content and the label obtained by a first extraction algorithm, where the first extraction algorithm is based on semantic matching; obtaining a second relevance between the content and the label obtained by a second extraction algorithm, where the second extraction algorithm is based on title matching; obtaining a third relevance between the content and the label obtained by a third extraction algorithm, where the third extraction algorithm is based on additional information matching.

3. The method according to claim 2, characterized in that, the step of obtaining a first relevance between the content and the label obtained by a first extraction algorithm, where the first extraction algorithm is based on semantic matching, includes: for the text content, after obtaining the keywords of the text content according to the keyword extraction algorithm, inputting the keywords into the word2vec model to obtain the vectors of the keywords; for the text content corresponding to the label, after obtaining the keywords of the text content according to the keyword extraction algorithm, then determining the core words and the related words thereof, inputting the core words and the related words into the word2vec model to obtain the word vectors of the core words and the related words, and performing a vector summation on the word vectors of the core words and the related words to obtain the vector of the label; performing a first preset operation on the vectors of the keywords and the vector of the label to obtain a first relevance between the keywords and the label.

4. The method according to claim 3, characterized in that: the related words are divided into highly related words and lowly related words according to the similarity.

5. The method according to claim 2, characterized in that, Obtaining the second relevance between the content obtained by the second extraction algorithm and the label, where the second extraction algorithm is based on title matching and includes: For the title of the text content, after obtaining the keywords of the title according to the keyword extraction algorithm, perform a second setting operation on the keywords of the title to obtain the second relevance between the title and the label.

6. The method according to claim 2, wherein, Obtaining the third relevance between the content obtained by the third extraction algorithm and the label, where the third extraction algorithm is based on additional information matching and includes: For the text content, after obtaining the additional information of the text content according to the keyword extraction algorithm, perform a third setting operation on the additional information to obtain the third relevance between the additional information and the label.

7. A label extraction device, wherein, comprising: An acquisition module for respectively obtaining the relevance between the content obtained by two or more set extraction algorithms and the label; The relevance is obtained by performing a setting operation on the keywords of the text content through the set extraction algorithm; An operation module for linearly summing the relevance between the content and the label respectively obtained by the acquisition module through a preset fusion algorithm to obtain a relevance operation value; wherein, the operation module includes: a relevance sub-module for obtaining a relevance matrix of the relevance between the content and the label according to the number of the labels and the number of the set extraction algorithms, the size of the relevance matrix is N×K, N is the number of the labels, K is the number of the set extraction algorithms, and the relevance matrix includes the relevance between the set extraction algorithm and the label; A weight sub-module for obtaining a preset weight matrix, where the weight matrix includes the weights of the set extraction algorithms in the labels, and wherein the weight matrix is preset in the following manner: for each set extraction algorithm, input a random evaluation content sample, obtain the relevance between the random evaluation content sample and the label, adjust the relevance threshold, and determine the weights of the set extraction algorithms in the labels according to the recall rate and accuracy rate of the label output evaluated each time; A calculation sub-module for linearly summing the relevance matrix and the weight matrix to obtain a relevance operation value; An output module for outputting the label corresponding to the relevance operation value when the relevance operation value obtained by the operation module is greater than the relevance threshold.

8. The device according to claim 7, wherein, The acquisition module includes: A first acquisition sub-module for obtaining the first relevance between the content obtained by the first extraction algorithm and the label, where the first extraction algorithm is based on semantic matching; A second acquisition sub-module for obtaining the second relevance between the content obtained by the second extraction algorithm and the label, where the second extraction algorithm is based on title matching; A third acquisition sub-module for obtaining the third relevance between the content obtained by the third extraction algorithm and the label, where the third extraction algorithm is based on additional information matching.

9. An electronic device, wherein, comprising: A processor; And A memory storing executable code thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1-6.

10. A computer-readable storage medium storing executable code thereon, which, when executed by a processor of an electronic device, causes the processor to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Label recommendation method and device, electronic equipment and storage medium

    CN112685642A