A keyword extraction, model training method, device, equipment and storage medium

By extracting and combining the global and local features of the target text, and using neural network models for keyword extraction, the problem of lack of semantic information and poor extraction of short text in the prior art is solved, and more accurate and extensive keyword extraction applications are achieved.

CN114201953BActive Publication Date: 2025-05-16BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111509488.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-05-16
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

The prior art lacks semantic information in keyword extraction, especially in the form of short text, making it difficult to accurately extract keywords, and depends on word frequency sorting or text clustering, limiting the breadth and depth of the application.

Method used

By extracting the global and local features of the target text, combining the neural network model for keyword extraction and model training, the overall semantics of the text are used to represent the overall semantics of the text, and the local features represent the character context semantics, improving the accuracy of keyword extraction.

Benefits of technology

It realizes the more accurate extraction of keywords in short text application scenarios, overcomes the problems of lack of semantic information and dependence on text quantity in the prior art, and improves the accuracy and application breadth of keyword extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114201953B_ABST
    Figure CN114201953B_ABST
Patent Text Reader

Abstract

The present disclosure provides a keyword extraction, model training method, device, equipment and storage medium, which relate to the field of data processing technology, especially to the field of content distribution technology and natural language processing technology. The above-mentioned keyword extraction scheme is: obtain the target text to be processed; extract the text features of the target text; perform global feature extraction on the text features of the target text to obtain the global features of the target text, wherein the global features of the target text are used to characterize the semantics expressed by the target text as a whole; perform local feature extraction on the text features of the target text to obtain the local features of the target text, wherein the local features of the target text are used to characterize the semantics expressed by the context of each character in the target text; extract the keywords of the target text based on the global features and local features of the target text. When the scheme provided by the embodiment of the present disclosure is applied to extract keywords, the accuracy of the extracted keywords can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to the field of content distribution technology and the field of natural language processing technology. Background Art

[0002] The keywords of a text refer to words related to the semantics expressed in the text. The keywords of a text can help people quickly understand the main content of the text. The keywords of a text play an important role in information retrieval, text clustering, etc. Summary of the invention

[0003] The present invention provides a keyword extraction, model training method, device, equipment and storage medium.

[0004] According to one aspect of the present disclosure, a keyword extraction method is provided, comprising:

[0005] Get the target text to be processed;

[0006] Extracting text features of the target text;

[0007] Performing global feature extraction on the text features of the target text to obtain the global features of the target text, wherein the global features of the target text are used to characterize the semantics expressed by the target text as a whole;

[0008] Performing local feature extraction on the text features of the target text to obtain the local features of the target text, wherein the local features of the target text are used to characterize the semantics expressed by the context in which each character in the target text is located;

[0009] Based on the global features and local features of the target text, keywords of the target text are extracted.

[0010] According to another aspect of the present disclosure, a model training method is provided, comprising:

[0011] Obtain sample text and real keywords of the sample text;

[0012] Input the sample text into a preset neural network model to obtain predicted keywords of the sample text, wherein the predicted keywords are keywords predicted based on global features and local features of the sample text, the global features of the sample text are features extracted from global features of the text features of the sample text, and the local features of the sample text are features extracted from local features of the text features of the sample text;

[0013] Based on the difference between the predicted keyword and the real keyword, the model parameters of the neural network model are adjusted.

[0014] According to another aspect of the present disclosure, there is provided a keyword extraction device, comprising:

[0015] A text acquisition module, used for acquiring a target text to be processed;

[0016] A feature extraction module, used to extract text features of the target text;

[0017] A global feature extraction module, used for performing global feature extraction on the text features of the target text to obtain the global features of the target text, wherein the global features of the target text are used to characterize the semantics expressed by the target text as a whole;

[0018] A local feature extraction module, used for extracting local features of the text features of the target text to obtain local features of the target text, wherein the local features of the target text are used to characterize the semantics expressed by the context in which each character in the target text is located;

[0019] The keyword extraction module is used to extract keywords of the target text based on the global features and local features of the target text.

[0020] According to another aspect of the present disclosure, a model training device is provided, comprising:

[0021] An information acquisition module is used to obtain sample text and real keywords of the sample text;

[0022] A keyword determination module, used for inputting the sample text into a preset neural network model to obtain predicted keywords of the sample text, wherein the predicted keywords are keywords predicted based on global features and local features of the sample text, the global features of the sample text are features extracted by performing global feature extraction on the text features of the sample text, and the local features of the sample text are features extracted by performing local feature extraction on the text features of the sample text;

[0023] A model parameter adjustment module is used to adjust the model parameters of the neural network model based on the difference between the predicted keyword and the real keyword.

[0024] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0025] at least one processor; and

[0026] a memory communicatively connected to the at least one processor; wherein,

[0027] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned keyword extraction method or model training method.

[0028] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the above-mentioned keyword extraction method or model training method.

[0029] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the above-mentioned keyword extraction method or model training method when executed by a processor.

[0030] By applying the solution provided by the embodiments of the present disclosure, the accuracy of the extracted keywords can be improved.

[0031] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0033] Figure 1 A schematic diagram of a flow chart of a first keyword extraction method provided in an embodiment of the present disclosure;

[0034] Figure 2 A schematic diagram of a flow chart of a second keyword extraction method provided in an embodiment of the present disclosure;

[0035] Figure 3 A schematic diagram of a flow chart of a third keyword extraction method provided in an embodiment of the present disclosure;

[0036] Figure 4a A flowchart of a fourth keyword extraction method provided in an embodiment of the present disclosure;

[0037] Figure 4b A schematic diagram of the structure of an encoder provided in an embodiment of the present disclosure;

[0038] Figure 5a A flowchart of a fifth keyword extraction method provided in an embodiment of the present disclosure;

[0039] Figure 5b A schematic diagram of the structure of the first keyword extraction model provided in the embodiment of the present disclosure;

[0040] Figure 6aA flowchart of a sixth keyword extraction method provided in an embodiment of the present disclosure;

[0041] Figure 6b A schematic diagram of the structure of a second keyword extraction model provided in an embodiment of the present disclosure;

[0042] Figure 7 A flowchart of a keyword recognition method provided by an embodiment of the present disclosure;

[0043] Figure 8 A flowchart of the first model training method provided in the embodiment of the present disclosure;

[0044] Fig. 9 A flowchart of a second model training method provided in an embodiment of the present disclosure;

[0045] Fig.10 A schematic diagram of the structure of a first keyword extraction device provided in an embodiment of the present disclosure;

[0046] Fig.11 A schematic diagram of the structure of a second keyword extraction device provided in an embodiment of the present disclosure;

[0047] Fig.12 A schematic diagram of the structure of a third keyword extraction device provided in an embodiment of the present disclosure;

[0048] Fig.13 A schematic diagram of the structure of a first model training device provided in an embodiment of the present disclosure;

[0049] Fig.14 A schematic diagram of the structure of a second model training device provided in an embodiment of the present disclosure;

[0050] Fig.15 It is a block diagram of an electronic device used to implement the keyword extraction method or model training method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0051] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0052] First, the application scenarios and execution entities of the embodiments of the present disclosure are described.

[0053] The application scenario of the embodiment of the present disclosure is: a scenario where keywords of a text need to be extracted. Specifically, it can be an application scenario of extracting keywords of a long text or an application scenario of extracting keywords of a short text.

[0054] Currently, keyword extraction is generally implemented based on word frequency sorting or text clustering. This method lacks semantic information and is highly dependent on the amount of text. It has significant limitations in its application to short text forms such as titles and short sentences.

[0055] The existing keyword extraction solutions are:

[0056] (1) Keyword extraction based on unsupervised methods. The unsupervised method is currently the most important method for keyword extraction. It refers to the use of unsupervised training methods to train a keyword extraction model. When extracting keywords, the trained keyword extraction model usually converts the task into a ranking problem.

[0057] (2) Keyword extraction based on supervised methods. Supervised methods refer to the use of supervised training methods to train keyword extraction models. The trained keyword extraction models usually regard keyword extraction as a classification problem and use the features and annotation information of candidate words to train the classifier.

[0058] In solution (1), keywords are usually extracted based on graph or subject clustering. The former uses candidate words as nodes and backward commonality relationships of words as edges, and recursively sorts them by calculating the weight scores of co-occurrence edges, thereby obtaining words with larger weights as the final keywords; the latter first clusters the text body and scores and sorts the candidate words based on the importance of the topic. Both are simple and easy to implement, but they rely heavily on the previous word segmentation model and text word frequency, and have poor application effects on short texts.

[0059] In solution (2), keyword extraction based on supervised methods relies on manually annotated a priori information, which is more complex and more expensive to implement than solution (1). This solution can currently be implemented by using a pre-training method based on neural networks or Transformer (encoder) to perform end-to-end joint modeling of sequence labeling tasks, avoiding the error propagation caused by multiple steps such as word segmentation. Although this solution is significantly improved over solution (1), the importance of keywords is heavily dependent on the annotated data, and tends to recognize words with greater weights and ignore words with smaller weights. The less important words are actually also in great demand in applications. In addition, existing keywords also lack attribute information. Therefore, the embodiment of the present disclosure provides a keyword extraction solution that can perform fine-grained keyword extraction on short texts. The method can achieve modeling of global and local information of the text, and integrate multi-task training to achieve keyword attribute, importance level and importance extraction.

[0060] The execution subject of the embodiment of the present disclosure is: an electronic device with a keyword extraction function, and the above electronic device can be a server, a terminal device, etc.

[0061] The keyword extraction method provided by the embodiment of the present disclosure is described in detail below.

[0062] See also Figure 1 , Figure 1 This is a flow chart of a first keyword extraction method provided in an embodiment of the present disclosure. The method includes the following steps S101-S105.

[0063] Step S101: Obtain target text to be processed.

[0064] The target text is a text that can be recognized by an electronic device. The target text is obtained by writing an original text in a natural language such as Chinese, English, or Japanese, converting the original text to obtain a converted text, and processing the converted text.

[0065] The so-called conversion of the original text means: setting the correspondence between characters and position identifiers in a preset dictionary, determining the position identifier corresponding to each character in the original text according to the correspondence, and replacing each character in the original text with the corresponding position identifier to obtain the converted text. The above position identifier can be a position serial number, a position name, etc.

[0066] The so-called processing of the converted text may include: removing punctuation marks in the converted text, processing the text length of the converted text, etc., so as to obtain the target text.

[0067] Specifically, after removing the punctuation marks of the conversion text, if the text length of the conversion text is less than the preset length, the conversion text without punctuation marks is padded with 0s so that the length of the padded text is the preset length, and the padded text is determined as the target text; after removing the punctuation marks of the conversion text, if the text length of the conversion text is greater than the preset length, the text before the preset length of the conversion text without punctuation marks is intercepted, and the intercepted text is determined as the target text; when after removing the punctuation marks of the conversion text, if the text length of the conversion text is equal to the preset length, the text length of the conversion text is not processed. The preset length can be preset by the staff, for example, the preset length can be 128, 256, etc.

[0068] In one embodiment of the present disclosure, the target text may be a short text. Therefore, the solution provided by the embodiment of the present disclosure can extract keywords from the short text.

[0069] Step S102: extracting text features of the target text.

[0070] The text features are used to characterize the features of the target text, and may include the semantic features of the target text, the semantic features of each character, and the like.

[0071] In one implementation, a text feature extraction algorithm may be used to extract text features from the target text to obtain text features of the target text.

[0072] The above-mentioned text feature extraction algorithms include: Word2Vec (word vector), Doc2Vec (sentence vector), etc.

[0073] Other implementations of extracting text features can be found in the following Figure 4a , Figure 5a The corresponding embodiments are not described in detail here.

[0074] Step S103: extracting global features of the target text to obtain global features of the target text.

[0075] The above-mentioned global features of the target text are used to characterize the semantics expressed by the target text as a whole.

[0076] The target text is a text formed by combining multiple sentences, and the target text as a whole refers to the entire text composed of the combination of multiple sentences included in the target text.

[0077] For example: suppose that the target text is "Keyword extraction generally refers to extracting the most relevant words in text expression from text data, and is a key step in knowledge extraction. Keyword extraction has important applications in information retrieval, text summarization, i.e. clustering, etc.". This target text includes two sentences. The overall text obtained by combining these two sentences represents the target text as a whole.

[0078] Based on the above analysis, the global features of the target text represent the semantics expressed by the target text from the perspective of the overall text of the target text.

[0079] In one implementation, a global feature extraction algorithm may be used to extract global features of the target text to obtain global features of the target text. The global feature extraction algorithm may be: IG (Information Gain), CE (Cross Entropy), WET (the Weight of Evidence for Tex), etc.

[0080] Other implementation methods for global feature extraction of text features can be found in the following Figure 4a , 5a The corresponding embodiments are not described in detail here.

[0081] Step S104: extracting local features of the target text to obtain local features of the target text.

[0082] The local features of the target text are used to characterize the semantics expressed by the context of each character in the target text. The context of each character is a local text in the target text, so the local features of the target text can reflect the semantic features of the local text in the target text.

[0083] In one implementation, a local feature extraction algorithm may be used to extract local features of the target text to obtain local features of the target text. The local feature extraction algorithm may be: MI (Mutual Information), χ 2 Estimation, etc.

[0084] Other implementations of local feature extraction for text features can be found in the following Figure 4a , Figure 5a The corresponding embodiments are not described in detail here.

[0085] Specifically, the above steps S103 - S104 may be steps executed in parallel, or may be steps executed in a preset order, which is not limited in the embodiment of the present invention.

[0086] Step S105: extracting keywords of the target text based on the global features and local features of the target text.

[0087] Since the global features reflect the semantics expressed by the target text as a whole, the overall semantic information of the target text is strengthened, but the local semantic information of the target text is weakened. Since the local features reflect the semantics expressed by the context of each character in the target text, and the context of each character is the local information in the target text, the local features can better reflect the local semantic information in the target text. In this way, when extracting keywords based on the global features and local features of the target text, the problem of the local semantic information of the target text being weakened by the global features can be overcome, and the keywords of the target text can be extracted based on richer and more comprehensive semantic information.

[0088] Furthermore, in the case where the target text is a short text, when the solution of the embodiment of the present disclosure is adopted, since the keywords of the target text are extracted based on the global features and local features of the target text, the global features and local features can reflect the semantic information of the target text in a rich and comprehensive manner, and the semantic information of the target text is fully considered when extracting the keywords of the short text. Compared with the prior art that relies on the frequency of occurrence of each word to extract the keywords of the target text, the accuracy of the extracted keywords can be effectively improved. Therefore, the solution provided by the embodiment of the present disclosure can more accurately extract the keywords of the short text in the application scenario of extracting the keywords of the short text.

[0089] In one implementation, when extracting keywords of a target text, the global features and local features of the target text may be fused to obtain a first fused feature; and the keywords of the target text may be extracted based on the first fused feature.

[0090] Specifically, when performing feature fusion, the global features and local features of the target text may be spliced ​​together, and the spliced ​​features may be used as the first fused features.

[0091] When extracting keywords of the target text, a keyword recognition algorithm may be used to perform keyword recognition on the first fusion feature to obtain keywords of the target text. The keyword recognition algorithm may be CRF (Conditional Random Field) or HMM (Hiden Markov Mode).

[0092] Since the first fusion feature is a feature obtained by fusing the global feature and the local feature of the target text, the first fusion feature can more completely include the global feature and the local feature of the target text. Therefore, when performing keyword extraction based on the first fusion feature, the global feature and the local feature of the target text can be better referenced.

[0093] Other ways to extract keywords from the target text can be found in the following Figure 5a The corresponding embodiments are not described in detail here.

[0094] When obtaining the keywords of the target text, since the word type of the keywords also needs to be known in some specific scenarios, based on this, in one embodiment of the present disclosure, the word type of the keywords of the target text can be determined from various preset word types based on the global features and local features of the target text.

[0095] The above-mentioned preset word types may be: names of people, places or organizational activities, etc.

[0096] The above-mentioned preset word types can also be further divided into fine-grained categories. For example, for names of people, the preset word types can be divided into names of specific references, names of general references, names of fuzzy references, etc.

[0097] In one implementation, a sequence labeling task may be used to determine the word type of the keyword of the target text based on the global features and local features of the target text.

[0098] Since the global and local features of the target text are considered when determining the word type of the keyword, the global and local features of the target text can fully reflect the word type information of each word in the target text. Therefore, based on the global and local features, a more accurate word type of the keyword can be obtained.

[0099] Moreover, in the solution provided by the embodiment of the present disclosure, the word type of the keyword can be determined more accurately based on the global features and local features of the target text. However, compared with extracting keywords based on the frequency of word appearance in the prior art, the frequency of word appearance is difficult to reflect the word type information of the word. Therefore, the prior art center finds it difficult to obtain the word type information of the keyword when extracting keywords based on the frequency of word appearance.

[0100] As can be seen from the above, in the scheme provided by the embodiment of the present disclosure, the keywords of the target text are extracted based on the global features and local features of the target text. Since the global features of the target text represent the semantics expressed by the target text as a whole, and the local features of the target text represent the semantics expressed by the context of each character in the target text, when extracting the keywords of the target text, both the semantics expressed by the target text as a whole and the semantics expressed by the context of each character in the target text are considered. Moreover, since the semantics expressed by the target text as a whole and the semantics expressed by the context of each character in the target text can fully and comprehensively represent the semantic information of the target text, therefore, by adopting the scheme provided by the embodiment of the present disclosure, the keywords of the target text are extracted based on the comprehensive and rich semantic information of the target text, thereby improving the accuracy of the extracted keywords.

[0101] In the case where the target text is a short text, the global features and local features of the target text are referenced when extracting the keywords of the target text, that is, the solution provided by the embodiment of the present disclosure obtains the keywords of the target text on the basis of fully referring to the semantic information of the target text. Compared with the prior art that extracts keywords based on the frequency of occurrence of each word, the accuracy of the extracted keywords is significantly improved. Therefore, the solution provided by the embodiment of the present disclosure can accurately extract the keywords of the short text in the application scenario of extracting the keywords of the short text.

[0102] In the case of obtaining multiple keywords, the importance of each keyword is different. In order to know the importance of each keyword, based on this, in one embodiment of the present disclosure, refer to Figure 2 , Figure 2 This is a flow chart of a second keyword extraction method provided in an embodiment of the present disclosure. The method includes the following steps S201-S206.

[0103] Step S201: Obtain target text to be processed.

[0104] Step S202: extracting text features of the target text.

[0105] Step S203: extracting global features of the target text to obtain global features of the target text.

[0106] The above-mentioned global features of the target text are used to characterize the semantics expressed by the target text as a whole.

[0107] Step S204: extracting local features of the target text to obtain local features of the target text.

[0108] The above local features of the target text are used to represent the semantics expressed by the context of each character in the target text.

[0109] Step S205: extracting keywords of the target text based on the global features and local features of the target text.

[0110] The above steps S201-S205 are respectively Figure 1 In the illustrated embodiment, steps S101 - S105 are the same and will not be described in detail herein.

[0111] Step S206: predicting the first importance of the keyword based on the local features of the target text.

[0112] The first importance reflects the importance of the keyword. When the first importance is higher, it means that the keyword is more important than other keywords; when the first importance is lower, it means that the keyword is less important than other keywords.

[0113] In one implementation, an importance prediction algorithm may be used to determine the first importance of a keyword based on local features of the target text. The importance prediction algorithm may include: WOE (Weight of Evidence), IV (Information Value), and the like.

[0114] Other implementations of predicting the first importance can be found in Figure 3 The corresponding embodiments are not described in detail here.

[0115] As can be seen from the above, when predicting the first importance of a keyword, it is predicted based on the local features of the target text. Since the local features of the target text are used to characterize the semantics expressed by the context of each character in the target text, the predicted first importance is related to the semantics expressed by the context of each character in the target text, so the first importance takes into account the local semantic information of the target text. And since the importance of each word in the target text is related to the local semantic information of the target text. Therefore, based on the local features of the target text, the accuracy of the predicted first importance of the keyword can be improved.

[0116] In the above Figure 2 In step S206 of the embodiment shown, in addition to using the importance prediction algorithm to predict the first importance of the keyword, Figure 3 In the illustrated embodiment, steps S306-S309 are implemented.

[0117] See also Figure 3 , Figure 3 This is a flow chart of a third keyword extraction method provided in an embodiment of the present disclosure. The method includes the following steps S301-S309.

[0118] Step S301: Obtain the target text to be processed.

[0119] Step S302: extracting text features of the target text.

[0120] Step S303: extracting global features of the target text to obtain global features of the target text.

[0121] The above-mentioned global features of the target text are used to characterize the semantics expressed by the target text as a whole.

[0122] Step S304: extracting local features of the text features of the target text to obtain local features of the target text.

[0123] The above local features of the target text are used to represent the semantics expressed by the context of each character in the target text.

[0124] Step S305: extracting keywords of the target text based on the global features and local features of the target text.

[0125] The above steps S301-S305 are respectively Figure 1 In the illustrated embodiment, steps S101 - S105 are the same and will not be described in detail herein.

[0126] Step S306: Based on the local features of the target text, predict the probability that the importance level of the keyword is each preset importance level.

[0127] The preset importance levels may be set by the staff based on experience, for example, the preset importance levels may include: very important, generally important, not very important, etc.

[0128] The above-mentioned probability represents the probability that the importance level of the keyword is each preset importance level. For example, when the probability that the importance level of the keyword is "very important level" is high, the probability corresponding to the keyword is high, indicating that the importance level of the keyword is likely to be "very important" level.

[0129] The implementation method of predicting the probability corresponding to a keyword can be found in the subsequent embodiments and will not be described in detail here.

[0130] Step S307: Based on the probability corresponding to the keyword, determine the target importance level of the keyword from various preset importance levels.

[0131] Specifically, the target importance level may be determined in the following two ways.

[0132] In one implementation, a preset importance level with the greatest possible degree may be determined as the target importance level of the keyword.

[0133] In another implementation, an average value of the likelihoods corresponding to the keywords may be calculated, and a preset importance level with a likelihood closest to the average value may be determined as the target importance level of the keyword.

[0134] Since the target importance level of the keyword is determined from each preset importance level, the target importance levels of different keywords may be different. For example, the importance level of a certain keyword is higher than the importance level of another keyword. Therefore, in the scheme provided by the embodiment of the present disclosure, keywords of various importance levels can be obtained. Since relatively unimportant keywords also have a large demand in practical applications, by determining the importance levels of different keywords, corresponding services can be implemented based on the importance levels of the keywords when the extracted keywords are subsequently applied.

[0135] Step S308: Based on the possibility corresponding to the keyword, calculate the second importance of the keyword at the target importance level.

[0136] The second importance reflects the importance of the keyword at the target importance level. When different keywords have the same importance level, the importance of each keyword can be effectively determined based on the second importance of the keyword.

[0137] Specifically, according to the weight corresponding to each preset importance level, the probability that the importance level of the keyword is each preset importance level may be weighted and summed, and the obtained sum value may be determined as the second importance.

[0138] For example, the second importance can be calculated according to the following formula:

[0139] score=(C[1,2]*a+C[1,1]*b)÷(a+b)

[0140] Among them, score represents the calculated second importance, C[1,2] and C[1,1] respectively represent the possibility that the importance level of the keyword is each preset importance level, and a and b respectively represent the weight corresponding to each preset importance level.

[0141] Step S309: Determine the first importance of the keyword based on the target importance level and the second importance of the keyword.

[0142] Since the target importance level of a keyword indicates the importance level of the keyword, and the second importance level of the keyword reflects the importance of the keyword under the target importance level, the importance of the keyword can be accurately reflected based on the importance level of the keyword and the importance level under the target importance level. Therefore, the first importance level of the keyword can be determined based on the target importance level and the second importance level of the keyword.

[0143] In one implementation, the level number corresponding to each preset importance level is determined in advance, the level number corresponding to the target importance level of the keyword is determined, and the determined level number is used as the integer part of the first importance; and the second importance is normalized, and the data obtained after the normalization is used as the decimal part of the first importance, so that the above integer part and decimal part are integrated to obtain the first importance.

[0144] For example: the target importance level of the keyword is "general importance level", the level number corresponding to "general importance level" is 1, which is the integer part of the first importance; the result of normalization of the second importance of the keyword is 0.067, which is the decimal part of the first importance. The above integer part and decimal part are integrated to get 1.067, which is the first importance.

[0145] From the above, it can be seen that the first importance of the keyword is determined based on the target importance level and the second importance of the keyword, wherein the target importance level of the keyword represents the importance level of the keyword, and the second importance of the keyword reflects the importance of the keyword at the target importance level. Since the importance of the keyword can be accurately reflected based on the importance level of the keyword and the importance at the target importance level, the accuracy of the determined first importance is improved.

[0146] Furthermore, since the first importance of the keyword is determined based on the target importance level and the second importance of the keyword, for different keywords of the same importance level, the importance of each keyword can be accurately determined based on the second importance of the keyword.

[0147] In the above Figure 3 When predicting the probability corresponding to the keyword in step S306 of the illustrated embodiment, it can be implemented according to the following steps A1-A3.

[0148] Step A1: Based on the local features of the target text, predict the probability that the importance level of each character in the target text is each preset importance level.

[0149] The above-mentioned likelihood reflects the likelihood that the importance level of each character in the target text is each preset importance level.

[0150] The local features of the target text contain importance information of each character in the target text. Based on this, the probability of the importance level of each character being each preset importance level can be predicted based on the importance information of each character contained in the local features of the target text.

[0151] Step A2: Determine the probability degree corresponding to each character in the keyword from the probability degree corresponding to each character in the target text.

[0152] In one embodiment, when the possibility degree corresponding to each character in the target text is predicted in the above step A1, the correspondence between the position identifier of each character in the target text at the location of the target text and the possibility degree can be determined. Based on this, when determining the possibility degree corresponding to each character in the keyword, the possibility degree corresponding to the position identifier of each character in the keyword at the location of the target text can be determined based on the above correspondence, as the possibility degree corresponding to each character in the keyword.

[0153] For example, the corresponding relationship between the position identifier and the possibility degree of each character in the target text is shown in Table 1 below.

[0154] Table 1

[0155] Location ID Very important General Important Not very important 0 0.6 0.8 0.1 1 0.7 0.8 0.2 2 0.7 0.9 0.2 3 0.8 0.7 0.3 4 0.3 0.4 0.8 5 0.5 0.6 0.9

[0156] Taking the first row of data in Table 1 as an example, a position marker of 0 indicates that the character is the first character in the target text, 0.6 indicates the possibility that the importance level of the character is "very important", 0.8 indicates the possibility that the importance level of the character is "generally important", and 0.1 indicates the possibility that the importance level of the character is "not very important".

[0157] The position identifiers of the positions of each character in the keyword in the target text are 2 and 3 respectively. Therefore, based on Table 1 above, it can be determined that the probability corresponding to the first character in the keyword is (0.7, 0.9, 0.2), and the probability corresponding to the second character in the keyword is (0.8, 0.7, 0.3).

[0158] Step A3: For each preset importance level, statistical analysis is performed on the probability that the importance level of each character in the keyword is the preset importance level, and the statistical analysis result is determined as the probability that the importance level of the keyword is each preset importance level.

[0159] Since a keyword includes multiple characters, the probability degree corresponding to the keyword is related to the probability degree corresponding to each character. Therefore, the probability degree corresponding to the keyword needs to be determined based on the probability degree corresponding to each character.

[0160] The above statistical analysis methods may include: calculating the average value, taking the median value, calculating the sum value, etc.

[0161] Taking the calculation of the sum value as an example, following the example shown in step A2, for the "very important level", the sum of 0.7 and 0.8 is calculated as the possibility that the keyword's importance level is "very important level", which is 1.5; for the "general importance level", the sum of 0.9 and 0.7 is calculated as the possibility that the keyword's importance level is "general importance level", which is 1.6; for the "not so important level", the sum of 0.2 and 0.3 is calculated as the possibility that the keyword's importance level is "not so important level", which is 0.5.

[0162] From the above, it can be seen that since the keyword contains multiple characters, the possibility corresponding to the keyword is related to the possibility corresponding to each character. Therefore, for each preset importance level, a statistical analysis is performed on the possibility that the importance level of each character in the keyword is the preset importance level. The result obtained by the statistical analysis can accurately reflect the possibility that the importance level of the keyword is the preset importance level, so as to obtain a more accurate possibility that the importance level of the keyword is each preset importance level.

[0163] In the aforementioned Figure 1In step S102 of the illustrated embodiment, in addition to using a preset text feature extraction algorithm to extract text features of the target text, the following Figure 4a In the illustrated embodiment, step S402 is implemented.

[0164] See also Figure 4a , Figure 4a This is a flow chart of a fourth keyword extraction method provided in an embodiment of the present disclosure. The method includes the following steps S401-S406.

[0165] Step S401: Obtain the target text to be processed.

[0166] The above step S401 is the same as the above Figure 1 Step S101 in the illustrated embodiment is the same and will not be described in detail here.

[0167] Step S402: Based on the context information of each character in the target text, each character is encoded using a multi-layer encoding method to obtain a preset number of layer feature vectors as text features of the target text.

[0168] In one implementation, the target text may be input into the encoder having multiple network layers, and the encoder encodes each character based on the context information of each character in the target text to obtain a feature vector of a preset number of layers.

[0169] by Figure 4b For example, Figure 4b A schematic diagram of an encoder structure is shown. The encoder includes N layers of network layers, and the input information of each network layer is the output information of the previous network layer of the network layer. The network layer character encodes the input information to obtain an encoding result. The encoding result obtained by multiple network layers is called a preset number of layer feature vectors.

[0170] The above encoder can be BERT (Bidirectional Encoder Representation from Transformers), ALBERT (A Lite BERT), etc.

[0171] Based on the above step S402, Figure 1 In the illustrated embodiment, S103 can be implemented according to the following steps S403.

[0172] Step S403: extracting the feature vector of the last layer among the feature vectors of a preset number of layers, and determining it as the global feature of the target text.

[0173] The above-mentioned global features of the target text are used to characterize the semantics expressed by the target text as a whole.

[0174] In the multi-layer feature vector, the feature vector of the last layer is used to represent the global features of the text. Therefore, the feature vector of the last layer in the feature vector of the preset number of layers can be determined as the global features of the target text.

[0175] For example: In the above Figure 4b In the encoder structure diagram shown, the Nth layer feature vector output by the Nth layer network layer is the feature vector of the last layer, and the feature vector output by the last layer network layer is determined as the global vector of the target text.

[0176] Based on the above step S402, Figure 1 In the illustrated embodiment, S104 may be implemented according to the following steps S404-S405.

[0177] Step S404: extracting feature vectors from feature vectors of a preset number of layers except the feature vector of the last layer, and performing feature fusion on the extracted feature vectors to obtain a second fused feature.

[0178] In the multi-layer feature vectors, feature vectors other than the feature vectors of the last layer can represent the local features of the text. Therefore, the local features of the target text can be determined based on the feature vectors other than the feature vectors of the last layer.

[0179] In one implementation, feature vectors other than the feature vector of the last layer may be extracted, and the extracted feature vectors may be concatenated to obtain a second fused feature. Since the second fused feature includes multiple layers of text features, and the text features may reflect the semantic information of the target text, the second fused feature may be called a hierarchical semantic feature.

[0180] For example: In the above Figure 4b In the encoder structure diagram shown, the feature vectors output by the 1st, 2nd, ..., N-1th network layers are: the feature vectors except the feature vector of the last layer are concatenated to obtain hierarchical semantic features.

[0181] Step S405: extract local features from the second fused features to obtain local features of the target text.

[0182] The above local features of the target text are used to represent the semantics expressed by the context of each character in the target text.

[0183] In one implementation, a convolutional neural network can be used to extract local features from the second fused features. The convolutional neural network has the characteristics of local perception, can extract local features from the second fused features, and can achieve parameter sharing, reducing the number of parameters and reducing the complexity of feature extraction.

[0184] Step S406: extracting keywords of the target text based on the global features and local features of the target text.

[0185] The above step S406 is similar to the above Figure 1 Step S105 in the illustrated embodiment is the same and will not be described in detail here.

[0186] From the above, it can be seen that since the text features of the target text are obtained based on the context information of each character in the target text, the context information of each character can accurately reflect the semantic information of each character. Since the semantic information of the target text is related to the semantic information of each character, the text features of the target text obtained can accurately reflect the semantic information of the target text.

[0187] Moreover, in the multi-layer feature vectors, the feature vectors of the last layer can accurately represent the global features of the text, and the feature vectors other than the feature vectors of the last layer can accurately represent the local features of the text. Therefore, by determining the feature vectors of the last layer as the global features of the target text, more accurate global features can be obtained, and by determining the local features of the target text based on the feature vectors other than the feature vectors of the last layer, more accurate local features can be obtained.

[0188] In the aforementioned Figure 1 In the steps S102, S103, S104 and S105 shown, each network layer in the pre-trained keyword extraction model can be used for implementation. Based on this, in one embodiment of the present disclosure, see Figure 5a , Figure 5a This is a flowchart of a fifth keyword extraction method provided in an embodiment of the present disclosure. The method includes the following steps S501-S505.

[0189] Step S501: Obtain the target text to be processed.

[0190] The above step S501 is the same as the above Figure 1 Step S101 in the illustrated embodiment is the same and will not be described in detail here.

[0191] Step S502: input the target text into the text feature extraction layer in the pre-trained keyword extraction model to obtain the text features of the target text.

[0192] The above text feature extraction layer is used to extract text features of text.

[0193] After the target text is input into the above text feature extraction layer, the text feature extraction layer may perform character encoding based on the context information of each character in the target text to obtain the text features of the target text.

[0194] Step S503: inputting the text features of the target text into the global feature extraction layer in the keyword extraction model to obtain the global features of the target text.

[0195] The above-mentioned global features of the target text are used to characterize the semantics expressed by the target text as a whole.

[0196] The global feature extraction layer is used to extract global features of text features.

[0197] After the text features of the target text are input into the above-mentioned global feature extraction layer, the global feature extraction layer can perform global feature extraction on the text features of the target text to obtain the global features of the target text.

[0198] Step S504: inputting the text features of the target text into the local feature extraction layer in the keyword extraction model to obtain the local features of the target text.

[0199] The above local features of the target text are used to represent the semantics expressed by the context of each character in the target text.

[0200] The local feature extraction layer is used to extract local features of text features.

[0201] After the text features of the target text are input into the above-mentioned local feature extraction layer, the local feature extraction layer can perform local feature extraction on the text features of the target text to obtain the local features of the target text.

[0202] Step S505: inputting the global features and local features of the target text into the keyword extraction layer in the keyword extraction model to obtain keywords of the target text.

[0203] The keyword extraction layer is used to extract keywords based on the global and local features of the text.

[0204] After the global features and local features of the target text are input into the keyword extraction layer, the keyword extraction layer can extract keywords of the target text based on the global features and local features of the target text.

[0205] From the above, it can be seen that since the pre-trained keyword extraction model is trained based on a large amount of sample text, the keyword extraction model learns the ability to identify keywords based on sample text. Therefore, when extracting keywords from the target text based on each network layer in the above keyword extraction model, the accuracy of the extracted keywords can be improved.

[0206] The following is Figure 5b Taking the structural diagram of the keyword extraction model shown in FIG. 1 as an example, the above keyword extraction process is explained.

[0207] Figure 5bThe keyword extraction model shown includes: a text feature extraction layer, a global feature extraction layer, a local feature extraction layer, and a keyword extraction layer.

[0208] Target text is entered first Figure 5b The text feature extraction layer in the keyword extraction model shown in the figure extracts text features from the target text and inputs the extracted text features into the global feature extraction layer and the local feature extraction layer;

[0209] The global feature extraction layer extracts global features from text features to obtain global features of the target text, and inputs the global features into the keyword extraction layer;

[0210] The local feature extraction layer is used to extract local features of text features, obtain local features of the target text, and input the local features into the keyword extraction layer;

[0211] The keyword extraction layer extracts keywords from the global features and local features of the target text, and outputs the extracted keywords as the keywords of the target text.

[0212] With the above Figure 2 Corresponding to the embodiment shown, when multiple keywords are obtained, the importance of each keyword is different. In order to know the importance of each keyword, in one embodiment of the present disclosure, refer to Figure 6a , Figure 6a This is a flow chart of a sixth keyword extraction method provided in an embodiment of the present disclosure. The method includes the following steps S601-S606.

[0213] Step S601: Obtain the target text to be processed.

[0214] Step S602: input the target text into the text feature extraction layer of the pre-trained keyword extraction model to obtain the text features of the target text.

[0215] Step S603: inputting the text features of the target text into the global feature extraction layer in the keyword extraction model to obtain the global features of the target text.

[0216] The above-mentioned global features of the target text are used to characterize the semantics expressed by the target text as a whole.

[0217] Step S604: inputting the text features of the target text into the local feature extraction layer in the keyword extraction model to obtain the local features of the target text.

[0218] The above local features of the target text are used to represent the semantics expressed by the context of each character in the target text.

[0219] Step S605: inputting the global features and local features of the target text into the keyword extraction layer in the keyword extraction model to obtain keywords of the target text.

[0220] The above steps S601-S605 are the same as the above steps S501-S505, and will not be described in detail here.

[0221] Step S606: Input the local features of the target text into the importance determination layer in the keyword extraction model to obtain the first importance of the keyword.

[0222] The importance determination layer is used to predict the importance of keywords in a text based on local features of the text.

[0223] When the local features of the target text are input into the importance determination layer, the importance determination layer is used to predict the first importance of the keyword based on the local features of the target text.

[0224] From the above, it can be seen that since the keyword extraction model also includes an importance determination layer, which is used to predict the keyword importance of the text based on the local features of the text, after the local features of the target text are input into the keyword extraction model, a more accurate first importance of the keyword can be obtained.

[0225] The following is Figure 6b Taking the structural diagram of the keyword extraction model shown in FIG. 1 as an example, the above keyword extraction process is explained.

[0226] Figure 6b The keyword extraction model shown includes: a text feature extraction layer, a global feature extraction layer, a local feature extraction layer, a keyword extraction layer, and an importance determination layer.

[0227] Target text is entered first Figure 6b The text feature extraction layer of the keyword extraction model extracts text features from the target text and inputs the extracted text features into the global feature extraction layer and the local feature extraction layer;

[0228] The global feature extraction layer extracts global features from text features to obtain global features of the target text, and inputs the global features into the keyword extraction layer;

[0229] The local feature extraction layer extracts local features of the text to obtain local features of the target text, and inputs the local features into the keyword extraction layer and the importance determination layer respectively;

[0230] The keyword extraction layer extracts keywords from the global features and local features of the target text, and outputs the extracted keywords as the keywords of the target text;

[0231] The importance determination layer determines the importance of keywords based on the local features of the target text.

[0232] The following combination Figure 7 , the process of the above keyword identification is explained in detail.

[0233] Figure 7 A flowchart of a keyword recognition method provided in an embodiment of the present disclosure. Figure 7 The following steps are included: S701-S712.

[0234] S701: Obtaining the original text to be processed;

[0235] S702: Convert the original text to obtain the token_id corresponding to each character in the original text.

[0236] The token_id above indicates the position identifier of each character in the original text in the preset dictionary library.

[0237] S703: Determine whether the length of the converted text is less than 128, if yes, execute S704, if no, execute S705.

[0238] S704: padding the converted text by adding 0s to obtain a text with a length of 128 as the target text.

[0239] S705: intercepting the first 128 characters of the converted text as the target text.

[0240] S706: Encode each character based on the context information of each character in the target text to obtain a feature vector of a preset number of layers.

[0241] S707: Extract the last layer of feature vectors from a preset number of layers of feature vectors as the global features of the target text.

[0242] S708: Extract feature vectors except the feature vector of the last layer from the feature vectors of a preset number of layers, and perform feature concatenation on the extracted feature vectors.

[0243] S709: Perform local feature extraction on the concatenated features to obtain local features of the target text.

[0244] S710: performing feature splicing on the global features and local features of the target text.

[0245] S711: Perform keyword recognition and keyword word type recognition on the features obtained by concatenating in S710 to obtain keywords and keyword word types of the target text.

[0246] S712: Based on the local features of the target text, determine the importance level of the keyword, and determine the importance of the keyword under the importance level, and determine the importance level and the importance under the importance level as the final importance of the keyword.

[0247] See also Figure 8 , Figure 8 A flowchart of the first model training method provided in an embodiment of the present disclosure, the method includes the following steps S801-S803.

[0248] Step S801: Obtain sample text and real keywords of the sample text.

[0249] Step S802: Input the sample text into a preset neural network model to obtain predicted keywords of the sample text.

[0250] The above-mentioned predicted keywords are keywords predicted based on the global features and local features of the sample text.

[0251] The global features of the sample text are: features obtained by performing global feature extraction on the text features of the sample text.

[0252] The local features of the sample text are: features obtained by performing local feature extraction on the text features of the sample text.

[0253] Step S803: Adjusting the model parameters of the neural network model based on the difference between the predicted keywords and the real keywords.

[0254] When adjusting the model parameters of the neural network model, the difference between the predicted keyword and the real keyword can be calculated, the loss value of the neural network model can be determined based on the calculated difference, and the model parameters of the neural network model can be adjusted based on the above loss value.

[0255] From the above, it can be seen that since the model parameters of the neural network model are adjusted based on the difference between the predicted keywords and the real keywords, the difference between the predicted keywords and the real keywords reflects the ability of the neural network model to extract keywords. Therefore, based on the difference between the predicted keywords and the real keywords, the ability of the neural network model to extract keywords can be effectively improved.

[0256] See also Fig. 9 , Fig. 9 A flowchart of a second model training method provided in an embodiment of the present disclosure, the method includes the following steps S901-S904.

[0257] Step S901: Obtain sample text and real keywords of the sample text.

[0258] Step S902: Obtain the real importance of the real keywords of the sample text.

[0259] Step S903: Input the sample text into a preset neural network model to obtain the predicted keywords of the sample and the predicted importance of the predicted keywords.

[0260] The above-mentioned predicted keywords are keywords predicted based on the global features and local features of the sample text.

[0261] The global features of the sample text are: features obtained by performing global feature extraction on the text features of the sample text.

[0262] The local features of the sample text are: features obtained by performing local feature extraction on the text features of the sample text.

[0263] The above-mentioned predicted importance is: the importance obtained by performing importance prediction based on the local features of the sample text.

[0264] Step S904: Adjust the model parameters of the neural network model based on the difference between the predicted keywords and the real keywords, and the difference between the predicted importance and the real importance of the same keyword.

[0265] In one implementation, a first loss value of the neural network model can be calculated based on the difference between the predicted keywords and the actual keywords, and a second loss value of the neural network model can be calculated based on the difference between the predicted importance and the actual importance of the same keyword. Based on the first loss value and the second loss value, the model parameters of the neural network model are adjusted.

[0266] Specifically, the first loss value and the second loss value may be weightedly summed according to the weight corresponding to the first loss value and the weight corresponding to the second loss value, and the model parameters of the neural network model may be adjusted based on the sum obtained by the weighted summation.

[0267] Since the model parameters of the neural network model are adjusted based on the difference between predicted keywords and real keywords, and the difference between the predicted importance and the real importance of the same keyword, the difference between the predicted keywords and the real keywords reflects the ability of the neural network model to extract keywords, and the difference between the predicted importance and the real importance of the same keywords reflects the ability of the neural network model to predict the importance of keywords. Therefore, based on the above two differences, the ability of the neural network model to extract keywords and predict the importance of keywords can be effectively improved.

[0268] Corresponding to the above-mentioned keyword extraction method, the embodiment of the present disclosure also provides a keyword extraction device.

[0269] See also Fig.10 , Fig.10This is a schematic diagram of the structure of a first keyword extraction device provided in an embodiment of the present disclosure. The device includes the following modules 1001-1005.

[0270] The text obtaining module 1001 is used to obtain the target text to be processed;

[0271] A feature extraction module 1002 is used to extract text features of the target text;

[0272] A global feature extraction module 1003 is used to extract global features of the text features of the target text to obtain global features of the target text, wherein the global features of the target text are used to characterize the semantics expressed by the target text as a whole;

[0273] A local feature extraction module 1004 is used to extract local features of the text features of the target text to obtain local features of the target text, wherein the local features of the target text are used to represent the semantics expressed by the context of each character in the target text;

[0274] The keyword extraction module 1005 is used to extract keywords of the target text based on the global features and local features of the target text.

[0275] As can be seen from the above, in the scheme provided by the embodiment of the present disclosure, the keywords of the target text are extracted based on the global features and local features of the target text. Since the global features of the target text represent the semantics expressed by the target text as a whole, and the local features of the target text represent the semantics expressed by the context of each character in the target text, when extracting the keywords of the target text, both the semantics expressed by the target text as a whole and the semantics expressed by the context of each character in the target text are considered. Moreover, since the semantics expressed by the target text as a whole and the semantics expressed by the context of each character in the target text can fully and comprehensively represent the semantic information of the target text, therefore, by adopting the scheme provided by the embodiment of the present disclosure, the keywords of the target text are extracted based on the comprehensive and rich semantic information of the target text, thereby improving the accuracy of the extracted keywords.

[0276] See also Fig.11 , Fig.11 This is a schematic diagram of the structure of a second keyword extraction device provided in an embodiment of the present disclosure. The device includes the following modules 1101-1106.

[0277] The text obtaining module 1101 is used to obtain the target text to be processed;

[0278] A feature extraction module 1102 is used to extract text features of the target text;

[0279] A global feature extraction module 1103 is used to extract global features of the text features of the target text to obtain global features of the target text, wherein the global features of the target text are used to characterize the semantics expressed by the target text as a whole;

[0280] A local feature extraction module 1104 is used to extract local features of the text features of the target text to obtain local features of the target text, wherein the local features of the target text are used to represent the semantics expressed by the context of each character in the target text;

[0281] A keyword extraction module 1105 is used to extract keywords of the target text based on the global features and local features of the target text;

[0282] The importance prediction module 1106 is used to predict the first importance of the keyword based on the local features of the target text.

[0283] As can be seen from the above, when predicting the first importance of a keyword, it is predicted based on the local features of the target text. Since the local features of the target text are used to characterize the semantics expressed by the context of each character in the target text, the predicted first importance is related to the semantics expressed by the context of each character in the target text, so the first importance takes into account the local semantic information of the target text. And since the importance of each word in the target text is related to the local semantic information of the target text. Therefore, based on the local features of the target text, the accuracy of the predicted first importance of the keyword can be improved.

[0284] See also Fig.12 , Fig.12 This is a schematic diagram of the structure of a third keyword extraction device provided in an embodiment of the present disclosure. The device includes the following modules 1201-1209.

[0285] The text obtaining module 1201 is used to obtain the target text to be processed;

[0286] A feature extraction module 1202 is used to extract text features of the target text;

[0287] A global feature extraction module 1203 is used to extract global features of the text features of the target text to obtain global features of the target text, wherein the global features of the target text are used to characterize the semantics expressed by the target text as a whole;

[0288] A local feature extraction module 1204 is used to extract local features of the text features of the target text to obtain local features of the target text, wherein the local features of the target text are used to represent the semantics expressed by the context of each character in the target text;

[0289] A keyword extraction module 1205 is used to extract keywords of the target text based on the global features and local features of the target text;

[0290] A possibility prediction submodule 1206 is used to predict the possibility that the importance level of the keyword is each preset importance level based on the local features of the target text;

[0291] The importance level determination submodule 1207 is used to determine the target importance level of the keyword from various preset importance levels based on the possibility corresponding to the keyword;

[0292] The importance calculation submodule 1208 is used to calculate the second importance of the keyword at the target importance level based on the possibility corresponding to the keyword;

[0293] The importance determination submodule 1209 is used to determine the first importance of the keyword based on the target importance level and the second importance of the keyword.

[0294] As can be seen from the above, the first importance of the keyword is determined based on the target importance level and the second importance of the keyword, wherein the target importance level of the keyword indicates the importance level of the keyword, and the second importance of the keyword reflects the importance of the keyword under the target importance level. Since the importance of the keyword based on the importance level and the importance under the target importance level can accurately reflect the importance of the keyword, the accuracy of the determined first importance is improved.

[0295] In one embodiment of the present disclosure, the possibility prediction submodule 1206 includes:

[0296] A probability prediction unit, used for predicting the probability that the importance level of each character in the target text is each preset importance level based on the local features of the target text;

[0297] A first possibility degree determination unit is used to determine the possibility degree corresponding to each character in the keyword from the possibility degree corresponding to each character in the target text;

[0298] The second possibility determination unit is used to perform statistical analysis on the possibility that the importance level of each character in the keyword is the preset importance level for each preset importance level, and determine the statistical analysis result as the possibility that the importance level of the keyword is each preset importance level.

[0299] From the above, it can be seen that since the keyword contains multiple characters, the possibility corresponding to the keyword is related to the possibility corresponding to each character. Therefore, for each preset importance level, a statistical analysis is performed on the possibility that the importance level of each character in the keyword is the preset importance level. The result obtained by the statistical analysis can accurately reflect the possibility that the importance level of the keyword is the preset importance level, so as to obtain a more accurate possibility that the importance level of the keyword is each preset importance level.

[0300] In one embodiment of the present disclosure, the keyword extraction module is specifically used to perform feature fusion on the global features and local features of the target text to obtain a first fused feature; and extract keywords of the target text based on the first fused feature.

[0301] Since the first fusion feature is a feature obtained by fusing the global feature and the local feature of the target text, the first fusion feature can more completely include the global feature and the local feature of the target text. Therefore, when extracting keywords based on the first fusion feature, the global feature and the local feature of the target text can be better referenced.

[0302] In one embodiment of the present disclosure, the above device further includes:

[0303] The word type determination module is used to determine the word type of the keyword of the target text from various preset word types based on the global features and local features of the target text.

[0304] Since the global and local features of the target text are considered when determining the word type of the keyword, the global and local features of the target text can fully reflect the word type information of each word in the target text. Therefore, based on the global and local features, a more accurate word type of the keyword can be obtained.

[0305] In one embodiment of the present disclosure, the text feature extraction module is specifically used to encode each character in the target text using a multi-layer encoding method based on the context information of each character in the target text to obtain a preset number of layer feature vectors as the text feature of the target text;

[0306] The global feature extraction module is specifically used to extract the feature vector of the last layer among the feature vectors of the preset number of layers, and determine it as the global feature of the target text;

[0307] The local feature extraction module is specifically used to extract feature vectors from the preset number of layers of feature vectors except the feature vector of the last layer, perform feature fusion on the extracted feature vectors to obtain a second fused feature; perform local feature extraction on the second fused feature to obtain the local feature of the target text.

[0308] From the above, it can be seen that since the text features of the target text are obtained based on the context information of each character in the target text, the context information of each character can accurately reflect the semantic information of each character. Since the semantic information of the target text is related to the semantic information of each character, the text features of the target text obtained can accurately reflect the semantic information of the target text.

[0309] Moreover, in the multi-layer feature vectors, the feature vectors of the last layer can accurately represent the global features of the text, and the feature vectors other than the feature vectors of the last layer can accurately represent the local features of the text. Therefore, by determining the feature vectors of the last layer as the global features of the target text, more accurate global features can be obtained, and by determining the local features of the target text based on the feature vectors other than the feature vectors of the last layer, more accurate local features can be obtained.

[0310] In one embodiment of the present disclosure, the text feature extraction module is specifically used to input the target text into the text feature extraction layer of the pre-trained keyword extraction model to obtain the text features of the target text;

[0311] The global feature extraction module is specifically used to input the text features of the target text into the global feature extraction layer in the keyword extraction model to obtain the global features of the target text;

[0312] The local feature extraction module is specifically used to input the text features of the target text into the local feature extraction layer in the keyword extraction model to obtain the local features of the target text;

[0313] The keyword extraction module is specifically used to input the global features and local features of the target text into the keyword extraction layer in the keyword extraction model to obtain the keywords of the target text.

[0314] From the above, it can be seen that since the pre-trained keyword extraction model is trained based on a large amount of sample text, the keyword extraction model learns the ability to identify keywords based on sample text. Therefore, when extracting keywords from the target text based on each network layer in the above keyword extraction model, the accuracy of the extracted keywords can be improved.

[0315] In one embodiment of the present disclosure, the above device further includes:

[0316] The importance determination module is used to input the local features of the target text into the importance determination layer in the keyword extraction model to obtain the first importance of the keyword.

[0317] From the above, it can be seen that since the keyword extraction model also includes an importance determination layer, which is used to predict the keyword importance of the text based on the local features of the text, after the local features of the target text are input into the keyword extraction model, a more accurate first importance of the keyword can be obtained.

[0318] Corresponding to the above-mentioned model training method, the embodiment of the present disclosure also provides a model training device. Fig.13 , Fig.13 This is a schematic diagram of the structure of the first model training device provided in an embodiment of the present disclosure, and the above-mentioned device includes the following modules 1301-1303.

[0319] An information acquisition module 1301 is used to obtain sample text and real keywords of the sample text;

[0320] The keyword determination module 1302 is used to input the sample text into a preset neural network model to obtain predicted keywords of the sample text, wherein the predicted keywords are keywords predicted based on global features and local features of the sample text, the global features of the sample text are features extracted from global features of the text features of the sample text, and the local features of the sample text are features extracted from local features of the text features of the sample text;

[0321] The model parameter adjustment module 1303 is used to adjust the model parameters of the neural network model based on the difference between the predicted keyword and the real keyword.

[0322] From the above, it can be seen that since the model parameters of the neural network model are adjusted based on the difference between the predicted keywords and the real keywords, the difference between the predicted keywords and the real keywords reflects the ability of the neural network model to extract keywords. Therefore, based on the difference between the predicted keywords and the real keywords, the ability of the neural network model to extract keywords can be effectively improved.

[0323] See also Fig.14 , Fig.14 This is a structural diagram of a second model training device provided in an embodiment of the present disclosure, and the device includes the following modules 1401-1404.

[0324] Information acquisition module 1401, used to obtain sample text and real keywords of the sample text;

[0325] An importance obtaining module 1402 is used to obtain the real importance of the real keywords of the sample text after obtaining the sample text and the real keywords of the sample text in the information obtaining module;

[0326] The keyword determination module 1403 is specifically used to input the sample text into a preset neural network model to obtain the predicted keywords of the sample and the predicted importance of the predicted keywords, wherein the predicted keywords are keywords predicted based on the global features and local features of the sample text, the global features of the sample text are features extracted from the global features of the text features of the sample text, the local features of the sample text are features extracted from the local features of the text features of the sample text, and the predicted importance is importance predicted based on the importance of the local features of the sample text;

[0327] The model parameter adjustment module 1404 is specifically used to adjust the model parameters of the neural network model based on the difference between the predicted keyword and the real keyword, and the difference between the predicted importance and the real importance of the same keyword.

[0328] Since the model parameters of the neural network model are adjusted based on the difference between predicted keywords and real keywords, and the difference between the predicted importance and the real importance of the same keyword, the difference between the predicted keywords and the real keywords reflects the ability of the neural network model to extract keywords, and the difference between the predicted importance and the real importance of the same keywords reflects the ability of the neural network model to predict the importance of keywords. Therefore, based on the above two differences, the ability of the neural network model to extract keywords and predict the importance of keywords can be effectively improved.

[0329] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0330] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0331] An embodiment of the present disclosure provides an electronic device, including:

[0332] at least one processor; and

[0333] a memory communicatively connected to the at least one processor; wherein,

[0334] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned keyword extraction method or model training method.

[0335] An embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the above-mentioned keyword extraction method or model training method.

[0336] The embodiments of the present disclosure provide a computer program product, including a computer program, which implements the above-mentioned keyword extraction method or model training method when executed by a processor.

[0337] Fig.15 A schematic block diagram of an example electronic device 1500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0338] like Fig.15 As shown, the device 1500 includes a computing unit 1501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1502 or a computer program loaded from a storage unit 1508 into a random access memory (RAM) 1503. In the RAM 1503, various programs and data required for the operation of the device 1500 can also be stored. The computing unit 1501, the ROM 1502, and the RAM 1503 are connected to each other via a bus 1504. An input / output (I / O) interface 1505 is also connected to the bus 1504.

[0339] A number of components in the device 1500 are connected to the I / O interface 1505, including: an input unit 1506, such as a keyboard, a mouse, etc.; an output unit 1507, such as various types of displays, speakers, etc.; a storage unit 1508, such as a disk, an optical disk, etc.; and a communication unit 1509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1509 allows the device 1500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0340] The computing unit 1501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1501 performs the various methods and processes described above, such as a keyword extraction method or a model training method. For example, in some embodiments, the keyword extraction method or the model training method may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 1508. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1500 via the ROM 1502 and / or the communication unit 1509. When the computer program is loaded into the RAM 1503 and executed by the computing unit 1501, one or more steps of the keyword extraction method or the model training method described above may be performed. Alternatively, in other embodiments, the computing unit 1501 may be configured to execute the keyword extraction method or the model training method in any other appropriate manner (eg, by means of firmware).

[0341] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0342] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0343] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0344] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0345] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0346] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0347] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0348] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A keyword extraction method, comprising: Obtaining a target text to be processed, including: obtaining an original text to be processed, converting the original text to obtain a converted text including a position identifier of a position of each character in the original text in a preset dictionary library, and processing the converted text into a target text of a preset length; Extracting text features of the target text; Performing global feature extraction on the text features of the target text to obtain the global features of the target text, wherein the global features of the target text are used to characterize the semantics expressed by the target text as a whole; Performing local feature extraction on the text features of the target text to obtain the local features of the target text, including: encoding each character based on the context information of each character in the target text to obtain a preset number of layers of feature vectors; fusing the feature vectors of the preset number of layers of feature vectors except the last layer of feature vectors to obtain hierarchical semantic features; performing local feature extraction on the hierarchical semantic features to obtain the local features of the target text; wherein the local features of the target text are used to characterize the semantics expressed by the context of each character in the target text; Based on the global features and local features of the target text, keywords of the target text are extracted.

2. The method according to claim 1, further comprising: Based on the local features of the target text, the first importance of the keyword is predicted.

3. The method according to claim 2, wherein: The predicting the first importance of the keyword based on the local features of the target text includes: Based on the local features of the target text, predicting the probability that the importance level of the keyword is each preset importance level; Determining a target importance level of the keyword from among various preset importance levels based on the probability corresponding to the keyword; Calculating the second importance of the keyword at the target importance level based on the possibility corresponding to the keyword; Based on the target importance level and the second importance of the keyword, a first importance of the keyword is determined.

4. The method according to claim 3, wherein: The predicting, based on the local features of the target text, that the importance level of the keyword is a probability of each preset importance level includes: Based on the local features of the target text, predicting the probability that the importance level of each character in the target text is each preset importance level; Determining the probability degree corresponding to each character in the keyword from the probability degree corresponding to each character in the target text; For each preset importance level, a statistical analysis is performed on the possibility that the importance level of each character in the keyword is the preset importance level, and the statistical analysis result is determined as the possibility that the importance level of the keyword is each preset importance level.

5. The method according to any one of claims 1 to 4, wherein: The extracting keywords of the target text based on the global features and local features of the target text includes: Performing feature fusion on the global features and local features of the target text to obtain a first fused feature; Based on the first fusion feature, keywords of the target text are extracted.

6. The method according to any one of claims 1 to 4, further comprising: Based on the global features and local features of the target text, the word type of the keyword of the target text is determined from various preset word types.

7. The method according to any one of claims 1 to 4, wherein: The extracting text features of the target text includes: Based on the context information of each character in the target text, each character is encoded using a multi-layer encoding method to obtain a preset number of layer feature vectors as text features of the target text; The step of extracting global features from the text features of the target text to obtain the global features of the target text includes: Extracting the feature vector of the last layer among the feature vectors of the preset number of layers and determining it as the global feature of the target text; The extracting local features of the text features of the target text to obtain the local features of the target text includes: Extracting feature vectors except the feature vector of the last layer from the feature vectors of the preset number of layers, and performing feature fusion on the extracted feature vectors to obtain a second fused feature; Perform local feature extraction on the second fused features to obtain local features of the target text.

8. The method according to any one of claims 1 to 4, wherein: The extracting text features of the target text includes: Inputting the target text into a text feature extraction layer in a pre-trained keyword extraction model to obtain text features of the target text; The step of extracting global features from the text features of the target text to obtain the global features of the target text includes: Inputting the text features of the target text into the global feature extraction layer in the keyword extraction model to obtain the global features of the target text; The extracting local features of the text features of the target text to obtain the local features of the target text includes: Inputting the text features of the target text into the local feature extraction layer in the keyword extraction model to obtain the local features of the target text; The extracting keywords of the target text based on the global features and local features of the target text includes: The global features and local features of the target text are input into the keyword extraction layer in the keyword extraction model to obtain the keywords of the target text.

9. The method according to claim 8, further comprising: The local features of the target text are input into the importance determination layer in the keyword extraction model to obtain the first importance of the keyword.

10. A model training method, comprising: Obtain sample text and real keywords of the sample text; The obtaining of the sample text comprises: obtaining a sample original text, converting the sample original text to obtain a sample converted text including a position identifier of a position of each character in the sample original text in a preset dictionary library, and processing the sample converted text into a sample text of a preset length; The sample text is input into a preset neural network model to obtain predicted keywords of the sample text, wherein the predicted keywords are keywords predicted based on global features and local features of the sample text, the global features of the sample text are features extracted from global features of the text features of the sample text, the local features of the sample text are features extracted from local features of the text features of the sample text, and the local features of the sample text are encoding each character based on context information of each character in the sample text to obtain a preset number of layers of feature vectors; the feature vectors of the preset number of layers of feature vectors except the last layer of feature vectors are merged to obtain hierarchical semantic features; the hierarchical semantic features are extracted from local features; Based on the difference between the predicted keyword and the real keyword, the model parameters of the neural network model are adjusted.

11. The method according to claim 10, After obtaining the sample text and the real keywords of the sample text, the method further includes: Obtaining the real importance of the real keywords of the sample text; The step of inputting the sample text into a preset neural network model to obtain predicted keywords of the sample text includes: Inputting the sample text into a preset neural network model to obtain predicted keywords of the sample and predicted importance of the predicted keywords, wherein the predicted importance is: the importance obtained by performing importance prediction based on local features of the sample text; The adjusting the model parameters of the neural network model based on the difference between the predicted keyword and the real keyword includes: Based on the difference between the predicted keyword and the real keyword, and the difference between the predicted importance and the real importance of the same keyword, the model parameters of the neural network model are adjusted.

12. A keyword extraction device, comprising: A text acquisition module, used for acquiring a target text to be processed; A feature extraction module, used to extract text features of the target text; A global feature extraction module, used for performing global feature extraction on the text features of the target text to obtain the global features of the target text, wherein the global features of the target text are used to characterize the semantics expressed by the target text as a whole; A local feature extraction module, used for extracting local features of the text features of the target text to obtain local features of the target text, wherein the local features of the target text are used to characterize the semantics expressed by the context in which each character in the target text is located; A keyword extraction module, used to extract keywords of the target text based on the global features and local features of the target text; The text acquisition module is specifically used to obtain the original text to be processed, convert the original text to obtain a converted text containing the position identifier of each character in the original text in a preset dictionary library, and process the converted text into a target text of a preset length; The local feature extraction module is specifically used to encode each character in the target text based on the context information of each character to obtain a preset number of layers of feature vectors; fuse the feature vectors of the preset number of layers of feature vectors except the last layer of feature vectors to obtain hierarchical semantic features; and perform local feature extraction on the hierarchical semantic features to obtain local features of the target text.

13. The apparatus according to claim 12, further comprising: The importance prediction module is used to predict the first importance of the keyword based on the local features of the target text.

14. The device according to claim 13, wherein: The importance prediction module comprises: A possibility prediction submodule, used for predicting the possibility that the importance level of the keyword is each preset importance level based on the local features of the target text; An importance level determination submodule, used to determine a target importance level of the keyword from various preset importance levels based on the probability corresponding to the keyword; An importance calculation submodule, used for calculating the second importance of the keyword at the target importance level based on the possibility corresponding to the keyword; The importance determination submodule is used to determine the first importance of the keyword based on the target importance level and the second importance of the keyword.

15. The device according to claim 14, wherein: The possibility prediction submodule includes: A probability prediction unit, used for predicting the probability that the importance level of each character in the target text is each preset importance level based on the local features of the target text; A first possibility degree determination unit is used to determine the possibility degree corresponding to each character in the keyword from the possibility degree corresponding to each character in the target text; The second possibility determination unit is used to perform statistical analysis on the possibility that the importance level of each character in the keyword is the preset importance level for each preset importance level, and determine the statistical analysis result as the possibility that the importance level of the keyword is each preset importance level.

16. The device according to any one of claims 12 to 15, wherein: The feature extraction module is specifically used to input the target text into the text feature extraction layer in the pre-trained keyword extraction model to obtain the text features of the target text; The global feature extraction module is specifically used to input the text features of the target text into the global feature extraction layer in the keyword extraction model to obtain the global features of the target text; The local feature extraction module is specifically used to input the text features of the target text into the local feature extraction layer in the keyword extraction model to obtain the local features of the target text; The keyword extraction module is specifically used to input the global features and local features of the target text into the keyword extraction layer in the keyword extraction model to obtain the keywords of the target text.

17. A model training device, comprising: An information acquisition module is used to obtain sample text and real keywords of the sample text; A keyword determination module is used to input the sample text into a preset neural network model to obtain predicted keywords of the sample text, wherein the predicted keywords are keywords predicted based on global features and local features of the sample text, the global features of the sample text are features extracted from global features of the text features of the sample text, the local features of the sample text are features extracted from local features of the text features of the sample text, and the local features of the sample text are: based on the context information of each character in the sample text, each character is encoded to obtain a preset number of layers of feature vectors; feature vectors other than the last layer of feature vectors in the preset number of layers of feature vectors are merged to obtain hierarchical semantic features; and the hierarchical semantic features are extracted from local features; A model parameter adjustment module, used to adjust the model parameters of the neural network model based on the difference between the predicted keyword and the real keyword; Among them, the information acquisition module is specifically used to obtain a sample original text, convert the sample original text, obtain a sample converted text containing a position identifier of the position of each character in the sample original text in a preset dictionary library, and process the sample converted text into a sample text of a preset length.

18. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9 or 10-11.

19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9 or 10-11.

20. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1-9 or 10-11.