Keyword extraction method and device
By extracting and combining the structural and semantic features of text and word, the problem of poor keyword extraction in short text is solved, and more efficient and accurate keyword extraction is achieved.
Patent Information
- Application Number
- CN202210760249.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-06-30
AI Technical Summary
The prior art has poor results in extracting keywords in short texts, mainly because the word frequency in short texts is generally once, and traditional algorithms are difficult to effectively handle.
By obtaining the target text, extracting its text structure and semantic features, as well as the word structure and word semantic features of each word, combining these features to determine the text and word characteristics, thereby accurately determining the keywords.
It improves the efficiency and accuracy of keyword extraction, avoids the problem of uneven distribution of high and low frequency words in the semantic space, and is especially suitable for keyword extraction in short texts.
Smart Images

Figure CN114943236B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and in particular to a keyword extraction method. The present application also relates to a keyword extraction device, a computing device, and a computer-readable storage medium. Background Art
[0002] With the development of artificial intelligence in the field of computer technology, the field of natural language processing has also developed rapidly. Information retrieval based on text is an important branch of natural language processing. Artificial intelligence (AI) refers to the ability of an engineered (i.e. designed and manufactured) system to perceive the environment, as well as the ability to acquire, process, apply and represent knowledge. The development status of key technologies in the field of artificial intelligence includes key technologies such as machine learning, knowledge graphs, natural language processing, computer vision, human-computer interaction, biometric recognition, virtual reality / augmented reality, etc. Natural language processing (NLP) is an important research direction in the field of computer science. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. The specific manifestations of natural language processing include machine translation, text summarization, text classification, text proofreading, information extraction, speech synthesis, speech recognition, etc. With the development of natural language processing technology and the accelerated pace of life, the effective information that needs to be delivered to users is becoming shorter and shorter. At this time, keyword extraction technology in natural language processing can be used to extract keywords from text to shorten the effective information.
[0003] Traditional general keyword extraction algorithms are mainly used to extract keywords from medium and long texts, such as the word frequency counting algorithm (TF-IDF) and the graph center point algorithm (TextRank). However, the above algorithms are not very effective in extracting keywords from short texts, mainly due to the particularity of short texts: the frequency of words in short texts is generally once, while the frequency of words in medium and long texts is more. Therefore, an effective solution is urgently needed to solve the above problems. Summary of the invention
[0004] In view of this, the embodiment of the present application provides a keyword extraction method to solve the technical defects existing in the prior art. The embodiment of the present application also provides a keyword extraction device, a computing device, and a computer-readable storage medium.
[0005] According to a first aspect of an embodiment of the present application, a keyword extraction method is provided, comprising:
[0006] Get the target text;
[0007] Extracting text structure features and text semantic features of the target text, as well as word structure features and word semantic features of each word in the target text;
[0008] Determine text features according to the text structure features and text semantic features, and determine word features according to the word structure features and word semantic features;
[0009] Keywords are determined from the words according to the text features and the word features of the words.
[0010] According to a second aspect of an embodiment of the present application, a keyword extraction device is provided, including:
[0011] A first acquisition module is configured to acquire a target text;
[0012] An extraction module, configured to extract text structure features and text semantic features of the target text, as well as word structure features and word semantic features of each word in the target text;
[0013] A first determination module is configured to determine text features according to the text structure features and text semantic features, and to determine word features according to the word structure features and word semantic features;
[0014] The second determination module is configured to determine keywords from the words according to the text features and the word features of the words.
[0015] According to a third aspect of an embodiment of the present application, a computing device is provided, including:
[0016] Memory and processor;
[0017] The memory is used to store computer-executable instructions, and the processor implements the steps of the keyword extraction method when executing the computer-executable instructions.
[0018] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the keyword extraction method are implemented.
[0019] According to a fifth aspect of the embodiments of the present application, a chip is provided, which stores a computer program, and the computer program implements the steps of the keyword extraction method when executed by the chip.
[0020] The keyword extraction method provided by the present application obtains a target text, extracts the text structure features and text semantic features of the target text, and the word structure features and word semantic features of each word in the target text; determines text features based on the text structure features and text semantic features, determines word features based on the word structure features and word semantic features; determines keywords from the words based on the text features and word features of each word. Determining text features through text structure features and text semantic features, and determining word features through word structure features and word semantic features can more accurately determine the semantic and structural level information of a text or word, that is, make the text features and word features more accurate, and then determine keywords based on text features and word features, thereby improving the efficiency of determining keywords. It avoids the problem of uneven distribution of high- and low-frequency words in the semantic space. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a structural schematic diagram of a keyword extraction method provided by an embodiment of the present application;
[0022] Figure 2 is a flow chart of a keyword extraction method provided by an embodiment of the present application;
[0023] Figure 3 It is a structural schematic diagram of a feature extraction model in a keyword extraction method provided in an embodiment of the present application;
[0024] Figure 4 It is a structural schematic diagram of another feature extraction model in a keyword extraction method provided in an embodiment of the present application;
[0025] Figure 5 It is a structural schematic diagram of another feature extraction model in a keyword extraction method provided in an embodiment of the present application;
[0026] Figure 6 is a processing flow chart of a keyword extraction method applied to short texts provided in an embodiment of the present application;
[0027] Figure 7 is a structural schematic diagram of a keyword extraction device provided in one embodiment of the present application;
[0028] Figure 8 It is a structural block diagram of a computing device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0029] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.
[0030] The terms used in one or more embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present application. The singular forms of "a", "said" and "the" used in one or more embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more associated listed items.
[0031] It should be understood that, although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first.
[0032] First, the terms involved in one or more embodiments of this specification are explained.
[0033] BERT (Bidirectional Encoder Representation from Transformers) pre-trained language model: learns deep bidirectional representations from unlabeled data through pre-training. After pre-training, it adds an additional output layer for fine-tuning, and finally achieves SOTA (state of the art) on multiple NLP (Natural Language Processing) tasks.
[0034] Contrastive Learning: Contrastive learning belongs to self-supervised learning, which is a type of unsupervised learning. It does not require manually labeled category label information, but directly uses the data itself as supervisory information to learn the feature expression of sample data and use it for downstream tasks.
[0035] Keyword extraction algorithm: In the field of natural language processing, keyword extraction algorithm is needed to extract relational information from long or short texts. Keyword extraction algorithm is widely used in recommendation systems and search engines, and is also an important component of text mining.
[0036] SimCSE (Simple Contrastive Learning of Sentence Embedding) uses a self-supervised approach to improve the sentence representation capabilities of the model. There are two main construction methods. For unsupervised, a Dropout layer is used to construct positive examples. A sample is passed through the encoder twice to obtain a positive example pair, and the negative examples are other sentences in the same batch. For supervised, the natural structure of the attention mechanism natural language inference (SNLI, The Stanford Natural Language Inference) dataset is used, in which the opposing category is the negative sample, and the other two categories are the positive samples.
[0037] Then, the keyword extraction method provided in this specification is briefly described.
[0038] With the development of artificial intelligence in the field of computer technology, the field of natural language processing has also developed rapidly. Information retrieval based on text is an important branch of natural language processing. Artificial intelligence (AI) refers to the ability of an engineered (i.e. designed and manufactured) system to perceive the environment, as well as the ability to acquire, process, apply and represent knowledge. The development status of key technologies in the field of artificial intelligence includes key technologies such as machine learning, knowledge graphs, natural language processing, computer vision, human-computer interaction, biometric recognition, virtual reality / augmented reality, etc. Natural language processing (NLP) is an important research direction in the field of computer science. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. The specific manifestations of natural language processing include machine translation, text summarization, text classification, text proofreading, information extraction, speech synthesis, speech recognition, etc. With the development of natural language processing technology and the accelerated pace of life, the effective information that needs to be delivered to users is becoming shorter and shorter. At this time, keyword extraction technology in natural language processing can be used to extract keywords from text to shorten the effective information.
[0039] There are three major categories of existing keyword extraction technologies: graph algorithms, word frequency statistics, and similarity calculations, as follows.
[0040] TextRank keyword extraction algorithm (graph algorithm): words are nodes in the graph, and the edges between words are determined by the "co-occurrence" relationship; "co-occurrence" refers to co-appearance within a sliding window of a given size. Construct an unweighted undirected graph and use the PageRank algorithm to obtain the corresponding high-weight words.
[0041] TF-IDF (term frequency-inverse document frequency) keyword extraction algorithm (word frequency statistics algorithm), the algorithm is divided into two parts: the first part calculates the word frequency, such as TF = (the number of times a word appears in a document) ÷ (the total number of words in the article); the second part calculates the inverse document frequency, IDF = log ((the total number of documents in the corpus) ÷ (the document containing the word) + 1); finally, the corresponding high-weighted words are calculated as the keywords of the document.
[0042] keyBERT extracts the document-level representation of the document through BERT, and then finds the first few subphrases that are most similar to the document by using cosine similarity.
[0043] Since TF-IDF is calculated relatively quickly, most search engines on websites use the TF-IDF keyword extraction algorithm.
[0044] Traditional general keyword extraction algorithms are mainly used to extract keywords from medium and long texts, such as the word frequency counting algorithm (TF-IDF) and the graph center point algorithm (TextRank). However, the above algorithms are not very effective in extracting keywords from short texts, mainly due to the particularity of short texts: the frequency of words in short texts is generally once, while the frequency of words in medium and long texts is more.
[0045] In addition, KeyBERT relies on BERT to provide document-level representations of documents. For example, the word vectors in BERT are not evenly distributed in space. High frequencies are close to the origin, while low frequencies are far away from the origin. The sparse distribution of low frequencies leads to insufficient training from low-frequency words, and ultimately BERT will have a certain degree of semantic non-smoothness in the sentence vector space. In addition, BERT's pre-trained language model for Chinese has the ERNIE (Enhanced Representation through Knowledge Integration) model, and the problem is obvious. For data sets in other specific fields, the characteristics of ERNIE's lack of generalization and migration capabilities will be magnified.
[0046] Therefore, the present application provides a keyword extraction method, which obtains a target text; extracts the text structure features and text semantic features of the target text, as well as the word structure features and word semantic features of each word in the target text; determines text features based on the text structure features and text semantic features, and determines word features based on the word structure features and word semantic features; determines keywords from the words based on the text features and word features of each word. Determining text features through text structure features and text semantic features, and determining word features through word structure features and word semantic features can more accurately determine the semantic and structural level information of a text or word, that is, make the text features and word features more accurate, and then determine keywords based on text features and word features, thereby improving the efficiency of determining keywords. It avoids the problem of uneven distribution of high- and low-frequency words in the semantic space.
[0047] In the present application, a keyword extraction method is provided. The present application also relates to a keyword extraction device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.
[0048] The execution subject of the keyword extraction method provided in the embodiment of the present application can be a server or a terminal, which is not limited in the embodiment of the present application. In addition, the terminal can be any electronic product that can interact with the user, such as a PC (Personal Computer), a mobile phone, a PPC (Pocket PC), a tablet computer, etc. The server can be a single server, a server cluster composed of multiple servers, or a cloud computing service center, which is not limited in the embodiment of the present application.
[0049] Figure 1 This is a structural diagram of a keyword extraction method provided by an embodiment of the present application. First, the target text is obtained; then the text structure features and text semantic features of the target text are extracted, and the word structure features and word semantic features of each word in the target text are extracted, such as extracting the word structure features and word semantic features of words 1 to X; then, the text features are determined according to the text structure features and text semantic features, and the word features of each word are determined according to the word structure features and word semantic features of each word. Finally, the keywords of the target text are determined from each word according to the text features and the features of each word.
[0050] The present application provides a keyword extraction method that determines text features through text structure features and text semantic features, and determines word features through word structure features and word semantic features, which can more accurately determine the semantic and structural level information of text or words, that is, the text features and word features are more accurate, and then keywords can be determined based on text features and word features, thereby improving the efficiency of determining keywords. The problem of uneven distribution of high-frequency and low-frequency words in the semantic space is avoided.
[0051] Figure 2 A flowchart of a keyword extraction method provided according to an embodiment of the present application is shown, which specifically includes the following steps:
[0052] Step 202: Obtain target text.
[0053] The core of the embodiments of the present application is to extract keywords. For texts in different fields or categories, such as texts in the medical field, texts in the astronomy field, long texts, and short texts, the process of extracting keywords is basically the same. The keyword extraction process is introduced in detail below.
[0054] Specifically, a text refers to a written language, usually a sentence or a combination of sentences with complete and systematic meanings. A text can be a sentence, a paragraph or a chapter, all of which belong to texts. A target text refers to a text from which keywords are to be extracted. Preferably, the target text is a short text.
[0055] In actual applications, there are many ways to obtain the target text. For example, the operator may send a keyword extraction instruction to the execution subject, or send an instruction to obtain the target text. Accordingly, after receiving the instruction, the execution subject starts to obtain the target text; or the server may automatically obtain the target text at preset time intervals. For example, after a preset time, a server with a keyword extraction function automatically obtains the target text; or after a preset time, a terminal with a keyword extraction function automatically obtains the target text. This specification does not limit the method of obtaining the target text.
[0056] Step 204: extracting text structure features and text semantic features of the target text, as well as word structure features and word semantic features of each word in the target text.
[0057] On the basis of acquiring the target text, further, text structure features and text semantic features of the target text, as well as word structure features and word semantic features of each word in the target text are extracted.
[0058] Specifically, structural features refer to the arrangement or position features of multiple text units, that is, the surface information features of multiple text units; semantic features refer to the features corresponding to the language meanings of multiple text units; text structural features refer to the structural features corresponding to the text; text semantic features refer to the semantic features corresponding to the text; word structural features refer to the structural features corresponding to words; word semantic features refer to the semantic features corresponding to words.
[0059] In a possible implementation of the embodiments of this specification, after acquiring the target text, the text structure features of the target text can be first extracted by a preset structural feature extraction tool, and then the word structure features of each word in the target text can be respectively extracted; the text semantic features of the target text can be first extracted by a preset semantic feature extraction tool, and then the word semantic features of each word in the target text can be respectively extracted. In this way, the accuracy of determining the extracted text structure features, text semantic features, word structure features, and word semantic features can be improved.
[0060] In another possible implementation of the embodiments of the present specification, after acquiring the target text, the text structure features and text semantic structure features of the target text can be first extracted by a preset structure semantic feature extraction tool, and then the word structure features and word semantic structure features of each word in the target text can be extracted. In this way, the speed of extracting text structure features, text semantic features, word structure features, and word semantic features can be improved.
[0061] It should be noted that in order to facilitate the processing of each word in the target text, before extracting the text structure features and text semantic features of the target text, as well as the word structure features and word semantic features of each word in the target text, it is also necessary to determine each word in the target text, or obtain each word in the target text.
[0062] In a possible implementation of the embodiments of this specification, before extracting the text structure features and text semantic features of the target text, and the word structure features and word semantic features of each word in the target text, the target text may be retrieved again to obtain each word in the pre-stored target text. In this way, each word in the target text may be quickly obtained.
[0063] In another possible implementation of the embodiment of the present specification, before extracting the text structure features and text semantic features of the target text, as well as the word structure features and word semantic features of each word in the target text, the target text may be segmented to obtain multiple words. In this way, the target text is analyzed to obtain each word in the target text, which can ensure the accuracy and comprehensiveness of each word obtained, thereby improving the accuracy and efficiency of keyword extraction.
[0064] Since the target text contains many function words, which are not necessarily keywords of the target text, in order to reduce the amount of data processing, these function words can be removed after the target text is segmented. That is, the target text is segmented to obtain multiple words. The specific implementation process can be as follows:
[0065] Perform word segmentation on the target text to obtain multiple candidate words, and determine the part of speech of each candidate word;
[0066] According to the part of speech, multiple words are selected from multiple candidate words.
[0067] Specifically, the word segmentation process of matching the character string in the sample problem can be a forward maximum matching method, a reverse maximum matching method, a shortest path word segmentation method or a bidirectional maximum matching method, and this application does not limit this; candidate words refer to all words obtained after word segmentation processing of the target text; part of speech refers to the basis for dividing word classes based on the characteristics of the word, and the part of speech is determined by its meaning, morphology and grammatical function in the language to which it belongs.
[0068] In practical applications, all candidate words are obtained by segmenting the target text, and then the part-of-speech tagging is performed on each candidate word to determine the part-of-speech of each candidate word. Then, according to the part-of-speech of each candidate word, the candidate words that are function words among the multiple candidate words are deleted to obtain multiple words.
[0069] Exemplarily, the target text is "white chalk is dancing on the blackboard, and white powder falls from time to time". After the target text is segmented, candidate words are obtained: "white", "of", "chalk", "in", "blackboard", "on", "dancing", "from time to time", "there is", "white powder" and "falling". Further, these candidate words are tagged with part of speech to obtain "white" (adjective), "of" (of), "chalk" (noun), "in" (preposition), "blackboard" (noun), "on" (directional word), "dancing" (state word), "wear" (wear), "from time to time" (adverb), "there is" (adverb), "white powder" (noun), and "falling" (verb). Among them, "of" (of), "in" (preposition), "on" (directional word), "wear" (wear), "from time to time" (adverb), and "there is" (adverb) are all function words. After deleting them, multiple words are obtained, namely "white", "chalk", "blackboard", "dancing", "white powder", and "falling".
[0070] In addition, the stop words in multiple candidate words or words can also be filtered, wherein stop words refer to that in information retrieval, in order to save storage space and improve search efficiency, some words or words are automatically filtered out before or after processing natural language text, and these words or words are called stop words, such as "I", "you" etc. These stop words are generally generated by manual input and non-automatic, and the stop words after generation can form a stop word list, such as a pause word library. It should be noted that filtering stop words can be performed before screening multiple words, or after screening multiple words. In this way, the data processing amount can be further reduced.
[0071] Step 206: Determine text features based on text structure features and text semantic features, and determine word features based on word structure features and word semantic features.
[0072] Based on the extracted text structure features and text semantic features of the target text, as well as the word structure features and word semantic features of each word in the target text, text features are further determined based on the text structure features and text semantic features, and word features are determined based on the word structure features and word semantic features.
[0073] Specifically, text features refer to the comprehensive features of the text, that is, the combination of features determined from multiple aspects, such as text structure features and text semantic features; word features refer to the comprehensive features of words, that is, the combination of features determined from multiple aspects, such as word structure features and word semantic features.
[0074] In a possible implementation of the embodiments of this specification, after extracting the text structure features and text semantic features of the target text, as well as the word structure features and word semantic features of each word in the target text, the text structure features and text semantic features of the target text can be spliced to obtain the text features of the target text; for any word, the word structure features and word semantic features of the word are spliced to obtain the word features of the word. In this way, the speed of determining text features and word features can be improved.
[0075] In another possible implementation of the embodiments of this specification, after extracting the text structure features and text semantic features of the target text, as well as the word structure features and word semantic features of each word in the target text, the text structure features and text semantic features of the target text can be input into a preset feature calculation formula for calculation to obtain the text features of the target text; for any word, the word structure features and word semantic features of the word are also input into a preset feature calculation formula for calculation to obtain the word features of the word. In this way, the feature calculation formula is set according to the needs so that the text features and word features can better reflect the characteristics of the text and words, that is, the accuracy of the text features and word features can be improved, thereby improving the accuracy of keyword extraction.
[0076] Exemplarily, the preset feature calculation formula is a weighted average formula, see Formula 1, where a is the weight of the preset structural feature, and b is the weight of the preset semantic feature. The text structural feature of the target text is multiplied by a, and the product of the text semantic feature and b is added to obtain the text feature of the target text; for any word, the word structural feature of the word is multiplied by a, and the product of the word semantic feature and b is added to obtain the word feature of the word.
[0077] Features = a*structural features + b*semantic features (Formula 1)
[0078] Therefore, when a and b are both 0.5, the text feature of the target text is the average of the text structure feature and the text semantic feature, and the word feature of each word is the average of the word structure feature and the word semantic feature.
[0079] Step 208: Determine keywords from each word based on the text features and the word features of each word.
[0080] Specifically, based on determining text features according to text structural features and text semantic features, and determining word features according to word structural features and word semantic features, keywords are further determined according to text features and word features of each word.
[0081] Specifically, keywords are used to express the subject content of a text, not only in scientific papers, but also in texts such as scientific reports and academic papers.
[0082] In a possible implementation of the embodiments of the present specification, the text features of the target text and the word features of each word can be grouped according to a preset classification rule, or the text features of the target text and the word features of each word can be clustered using a preset clustering algorithm, and the words that are used in the same groups or in the same category as the text features can be determined as keywords of the target text. In this way, the efficiency of determining the target keywords can be improved.
[0083] For example, the text feature of the target text is feature 1, the target text has 5 words, the word feature of the first word is feature 2, the word feature of the second word is feature 3, the word feature of the third word is feature 4, the word feature of the fourth word is feature 5, and the word feature of the fifth word is feature 6. Use a clustering algorithm to cluster features 1 to 6 to obtain two classes, where the first class includes features 3 and 5, and the second class includes features 1, 2, 4, and 6. Then, the first word, the third word, and the fifth word corresponding to features 2, 4, and 6, respectively, are determined as keywords of the target text.
[0084] In a possible implementation of the embodiments of this specification, keywords can also be determined from each word based on the similarity between the text feature and the word feature. That is, keywords are determined from each word based on the text feature and the word feature of each word. The specific implementation process can be as follows:
[0085] Determine the first similarity between the word feature and the text feature of each word respectively;
[0086] According to each first similarity, a keyword of the target text is determined from the plurality of words.
[0087] Specifically, similarity is used to describe the similarity between two things; the first similarity refers to the similarity between word features and text features.
[0088] In practical applications, after determining the text features of the target text and the word features of each word, the first similarity between the word features and the text features of each word can be calculated according to a preset similarity algorithm. The preset similarity algorithm can be any one of the Euclidian distance algorithm, the Manhattan distance algorithm, the Minkowski distance algorithm, the cosine similarity algorithm, etc. After determining the first similarity between the word features and the text features of each word, the first similarities can be arranged from large to small, and the words corresponding to the first N similarities in the arrangement are determined as keywords of the target text, where N is a pre-set positive integer; a similarity threshold can also be set, and the words corresponding to the first similarity greater than the similarity threshold are determined as keywords of the target text. In this way, the keywords are determined according to the first similarity between the word features and the text features, that is, the keywords are determined based on the degree of association between the words and the target text, which can improve the accuracy of extracting keywords.
[0089] Exemplarily, the text feature of the target text is feature 1, the target text has 5 words, the word feature of the first word is feature 2, the word feature of the second word is feature 3, the word feature of the third word is feature 4, the word feature of the fourth word is feature 5, and the word feature of the fifth word is feature 6. Using the cosine similarity algorithm, the first similarity s1 between feature 1 and feature 2 is calculated to be 0.2, the first similarity s2 between feature 1 and feature 3 is 0.7, the first similarity s3 between feature 1 and feature 4 is 0.9, the first similarity s4 between feature 1 and feature 5 is 0.1, and the first similarity s5 between feature 1 and feature 6 is 0.5. Assuming that the similarity threshold is 0.4, the first similarity s2, the first similarity s3, and the first similarity s5 are all greater than the similarity threshold, and the second word, the third word, and the fifth word are determined as keywords of the target text.
[0090] In one or more optional embodiments of the present specification, before extracting and determining text structure features and text semantic features, as well as word structure features and word semantic features, a pre-trained feature extraction model may be obtained, and then the target text and each word in the target text are input into the feature extraction model, and the feature extraction model performs feature extraction on the target text and each word in the target text to obtain text features and word features. That is, before extracting the text structure features and text semantic features of the target text, as well as the word structure features and word semantic features of each word in the target text, the following may also be included:
[0091] Obtaining a pre-trained feature extraction model, wherein the feature extraction model includes a structural feature extraction sub-model, a semantic feature extraction sub-model, and an output layer;
[0092] Accordingly, the text structure features and text semantic features of the target text, as well as the word structure features and word semantic features of each word in the target text are extracted. The specific implementation process can be as follows:
[0093] Input the target text and each word in the target text into the structural feature extraction sub-model respectively to obtain the text structural features of the target text and the word structural features of each word;
[0094] Input the target text and each word into the semantic feature extraction sub-model respectively to obtain the text semantic features of the target text and the word semantic features of each word;
[0095] Accordingly, the text features are determined according to the text structure features and the text semantic features, and the word features are determined according to the word structure features and the word semantic features, including:
[0096] For the target text, the text structure features and text semantic features are input into the output layer for processing, and the text features of the target text are output;
[0097] For any word, the word structure features and word semantic features of the word are input into the output layer for processing, and the word features of the word are output.
[0098] Specifically, the feature extraction model refers to a pre-trained neural network model, such as a neural network model, a probabilistic neural network model, such as a BERT model, a Transformer model, etc.; the structural feature extraction sub-model refers to the part of the feature extraction model that extracts structural features of text or words, which can better express the surface information of the text or words; the semantic feature extraction sub-model refers to the part of the feature extraction model that extracts semantic features of text or words, which can better express the semantic level information of the text or words; the output layer refers to the part of the feature extraction model that processes the semantic features and structural features to obtain and output the results.
[0099] In practical applications, after obtaining the target text and each word in the target text, a pre-trained feature extraction model including a structural feature extraction sub-model, a semantic feature extraction sub-model and an output layer is obtained. Then the target text and each word in the target text are input into the structural feature extraction sub-model, and the structural feature extraction sub-model extracts structural features of the target text and each word, and outputs the text structural features of the target text and the word structural features of each word. The target text and each word are input into the semantic feature extraction sub-model, and the semantic feature extraction sub-model extracts semantic features, and outputs the text semantic features of the target text and the word semantic features of each word. Finally, the text semantic features and text structural features of the target text, the word semantic features and word structural features of each word are input into the output layer, and the output layer analyzes and processes the text semantic features and text structural features of the target text, obtains and outputs the text features of the target text, and analyzes and processes the word semantic features and word structural features of each word, and obtains and outputs the word features of each word. By extracting features from the target text and each word through the pre-trained feature extraction model, the acquisition rate and accuracy of text features and word features can be improved.
[0100] See also Figure 3 , Figure 3 A structural schematic diagram of a feature extraction model in a keyword extraction method provided by an embodiment of the present application is shown: the feature extraction model includes a structural feature extraction submodel, a semantic feature extraction submodel and an output layer, wherein the structural feature extraction submodel is used to receive a target text and each word, and extract the structural features of the target text and each word, and obtain text structural features and structural features of each word; the semantic feature extraction submodel is used to receive a target text and each word, and extract the semantic features of the target text and each word, and obtain text semantic features and semantic features of each word; the output layer is used to receive and process text structural features and text semantic features, obtain and output text features, and receive and process structural features and semantic features of each word, and obtain and output features of each word.
[0101] In a possible implementation of the embodiments of this specification, the structural feature extraction submodel may be a first coding layer, and the target text and each word in the target text may be respectively input into the first coding layer to obtain the text structural features of the target text and the word structural features of each word. That is, when the structural feature extraction submodel includes the first coding layer, the target text and each word in the target text may be respectively input into the structural feature extraction submodel to obtain the text structural features of the target text and the word structural features of each word. The specific implementation process may be as follows:
[0102] Input the target text into the first encoding layer for feature extraction to obtain the text structure features of the target text;
[0103] For any word, the word is input into the first coding layer for feature extraction to obtain the word structure features of the candidate word.
[0104] In practical applications, the structural feature extraction submodel includes a first coding layer, the target text can be input into the first coding layer, the first coding layer extracts the structural features of the target text, and outputs the text structural features of the target text; then each word is input into the first coding layer, the first coding layer extracts the structural features of each word, and outputs the word structural features of each word. In this way, as the number of coding layers increases, the structural information at the text level gradually disappears, that is, a single coding layer can well obtain the structural information of the text or word. Therefore, using one coding layer as a structural feature extraction submodel to extract the structural features of the target text and each word can make the obtained text structural features and word structural features more accurate.
[0105] In a possible implementation of the embodiments of this specification, the semantic feature extraction submodel may include multiple second coding layers, and the target text and each word in the target text may be respectively input into the multiple second coding layers connected in series for processing to obtain the text semantic features of the target text and the word semantic features of each word. That is, when the semantic feature extraction submodel includes N second coding layers, where N is a positive integer greater than 2, the target text and each word are respectively input into the semantic feature extraction submodel to obtain the text semantic features of the target text and the word semantic features of each word. The specific implementation process may be as follows:
[0106] For the target text, starting from the first second coding layer, the output of the previous second coding layer is used as the input of the current second coding layer for feature extraction, until the Nth second coding layer, and the text semantic features of the target text are output, wherein the input of the first second coding layer is the target text;
[0107] For any word, starting from the first second coding layer, the output of the previous second coding layer is used as the input of the current second coding layer for feature extraction until the Nth second coding layer, and the semantic features of the word are output, where the input of the first second coding layer is the word.
[0108] In practical applications, the semantic feature extraction submodel includes multiple second coding layers. The target text can be input into the first second coding layer to obtain the first output of the target text, and then the first output of the target text is input into the second second coding layer to obtain the second output of the target text, and so on, until the (N-1)th output of the target text is input into the Nth second coding layer to obtain the text semantic features of the target text. Similarly, for any word, the word is input into the first second coding layer to obtain the first output of the word, and then the first output of the word is input into the second second coding layer to obtain the second output of the word, and so on, until the (N-1)th output of the word is input into the Nth second coding layer to obtain the word semantic features of the word. In this way, as the number of coding layers increases, semantic information can be extracted more accurately. Therefore, using multiple coding layers as semantic feature extraction submodels to extract semantic features of the target text and each word can make the obtained text semantic features and word semantic features more accurate.
[0109] See also Figure 4 , in a keyword extraction method provided by an embodiment of the present application, a structural schematic diagram of another feature extraction model: the feature extraction model includes a structural feature extraction submodel, a semantic feature extraction submodel and an output layer, and the structural feature extraction submodel and the semantic feature extraction submodel are connected in parallel. The structural feature extraction submodel includes a first encoding layer for receiving the target text and each word, and extracting the structural features of the target text and each word, and obtaining the text structural features and the structural features of each word; the semantic feature extraction submodel includes a plurality of second encoding layers connected in series, for receiving the target text and each word, and extracting the semantic features of the target text and each word, and obtaining the text semantic features and the semantic features of each word; the output layer is used to receive and process the text structural features and the text semantic features, obtain and output the text features, and receive and process the structural features and the semantic features of each word, and obtain and output the features of each word. The feature extraction model using the structural feature extraction submodel and the semantic feature extraction submodel can extract the features of the target text and each word simultaneously, and can improve the model processing efficiency.
[0110] In a possible implementation of the embodiments of this specification, the feature extraction model includes multiple third coding layers, wherein the first third coding layer constitutes a structural feature extraction sub-model, and the target text and each word in the target text can be respectively input into the first coding layer to obtain the text structure features of the target text and the word structure features of each word. That is, when the feature extraction model includes M third coding layers, M is a positive integer greater than 2, and the structural feature extraction sub-model includes the first third coding layer, correspondingly, the target text and each word in the target text are respectively input into the structural feature extraction sub-model to obtain the text structure features of the target text and the word structure features of each word, and the specific implementation process can be as follows:
[0111] For the target text, the target text is input into the first third coding layer for feature extraction to obtain the first text coding feature of the target text, and the first text coding feature is determined as the text structure feature of the target text;
[0112] For any word, the word is input into the first third coding layer for feature extraction to obtain the first word coding feature of the word, and the first word coding feature is determined as the word structure feature of the word.
[0113] In practical applications, the feature extraction model includes multiple third coding layers, of which the first third coding layer is a structural feature extraction sub-model. The target text can be input into the first third coding layer, and the first third coding layer extracts the structural features of the target text and outputs the first text coding feature, that is, the text structural feature of the target text; each word is then input into the first third coding layer, and the first third coding layer extracts the structural features of each word and outputs the first word coding feature of each word, that is, the word structural feature of each word. In this way, as the number of coding layers increases, the text-level structural information gradually disappears, that is, a single coding layer can well obtain the structural information of the text or word. Therefore, using the first third coding layer as a structural feature extraction sub-model to extract the structural features of the target text and each word can make the obtained text structural features and word structural features more accurate.
[0114] In a possible implementation of the embodiments of the present specification, the feature extraction model includes multiple third coding layers, wherein the first third coding layer constitutes a structural feature extraction sub-model, and the first third coding layer to the last third coding layer constitute a semantic feature extraction sub-model. At this time, on the basis of inputting the target text into the first third coding layer for feature extraction to obtain the first text coding feature of the target text, and for any word, inputting the word into the first third coding layer for feature extraction to obtain the first word coding feature of the word, the first text coding feature and the first word coding feature can be respectively input into the remaining third coding layers in series for processing to obtain the text semantic features of the target text and the word semantic features of each word. That is, on the basis that the semantic feature extraction sub-model includes the first to Mth third coding layers, the target text and each word are respectively input into the semantic feature extraction sub-model to obtain the text semantic features of the target text and the word semantic features of each word. The specific implementation process can be as follows:
[0115] For the target text, the first text encoding feature is input into the second third encoding layer for feature extraction to obtain the second text encoding feature of the target text, until the (M-1)th text encoding feature is input into the Mth third encoding layer for feature extraction to obtain the Mth text encoding feature of the target text, and the Mth text encoding feature is determined as the text semantic feature of the target text;
[0116] For any word, the first word coding feature of the word is input into the second third coding layer for feature extraction to obtain the second word coding feature of the word, until the (M-1)th word coding feature is input into the Mth third coding layer for feature extraction to obtain the Mth word coding feature of the word, and the Mth word coding feature is determined as the word semantic feature of the word.
[0117] In practical applications, the feature extraction model includes M third coding layers, wherein the structural feature extraction submodel includes the first third coding layer, and the semantic feature extraction submodel includes the first third coding layer to the Nth third coding layer. Since when performing structural feature extraction, the target text is input into the first third coding layer for feature extraction to obtain the first text coding feature of the target text, and each word is input into the first third coding layer for feature extraction to obtain the first word coding feature of each word, at this time, the first text coding feature and each first word coding feature can be respectively input into the second third coding layer for processing to obtain the second text coding feature of the target text and the second word coding feature of each word; then the second text coding feature and each second word coding feature are respectively input into the third third coding layer for processing, and so on, until the (M-1)th text coding feature and each (M-1)th word coding feature are respectively input into the Mth third coding layer for processing to obtain the Mth text coding feature of the target text and the Mth word coding feature of each word, that is, the text semantic feature of the target text and the word semantic feature of each word. In this way, as the number of coding layers increases, semantic information can be extracted more accurately. Therefore, using multiple coding layers as semantic feature extraction sub-models to extract semantic features of the target text and each word can make the obtained text semantic features and word semantic features more accurate.
[0118] See also Figure 5 In a keyword extraction method provided in an embodiment of the present application, a structural schematic diagram of another feature extraction model is provided: the feature extraction model includes multiple third coding layers and an output layer, wherein the first third coding layer constitutes a structural feature extraction sub-model, and the first third coding layer to the last third coding layer constitute a semantic feature extraction sub-model. The structural feature extraction sub-model is used to receive the target text and each word, and extract the structural features of the target text and each word, and obtain the text structural features and the structural features of each word; the semantic feature extraction sub-model is used to receive the target text and each word, and extract the semantic features of the target text and each word, and obtain the text semantic features and the semantic features of each word; the output layer is used to receive and process the text structural features and the text semantic features, obtain and output the text features, and receive and process the structural features and the semantic features of each word, and obtain and output the features of each word. By making the structural feature extraction sub-model and the semantic feature extraction sub-model have an intersection, the accuracy and efficiency of feature extraction can be guaranteed while reducing the coding layers included in the feature extraction model.
[0119] Before obtaining the pre-trained feature extraction model, the language representation model needs to be trained in order to obtain a feature extraction model with feature extraction function. That is, before obtaining the pre-trained feature extraction model, it also includes:
[0120] Obtaining a sample text set and a preset language representation model, wherein the language representation model includes a structural feature extraction sub-model, a semantic feature extraction sub-model and an output layer;
[0121] Extract at least two sample texts from the sample text set, input each sample text into the structural feature extraction sub-model, and obtain the predicted structural features of each sample text;
[0122] The predicted structural features of each sample text are respectively input into the semantic feature extraction sub-model to obtain the predicted semantic features of each sample text;
[0123] The predicted structural features and predicted semantic features of each sample text are respectively input into the output layer for processing, and the predicted text features of each sample text are output;
[0124] Calculate the loss value based on the predicted text features of each sample text;
[0125] According to the loss value, the model parameters of the structural feature extraction submodel and the semantic feature extraction submodel in the language representation model are adjusted, and the step of extracting at least two sample texts from the sample text set is continued. When the preset training stop condition is reached, the trained language representation model is determined as the feature extraction model.
[0126] Specifically, a language representation model refers to a pre-specified pre-trained neural network model, such as a RoBERTa model; sample text refers to samples used for a language representation model, which may be sentences, words, articles, etc.; a sample text set refers to a collection of multiple sample texts; predicted structural features refer to structural features of sample texts extracted by a structural feature extraction submodel; predicted semantic features refer to semantic features of sample texts extracted by a semantic feature extraction submodel; predicted text features refer to features that are a combination of predicted structural features and predicted semantic features; the training stop condition may be that the loss value is less than or equal to a preset threshold, or that the number of iterative training reaches a preset iteration value, or that the loss value converges, that is, the loss value no longer decreases as training continues.
[0127] In practical applications, there are many ways to obtain sample text sets and preset language representation models. For example, the operator may send a training instruction for the language representation model to the execution subject, or send an instruction to obtain the sample text set and the preset language representation model. Accordingly, after receiving the instruction, the execution subject begins to obtain the sample text set and the preset language representation model; or the server may automatically obtain the sample text set and the preset language representation model at preset time intervals. For example, after a preset time, the server with a model training function automatically obtains the sample text set and the preset language representation model in the specified access area; or after a preset time, the terminal with a model training function automatically obtains the sample text set and the preset language representation model stored locally. This specification does not limit the method of obtaining the sample text set and the preset language representation model.
[0128] After obtaining the sample text set and the preset language representation model, extract multiple sample texts from the sample text set, then input the multiple sample texts into the structural feature extraction sub-model, the structural feature extraction sub-model extracts the structural features of each sample text, and outputs the predicted structural features of each sample text, and inputs each sample text into the semantic feature extraction sub-model, the semantic feature extraction sub-model extracts the semantic features of each sample text, and outputs the predicted semantic features of each sample text. Then, the predicted semantic features and predicted structural features of each sample text are input into the output layer, and the output layer analyzes and processes the predicted semantic features and predicted structural features of each sample text to obtain and output the predicted text features of each sample text. Next, determine the loss value according to the predicted text features of each sample text and the preset loss function, and adjust the model parameters of the language representation model according to the loss value when the preset training stop condition is not reached, that is, the model parameters of the structural feature extraction sub-model and the semantic feature extraction sub-model, and then extract multiple sample texts from the sample text set again to perform the next round of training; when the preset training stop condition is reached, determine the trained language representation model as the feature extraction model. In this way, by training the language representation model with multiple sample texts, the accuracy and speed of feature extraction by the feature extraction model can be improved, and the robustness of the feature extraction model can be improved.
[0129] It should be noted that the training stop condition may be to traverse all sample texts in the sample text set K times, where K is a preset value, that is, each sample text in the sample text set is used K times to train the feature extraction model.
[0130] In addition, in order to achieve keyword extraction for target text in a specific field, sample texts in the specific field can be obtained for the specific field to train the language representation model, and a feature extraction model dedicated to the specific field can be obtained. For example, sample texts in the medical field are used to train the language representation model to obtain a feature extraction model dedicated to the medical field. For another example, sample texts in the geography field are used to train the language representation model to obtain a feature extraction model dedicated to the geography field.
[0131] In a possible implementation of the embodiments of this specification, in order to further improve the robustness of the feature extraction model, when determining the loss value, it can be determined that the loss value is determined by a contrastive learning method. That is, the loss value is calculated based on the predicted text features of each sample text. The specific implementation process can be as follows:
[0132] For any predicted text feature of a sample text, the predicted text feature of the sample text is input twice into a preset random deactivation layer for processing to obtain a first sample text feature and a second sample text feature of the sample text;
[0133] Calculating a second similarity between each sample text feature, where the sample text features include a first sample text feature and a second sample text feature;
[0134] According to each second similarity, a loss value is calculated.
[0135] Specifically, the random dropout layer refers to a random hidden code processing layer; the first sample text feature refers to the output result obtained by inputting a certain predicted text feature into the random dropout layer for the first time; the second sample text feature refers to the output result obtained by inputting a certain predicted text feature into the random dropout layer for the second time; the second similarity refers to the similarity between the first sample text feature and the second sample text feature.
[0136] In practical applications, each predicted text feature can be repeatedly input into the predicted text feature twice to obtain the first sample text feature and the second sample text feature of each predicted text feature; then, the second similarity between all the first sample text features and the second sample text features is calculated; then, the second similarity is input into the preset loss function to obtain the loss value.
[0137] For example, there are two sample texts. The predicted text features of the first sample text are input to the random deactivation layer for processing for the first time to obtain sample text features m1, and the predicted text features of the first sample text are input to the random deactivation layer for processing for the second time to obtain sample text features m2; the predicted text features of the second sample text are input to the random deactivation layer for processing for the first time to obtain sample text features m3, and the predicted text features of the second sample text are input to the random deactivation layer for processing for the second time to obtain sample text features m4. Then, the second similarities among the sample text features m1, sample text features m2, sample text features m3, and sample text features m4 are calculated, and then each second similarity is input to the loss function shown in Formula 2 for calculation to obtain a loss value.
[0138]
[0139] In formula 2, L i represents the loss value corresponding to the sample text feature; P i It represents the similarity between sample text feature i and its positive sample (another sample text feature of the sample text corresponding to sample text feature i), and it represents the similarity between sample text feature i and negative sample j (any sample text feature of other sample texts other than the sample text corresponding to sample text feature i); N represents the number of sample text features.
[0140] It should be noted that the SimCSE unsupervised algorithm can be used to train the language representation model. In this way, a small amount of unlabeled sample text in a specific field can be used to train the language representation model (such as pre-trained BERT) and the model suitable for the specific field after unsupervised training, avoiding the problem of uneven distribution caused by directly using the anisotropy in the BERT sentence vector.
[0141] In a possible implementation of the embodiments of the present specification, in order to ensure that all sample texts in the sample text set can be used for training the feature extraction model, the sample texts in the sample text set can be grouped and the sample text groups can be extracted in a certain order. That is, before extracting at least two sample texts from the sample text set, the method further includes:
[0142] Grouping the sample texts in the sample text set to obtain at least one sample text group, wherein the sample text group contains at least two sample texts;
[0143] Accordingly, at least two sample texts are extracted from the sample text set, including:
[0144] A sample text group is extracted from the sample text set according to a preset extraction order.
[0145] Specifically, grouping refers to dividing multiple sample texts into several groups; a sample text group refers to a group containing multiple sample texts after division; and a preset extraction order refers to a pre-set order for extracting sample text groups, such as sequential extraction or reverse extraction.
[0146] In practical applications, after obtaining a sample text set, the sample texts in the sample text set are differentiated into multiple sample text groups containing at least two sample texts, and then based on a preset extraction order, a sample text group is extracted from the sample text set each time, that is, at least two sample texts are extracted from the sample text set. In this way, it can be ensured that all sample texts in the sample text set can be used for training the feature extraction model, which is conducive to improving the robustness of the feature extraction model. In addition, taking a sample text group each time can effectively avoid the problem of extracting too many sample texts, which causes the language representation model to crash due to excessive data processing.
[0147] For example, the sample text set contains 100 sample texts, and according to the principle of even distribution, the 100 sample texts are divided into 20 sample text groups, each of which contains 5 sample texts. Then, in the order from the first sample text group to the twentieth sample text, that is, sequential extraction, one sample text group is extracted in turn to train the language representation model.
[0148] The present application provides a keyword extraction method, which obtains a target text; extracts the text structure features and text semantic features of the target text, as well as the word structure features and word semantic features of each word in the target text; determines text features according to the text structure features and text semantic features, and determines word features according to the word structure features and word semantic features; determines keywords from the words according to the text features and word features of each word. Determining text features through text structure features and text semantic features, and determining word features through word structure features and word semantic features can more accurately determine the semantic and structural level information of the text or word, that is, the text features and word features are more accurate, and then keywords can be determined based on text features and word features, thereby improving the efficiency of determining keywords. The problem of uneven distribution of high- and low-frequency words in the semantic space is avoided. In addition, compared with traditional word frequency statistics and graph algorithms, the keyword extraction method of the present application focuses more on the relevant technologies of keyword extraction of short texts.
[0149] The following combination Figure 6 Taking the application of the keyword extraction method provided in this application to a short text as an example, the keyword extraction method is further described. Figure 6 A processing flow chart of a keyword extraction method applied to short text provided by an embodiment of the present application is shown, which specifically includes the following steps:
[0150] Step 602: Obtain a sample short text set and a preset language representation model, wherein the language representation model includes a structural feature extraction sub-model, a semantic feature extraction sub-model and an output layer.
[0151] Step 604: extract at least two sample short texts from the sample short text set, input each sample short text into the structural feature extraction sub-model, and obtain the predicted structural features of each sample short text.
[0152] In one or more optional embodiments of the present specification, before extracting at least two sample short texts from the sample short text set, the method further includes:
[0153] Grouping the sample short texts in the sample short text set to obtain at least one sample short text group, wherein the sample short text group contains at least two sample short texts;
[0154] Extract at least two sample short texts from the sample short text set, including:
[0155] A sample short text group is extracted from the sample short text set according to a preset extraction order.
[0156] Step 606: input the predicted structural features of each sample short text into the semantic feature extraction sub-model respectively to obtain the predicted semantic features of each sample short text.
[0157] Step 608: The predicted structural features and predicted semantic features of each sample short text are respectively input into the output layer for processing, and the predicted text features of each sample short text are output.
[0158] Step 610: for any predicted text feature of a sample short text, the predicted text feature of the sample short text is input twice into a preset random deactivation layer for processing to obtain a first sample text feature and a second sample text feature of the sample short text.
[0159] Step 612: Calculate the second similarity between each sample text feature, where the sample text features include the first sample text feature and the second sample text feature.
[0160] Step 614: Calculate a loss value according to each second similarity.
[0161] Step 616: According to the loss value, adjust the model parameters of the structural feature extraction submodel and the semantic feature extraction submodel in the language representation model, continue to execute the step of extracting at least two sample short texts from the sample short text set, and when the preset training stop condition is reached, determine the trained language representation model as the feature extraction model.
[0162] Step 618: Obtain the target short text.
[0163] Step 620: Perform word segmentation processing on the target short text to obtain multiple candidate words, and determine the part of speech of each candidate word.
[0164] Step 622: Filter out multiple words from multiple candidate words based on parts of speech.
[0165] Step 624: Input the target short text and each word into the structural feature extraction sub-model respectively to obtain the text structural features of the target short text and the word structural features of each word.
[0166] In one or more optional embodiments of the present specification, the structural feature extraction sub-model includes a first encoding layer;
[0167] Accordingly, the target short text and each word in the target short text are respectively input into the structural feature extraction sub-model to obtain the text structure features of the target short text and the word structure features of each word, including:
[0168] Input the target short text into the first encoding layer for feature extraction to obtain the text structure features of the target short text;
[0169] For any word, the word is input into the first coding layer for feature extraction to obtain the word structure features of the candidate word.
[0170] In one or more optional embodiments of the present specification, the feature extraction model includes M third coding layers, M is a positive integer greater than 2, and the structural feature extraction sub-model includes the first third coding layer;
[0171] Accordingly, the target short text and each word in the target short text are respectively input into the structural feature extraction sub-model to obtain the text structure features of the target short text and the word structure features of each word, including:
[0172] For the target short text, the target short text is input into the first third coding layer for feature extraction to obtain a first text coding feature of the target short text, and the first text coding feature is determined as a text structure feature of the target short text;
[0173] For any word, the word is input into the first third coding layer for feature extraction to obtain the first word coding feature of the word, and the first word coding feature is determined as the word structure feature of the word.
[0174] Step 626: Input the target short text and each word into the semantic feature extraction sub-model respectively to obtain the text semantic features of the target short text and the word semantic features of each word.
[0175] In one or more optional embodiments of the present specification, the semantic feature extraction sub-model includes N second encoding layers, where N is a positive integer greater than 2;
[0176] Accordingly, the target short text and each word are respectively input into the semantic feature extraction sub-model to obtain the text semantic features of the target short text and the word semantic features of each word, including:
[0177] For the target short text, starting from the first second coding layer, the output of the previous second coding layer is used as the input of the current second coding layer for feature extraction, until the Nth second coding layer, and the text semantic features of the target short text are output, wherein the input of the first second coding layer is the target short text;
[0178] For any word, starting from the first second coding layer, the output of the previous second coding layer is used as the input of the current second coding layer for feature extraction until the Nth second coding layer, and the semantic features of the word are output, where the input of the first second coding layer is the word.
[0179] In one or more optional embodiments of the present specification, the semantic feature extraction sub-model includes 1st to Mth third encoding layers;
[0180] Accordingly, the target short text and each word are respectively input into the semantic feature extraction sub-model to obtain the text semantic features of the target short text and the word semantic features of each word, including:
[0181] For the target short text, the first text encoding feature is input into the second third encoding layer for feature extraction to obtain the second text encoding feature of the target short text, until the (M-1)th text encoding feature is input into the Mth third encoding layer for feature extraction to obtain the Mth text encoding feature of the target short text, and the Mth text encoding feature is determined as the text semantic feature of the target short text;
[0182] For any word, the first word coding feature of the word is input into the second third coding layer for feature extraction to obtain the second word coding feature of the word, until the (M-1)th word coding feature is input into the Mth third coding layer for feature extraction to obtain the Mth word coding feature of the word, and the Mth word coding feature is determined as the word semantic feature of the word.
[0183] Step 628: For the target short text, the text structure features and text semantic features are input into the output layer for processing, and the text features of the target short text are output.
[0184] Step 630: For any word, the word structure feature and word semantic feature of the word are input into the output layer for processing, and the word feature of the word is output.
[0185] Step 632: Determine the first similarity between the word feature and the text feature of each word.
[0186] Step 634: Determine keywords of the target short text from the plurality of words according to the first similarities.
[0187] The present application provides a keyword extraction method that determines text features through text structure features and text semantic features, and determines word features through word structure features and word semantic features, which can more accurately determine the semantic and structural level information of text or words, that is, the text features and word features are more accurate, and then keywords can be determined based on text features and word features, thereby improving the efficiency of determining keywords. The problem of uneven distribution of high-frequency and low-frequency words in the semantic space is avoided.
[0188] Corresponding to the above method embodiment, the present application also provides a keyword extraction device embodiment, Figure 7 FIG. 1 is a schematic diagram showing the structure of a keyword extraction device provided by an embodiment of the present application. Figure 7 As shown, the device comprises:
[0189] A first acquisition module 702 is configured to acquire a target text;
[0190] An extraction module 704 is configured to extract text structure features and text semantic features of the target text, as well as word structure features and word semantic features of each word in the target text;
[0191] The first determination module 706 is configured to determine text features according to text structure features and text semantic features, and determine word features according to word structure features and word semantic features;
[0192] The second determination module 708 is configured to determine keywords from each word according to the text feature and the word feature of each word.
[0193] In one or more optional embodiments of the present specification, the second determining module 708 is further configured to:
[0194] Determine the first similarity between the word feature and the text feature of each word respectively;
[0195] According to each first similarity, a keyword of the target text is determined from the plurality of words.
[0196] In one or more optional embodiments of the present specification, the device further includes a second acquisition model configured to:
[0197] Obtaining a pre-trained feature extraction model, wherein the feature extraction model includes a structural feature extraction sub-model, a semantic feature extraction sub-model, and an output layer;
[0198] Accordingly, the extraction module 704 is further configured to:
[0199] Input the target text and each word in the target text into the structural feature extraction sub-model respectively to obtain the text structural features of the target text and the word structural features of each word;
[0200] Input the target text and each word into the semantic feature extraction sub-model respectively to obtain the text semantic features of the target text and the word semantic features of each word;
[0201] Accordingly, the first determining module 706 is further configured to:
[0202] For the target text, the text structure features and text semantic features are input into the output layer for processing, and the text features of the target text are output;
[0203] For any word, the word structure features and word semantic features of the word are input into the output layer for processing, and the word features of the word are output.
[0204] In one or more optional embodiments of the present specification, the structural feature extraction sub-model includes a first encoding layer;
[0205] Accordingly, the extraction module 704 is further configured to:
[0206] Input the target text into the first encoding layer for feature extraction to obtain the text structure features of the target text;
[0207] For any word, the word is input into the first coding layer for feature extraction to obtain the word structure features of the candidate word.
[0208] In one or more optional embodiments of the present specification, the semantic feature extraction sub-model includes N second encoding layers, where N is a positive integer greater than 2;
[0209] Accordingly, the extraction module 704 is further configured to:
[0210] For the target text, starting from the first second coding layer, the output of the previous second coding layer is used as the input of the current second coding layer for feature extraction, until the Nth second coding layer, and the text semantic features of the target text are output, wherein the input of the first second coding layer is the target text;
[0211] For any word, starting from the first second coding layer, the output of the previous second coding layer is used as the input of the current second coding layer for feature extraction until the Nth second coding layer, and the semantic features of the word are output, where the input of the first second coding layer is the word.
[0212] In one or more optional embodiments of the present specification, the feature extraction model includes M third coding layers, M is a positive integer greater than 2, and the structural feature extraction sub-model includes the first third coding layer;
[0213] Accordingly, the extraction module 704 is further configured to:
[0214] For the target text, the target text is input into the first third coding layer for feature extraction to obtain the first text coding feature of the target text, and the first text coding feature is determined as the text structure feature of the target text;
[0215] For any word, the word is input into the first third coding layer for feature extraction to obtain the first word coding feature of the word, and the first word coding feature is determined as the word structure feature of the word.
[0216] In one or more optional embodiments of the present specification, the semantic feature extraction sub-model includes 1st to Mth third encoding layers;
[0217] Accordingly, the extraction module 704 is further configured to:
[0218] For the target text, the first text encoding feature is input into the second third encoding layer for feature extraction to obtain the second text encoding feature of the target text, until the (M-1)th text encoding feature is input into the Mth third encoding layer for feature extraction to obtain the Mth text encoding feature of the target text, and the Mth text encoding feature is determined as the text semantic feature of the target text;
[0219] For any word, the first word coding feature of the word is input into the second third coding layer for feature extraction to obtain the second word coding feature of the word, until the (M-1)th word coding feature is input into the Mth third coding layer for feature extraction to obtain the Mth word coding feature of the word, and the Mth word coding feature is determined as the word semantic feature of the word.
[0220] In one or more optional embodiments of the present specification, the device further includes a training module configured to:
[0221] Obtaining a sample text set and a preset language representation model, wherein the language representation model includes a structural feature extraction sub-model, a semantic feature extraction sub-model and an output layer;
[0222] Extract at least two sample texts from the sample text set, input each sample text into the structural feature extraction sub-model, and obtain the predicted structural features of each sample text;
[0223] The predicted structural features of each sample text are respectively input into the semantic feature extraction sub-model to obtain the predicted semantic features of each sample text;
[0224] The predicted structural features and predicted semantic features of each sample text are respectively input into the output layer for processing, and the predicted text features of each sample text are output;
[0225] Calculate the loss value based on the predicted text features of each sample text;
[0226] According to the loss value, the model parameters of the structural feature extraction submodel and the semantic feature extraction submodel in the language representation model are adjusted, and the step of extracting at least two sample texts from the sample text set is continued. When the preset training stop condition is reached, the trained language representation model is determined as the feature extraction model.
[0227] In one or more optional embodiments of the present specification, the training module is further configured to:
[0228] For any predicted text feature of a sample text, the predicted text feature of the sample text is input twice into a preset random deactivation layer for processing to obtain a first sample text feature and a second sample text feature of the sample text;
[0229] Calculating a second similarity between each sample text feature, where the sample text features include a first sample text feature and a second sample text feature;
[0230] According to each second similarity, a loss value is calculated.
[0231] In one or more optional embodiments of the present specification, the training module is further configured to:
[0232] Grouping the sample texts in the sample text set to obtain at least one sample text group, wherein the sample text group contains at least two sample texts;
[0233] A sample text group is extracted from the sample text set according to a preset extraction order.
[0234] In one or more optional embodiments of the present specification, the device further includes a word segmentation module configured to:
[0235] The target text is segmented to obtain multiple words.
[0236] In one or more optional embodiments of the present specification, the word segmentation module is further configured as follows:
[0237] Perform word segmentation on the target text to obtain multiple candidate words, and determine the part of speech of each candidate word;
[0238] According to the part of speech, multiple words are selected from multiple candidate words.
[0239] The present application provides a keyword extraction device that determines text features through text structure features and text semantic features, and determines word features through word structure features and word semantic features, which can more accurately determine the semantic and structural level information of text or words, that is, the text features and word features are more accurate, and then keywords can be determined based on text features and word features, thereby improving the efficiency of determining keywords. The problem of uneven distribution of high-frequency and low-frequency words in the semantic space is avoided.
[0240] The above is a schematic scheme of a keyword extraction device of the present embodiment. It should be noted that the technical scheme of the keyword extraction device and the technical scheme of the keyword extraction method mentioned above belong to the same concept. For details not described in detail in the technical scheme of the keyword extraction device, please refer to the description of the technical scheme of the keyword extraction method mentioned above. In addition, each component in the device embodiment should be understood as a functional module that must be established to implement each step of the program flow or each step of the method, and each functional module is not an actual functional division or separation definition. The device claim defined by such a group of functional modules should be understood as a functional module architecture that mainly implements the solution through the computer program recorded in the specification, and should not be understood as a physical device that mainly implements the solution through hardware.
[0241] Figure 8 The structure block diagram of a computing device 800 provided according to an embodiment of the present application is shown. The components of the computing device 800 include but are not limited to a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and the database 850 is used to store data.
[0242] The computing device 800 also includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.
[0243] In one embodiment of the present application, the above components of the computing device 800 and Figure 8 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 8 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of the present application. Those skilled in the art may add or replace other components as needed.
[0244] The computing device 800 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or PC. The computing device 800 may also be a mobile or stationary server.
[0245] The processor 820 is used to execute computer executable instructions of the keyword extraction method.
[0246] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above keyword extraction method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above keyword extraction method.
[0247] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which are used in a keyword extraction method when executed by a processor.
[0248] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the keyword extraction method described above are of the same concept, and the details not described in detail in the technical scheme of the storage medium can be found in the description of the technical scheme of the keyword extraction method described above.
[0249] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0250] An embodiment of the present application further provides a chip storing a computer program, which implements the steps of the keyword extraction method when executed by the chip.
[0251] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0252] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0253] The preferred embodiments of the present application disclosed above are only used to help explain the present application. The optional embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the present application. The present application selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present application, so that those skilled in the art can understand and use the present application well. The present application is only limited by the claims and their full scope and equivalents.
Claims
1. A keyword extraction method, characterized in that: include: Acquire a target text, wherein the target text is a short text; Extracting text structure features and text semantic features of the target text, as well as word structure features and word semantic features of each word in the target text, wherein the text structure features refer to the structural features corresponding to the text; the text semantic features refer to the semantic features corresponding to the text; the word structure features refer to the structural features corresponding to the words; the word semantic features refer to the semantic features corresponding to the words; the structural features refer to the surface information features of multiple text units; the semantic features refer to the features corresponding to the language meanings of multiple text units; Determine a text feature according to the text structure feature and the text semantic feature, and determine a word feature according to the word structure feature and the word semantic feature, wherein the text feature is a combination of the text structure feature and the text semantic feature, and the word feature is a combination of the word structure feature and the word semantic feature; Determine keywords from the words based on the text features and the word features of the words, wherein determining keywords from the words based on the text features and the word features of the words includes: respectively determining a first similarity between the word feature of each word and the text feature; and determining keywords of the target text from the multiple words based on the first similarities.
2. The method according to claim 1, characterized in that Before extracting the text structure features and text semantic features of the target text, and the word structure features and word semantic features of each word in the target text, the method further includes: Obtaining a pre-trained feature extraction model, wherein the feature extraction model includes a structural feature extraction sub-model, a semantic feature extraction sub-model and an output layer; Accordingly, the extracting of the text structure features and text semantic features of the target text, as well as the word structure features and word semantic features of each word in the target text, includes: Inputting the target text and each word in the target text into the structural feature extraction sub-model respectively to obtain the text structural features of the target text and the word structural features of each word; Inputting the target text and each word into the semantic feature extraction sub-model respectively to obtain the text semantic features of the target text and the word semantic features of each word; Correspondingly, determining text features according to the text structure features and text semantic features, and determining word features according to the word structure features and word semantic features, includes: For the target text, the text structure features and the text semantic features are input into the output layer for processing, and the text features of the target text are output; For any word, the word structure feature and the word semantic feature of the word are input into the output layer for processing, and the word feature of the word is output.
3. The method according to claim 2, characterized in that The structural feature extraction sub-model includes a first encoding layer; Correspondingly, the target text and each word in the target text are respectively input into the structural feature extraction sub-model to obtain the text structural features of the target text and the word structural features of each word, including: Inputting the target text into the first coding layer for feature extraction to obtain text structure features of the target text; For any word, the word is input into the first coding layer for feature extraction to obtain the word structure features of the candidate word.
4. The method according to claim 2 or 3, characterized in that: The semantic feature extraction sub-model includes N second encoding layers, where N is a positive integer greater than 2; Correspondingly, the target text and each word are respectively input into the semantic feature extraction sub-model to obtain the text semantic features of the target text and the word semantic features of each word, including: For the target text, starting from the first second coding layer, the output of the previous second coding layer is used as the input of the current second coding layer to perform feature extraction, until the Nth second coding layer, and the text semantic features of the target text are output, wherein the input of the first second coding layer is the target text; For any word, starting from the first second coding layer, the output of the previous second coding layer is used as the input of the current second coding layer for feature extraction until the Nth second coding layer, and the semantic features of the word are output, wherein the input of the first second coding layer is the word.
5. The method according to claim 2, characterized in that: The feature extraction model includes M third coding layers, where M is a positive integer greater than 2, and the structural feature extraction sub-model includes the first third coding layer; Correspondingly, the target text and each word in the target text are respectively input into the structural feature extraction sub-model to obtain the text structural features of the target text and the word structural features of each word, including: For the target text, input the target text into the first third coding layer to perform feature extraction to obtain a first text coding feature of the target text, and determine the first text coding feature as a text structure feature of the target text; For any word, the word is input into the first third coding layer for feature extraction to obtain the first word coding feature of the word, and the first word coding feature is determined as the word structure feature of the word.
6. The method according to claim 5, characterized in that The semantic feature extraction sub-model includes the 1st to the Mth third encoding layers; Correspondingly, the target text and each word are respectively input into the semantic feature extraction sub-model to obtain the text semantic features of the target text and the word semantic features of each word, including: For the target text, the first text encoding feature is input into the second third encoding layer for feature extraction to obtain the second text encoding feature of the target text, until the (M-1)th text encoding feature is input into the Mth third encoding layer for feature extraction to obtain the Mth text encoding feature of the target text, and the Mth text encoding feature is determined as the text semantic feature of the target text; For any word, the first word coding feature of the word is input into the second third coding layer for feature extraction to obtain the second word coding feature of the word, until the (M-1)th word coding feature is input into the Mth third coding layer for feature extraction to obtain the Mth word coding feature of the word, and the Mth word coding feature is determined as the word semantic feature of the word.
7. The method according to claim 2, characterized in that Before obtaining the pre-trained feature extraction model, the method further includes: Acquire a sample text set and a preset language representation model, wherein the language representation model includes a structural feature extraction sub-model, a semantic feature extraction sub-model and an output layer; Extracting at least two sample texts from the sample text set, and inputting each sample text into the structural feature extraction sub-model to obtain predicted structural features of each sample text; Inputting the predicted structural features of each sample text into the semantic feature extraction sub-model respectively to obtain the predicted semantic features of each sample text; Inputting the predicted structural features and predicted semantic features of each sample text into the output layer for processing, and outputting the predicted text features of each sample text; Calculate the loss value based on the predicted text features of each sample text; According to the loss value, the model parameters of the structural feature extraction submodel and the semantic feature extraction submodel in the language representation model are adjusted, and the step of extracting at least two sample texts from the sample text set is continued. When a preset training stop condition is reached, the trained language representation model is determined as a feature extraction model.
8. The method according to claim 7, characterized in that The step of calculating the loss value according to the predicted text features of each sample text includes: For any predicted text feature of a sample text, the predicted text feature of the sample text is input twice into a preset random deactivation layer for processing to obtain a first sample text feature and a second sample text feature of the sample text; Calculating a second similarity between each sample text feature, wherein the sample text features include a first sample text feature and a second sample text feature; According to each second similarity, a loss value is calculated.
9. The method according to claim 7, characterized in that: Before extracting at least two sample texts from the sample text set, the method further includes: Grouping the sample texts in the sample text set to obtain at least one sample text group, wherein the sample text group contains at least two sample texts; Accordingly, extracting at least two sample texts from the sample text set includes: A sample text group is extracted from the sample text set according to a preset extraction order.
10. The method according to claim 1, characterized in that Before extracting the text structure features and text semantic features of the target text, and the word structure features and word semantic features of each word in the target text, the method further includes: The target text is segmented to obtain a plurality of words.
11. The method according to claim 10, characterized in that The target text is segmented to obtain a plurality of words, including: Performing word segmentation processing on the target text to obtain multiple candidate words, and determining the part of speech of each candidate word; According to the part of speech, a plurality of words are selected from the plurality of candidate words.
12. A keyword extraction device, characterized in that: include: A first acquisition module is configured to acquire a target text, wherein the target text is a short text; The extraction module is configured to extract text structure features and text semantic features of the target text, as well as word structure features and word semantic features of each word in the target text, wherein the text structure features refer to the structural features corresponding to the text; the text semantic features refer to the semantic features corresponding to the text; the word structure features refer to the structural features corresponding to the words; the word semantic features refer to the semantic features corresponding to the words; the structural features refer to the surface information features of multiple text units; the semantic features refer to the features corresponding to the language meanings of multiple text units; A first determination module is configured to determine a text feature according to the text structure feature and the text semantic feature, and to determine a word feature according to the word structure feature and the word semantic feature, wherein the text feature is a combination of the text structure feature and the text semantic feature, and the word feature is a combination of the word structure feature and the word semantic feature; The second determination module is configured to determine keywords from the words based on the text features and the word features of the words, wherein the determining keywords from the words based on the text features and the word features of the words includes: respectively determining a first similarity between the word feature of each word and the text feature; and determining the keywords of the target text from the multiple words based on each of the first similarities.
13. A computing device, characterized in that: include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the keyword extraction method according to any one of claims 1 to 11.
14. A computer-readable storage medium storing computer instructions, characterized in that: When the instruction is executed by the processor, the steps of the keyword extraction method described in any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Keyword extraction method and device, equipment and storage medium
CN114282528A