Keyword classification method and device based on text information, equipment and storage medium

Through preset language model scoring and filtering, embedded vector calculation and clustering algorithm classification methods, the shortcomings of traditional keyword extraction methods in capturing contextual relationships and semantic information are solved, and more accurate and relevant keyword extraction effects are achieved.

CN120067334APending Publication Date: 2025-05-30INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510227917.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional keyword extraction methods such as TF-IDF and TextRank are difficult to capture the contextual relationship and semantic information between words, resulting in the extracted keywords being one-sided, especially when dealing with short or complex texts.

Method used

The text is scored and selected and selected through the preset language model to obtain high-quality target text; the word embedding vector is calculated using the preset model, and keywords are extracted through semantic similarity evaluation; and keywords are classified through clustering algorithms to generate keyword summary and classification results.

Benefits of technology

It improves the accuracy and relevance of keyword extraction, can capture subtle semantic differences between texts more accurately, and enhances the accessibility and practicality of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067334A_ABST
    Figure CN120067334A_ABST
Patent Text Reader

Abstract

The invention discloses a keyword classification method and device based on text information, equipment and a storage medium, and relates to the technical field of computers.The method comprises the steps that original text information is obtained, a preset language model is utilized to score at least one text, the at least one text is preferentially screened, and one or more target texts are obtained; calculating at least one embedded vector corresponding to the one or more target texts by using a preset model, and inputting the at least one embedded vector into a text processing algorithm for semantic similarity evaluation to obtain at least one keyword and a corresponding relationship between each keyword and the target text; and classifying the at least one keyword through a clustering algorithm to generate a keyword abstract and a classification result of the one or more target texts. According to the method, the semantic similarity between the words is calculated by embedding the vectors, and the subtle semantic difference between the texts is more accurately captured, so that the keywords highly related to the text content are extracted, and the accuracy and correlation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a keyword classification method, apparatus, device, and storage medium based on text information. Background Art

[0002] With the development of the Internet, a vast amount of text data has been continuously generated and accumulated. To process and analyze this text data efficiently and accurately, it is necessary to extract key information or keywords from the vast amount of text data. Traditional keyword extraction methods such as TF-IDF (Term Frequency-Inverse Document Frequency, a commonly used weighting technique in information retrieval and data mining), word frequency statistics, etc., although simple and easy to implement, often have unsatisfactory effects when dealing with complex texts and are difficult to accurately capture the core content. Because the TF-IDF method cannot capture the context relationship and semantic information between words, the extracted keywords are often one-sided, especially when dealing with short texts or complex texts, the effect is poor. Summary of the Invention

[0003] This application provides a keyword classification method, apparatus, device, and storage medium based on text information to solve the problem of one-sided and poor keyword extraction content in text information.

[0004] This application provides a keyword classification method based on text information, and the method includes:

[0005] Obtain the original text information, where the original text information includes at least one piece of text;

[0006] Use a preset language model to score at least one piece of text, and preferentially screen at least one piece of text according to at least one scoring result to obtain one or more target texts;

[0007] Use a preset model to calculate at least one embedding vector corresponding to one or more target texts, and input at least one embedding vector into a text processing algorithm for semantic similarity evaluation to obtain at least one keyword corresponding to one or more target texts, and the corresponding relationship between each keyword and the target text where it is located;

[0008] According to at least one keyword and the corresponding relationship, classify at least one keyword through a clustering algorithm to generate one or more keyword summaries and the classification results of one or more target texts.

[0009] This application also provides a keyword classification apparatus based on text information, and the apparatus includes:

[0010] An acquisition module, configured to acquire the original text information, where the original text information includes at least one piece of text;

[0011] A screening module, configured to score at least one piece of text by using a preset language model, and preferentially screen the at least one piece of text according to at least one scoring result to obtain one or more target texts;

[0012] A processing module, configured to calculate at least one embedding vector corresponding to one or more target texts by using a preset model, and input the at least one embedding vector into a text processing algorithm for semantic similarity evaluation, to obtain at least one keyword corresponding to one or more target texts, and the corresponding relationship between each keyword and the target text where it is located;

[0013] A clustering module, configured to classify at least one keyword according to at least one keyword and the corresponding relationship by using a clustering algorithm, to generate one or more keyword summaries and classification results for one or more target texts.

[0014] This application also provides a computer device, including a memory and a processor, wherein the memory is used to store a computer program; the processor is used to implement the steps of any one of the above keyword classification methods based on text information when executing the computer program.

[0015] This application also provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program implements the steps of any one of the above keyword classification methods based on text information when executed by a processor.

[0016] In addition, this application also provides a computer program product, including a computer program, and the computer program implements the steps of any one of the above keyword classification methods based on text information when executed by a processor.

[0017] The keyword classification method provided in this embodiment scores and preferentially screens the original text information through a preset language model, can quickly identify and retain the most valuable or highly relevant text, effectively reduces the amount of data for subsequent processing, and improves the overall processing efficiency.

[0018] This method also calculates word embedding vectors by using a preset model, calculates the semantic similarity between words through the embedding vectors, can more accurately capture the subtle semantic differences between texts, thereby extracting keywords highly relevant to the text content, and improves the accuracy and relevance of the keywords. The network structure of this method can capture the semantic relevance between words, rather than relying solely on word co-occurrence relationships, thereby significantly improving the accuracy of keyword extraction.

[0019] In addition, classifying the extracted keywords through a clustering algorithm not only makes keyword management more orderly and systematic, but also helps users quickly understand and locate information on specific topics or categories, enhancing the accessibility and practicality of the information.

[0020] In addition, at least one preprocessing is performed on the original text information, including removing redundant characters, duplicate content, privacy information, violent content, etc., to exclude redundant or invalid content and reduce the computational workload. Based on the text quality scoring method of the open-source large language model, the text is comprehensively scored from dimensions such as content quality, logic, and information density, and high-quality texts are screened out by setting a scoring threshold to ensure the effectiveness of subsequent analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 Schematic diagram of a keyword classification method based on text information provided by an embodiment of the present application;

[0023] Figure 2 Flowchart of a keyword classification method based on text information provided by an embodiment of the present application;

[0024] Figure 3 Flowchart of another keyword classification method based on text information provided by an embodiment of the present application;

[0025] Figure 4 Flowchart of yet another keyword classification method based on text information provided by an embodiment of the present application;

[0026] Figure 5 Flowchart of yet another keyword classification method based on text information provided by an embodiment of the present application;

[0027] Figure 6 Block diagram of a keyword classification method provided by an embodiment of the present application;

[0028] Figure 7 Schematic diagram of the hardware structure of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the protection scope of the present application.

[0030] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0031] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] First, the technical field and related technical problems involved in the present application will be introduced.

[0033] The technical solution of the present application mainly involves feature extraction of text data or text information, including preprocessing of text data, keyword extraction and screening of text content, etc.

[0034] The current keyword extraction technologies are mainly divided into two categories: one is TF-IDF (Term Frequency Inverse Document Frequency). TF-IDF is a commonly used weighting technology in information retrieval and data mining, and it is also a statistical method used to evaluate the importance of a word in a document set or a corpus. TF (Term Frequency, word frequency) represents the frequency of a word appearing in a document. This number is usually normalized (generally, the word frequency is divided by the total number of words in the document) to prevent it from biasing towards long documents. IDF (Inverse Document Frequency) represents the rarity of a word in the entire document collection. IDF reduces the weight of words that appear frequently in many documents by calculating the number of documents containing a certain word. The calculation formula of IDF is: IDF(t) = log(N / df(t)), where N is the total number of documents in the corpus, and df(t) is the number of documents containing the word t. The TF-IDF technology: multiplies TF and IDF to obtain the weight value. The larger the TF-IDF value, the higher the importance of the word in the document.

[0035] TF-IDF evaluates the importance of words by calculating the term frequency of words in a text and their inverse document frequency in the entire corpus. However, TF-IDF cannot capture the context relationships and semantic information between words, resulting in one-sided keyword extraction, especially with poor performance when dealing with short texts or complex texts.

[0036] Another type of technology is TextRank (a graph-based ranking algorithm). The TextRank algorithm is an improved version based on the PageRank algorithm and is used to extract key information from texts. In the TextRank algorithm, a text is represented as a graph structure, where nodes can represent words or sentences, and edges represent the relationships between nodes. By iteratively calculating the weights of each node, the importance of the nodes in the text can be evaluated, thereby extracting keywords or generating summaries.

[0037] The core formula of the TextRank algorithm is similar to PageRank, using a damping factor to control the transfer of importance between nodes. Specifically, the weight of a node depends on the weights of the nodes connected to it and the weights of the edges between them. By continuously updating the weights of the nodes until convergence, the final ranking of the node weights can be obtained.

[0038] Although the TextRank algorithm has made great improvements in unsupervised keyword extraction compared to the TF-IDF algorithm, its graph construction method only considers the co-occurrence relationship of words, ignoring semantic similarity and context information, which limits the accuracy of the extraction results. Therefore, there are still significant deficiencies in the preprocessing, quality assessment, and keyword extraction of text data, especially when dealing with texts with complex semantics, lacking an efficient analysis method that can combine semantic information and context relevance. These defects are specifically manifested as follows:

[0039] Problems such as uneven quality of the original text data, lack of a scoring and screening mechanism for text quality, insufficient accuracy of keyword extraction, and low efficiency in the analysis and classification of large-scale data. Among them, the insufficient accuracy of keyword extraction is understood as: keyword extraction algorithms (such as TF-IDF and the original TextRank algorithm) often only utilize term frequency information or simple word co-occurrence relationships when dealing with texts, ignoring the semantic information and context relevance between words, resulting in keywords that are difficult to accurately reflect the core content of the text.

[0040] To solve the above technical problems, embodiments of the present application provide a keyword classification method based on text information. See Figure 1 , which is the overall flowchart of the keyword classification method provided in this embodiment. The method includes:

[0041] Obtain the original complex text data and / or text information as input data, and first perform the first-step preprocessing. This step is an optional step. During this process, redundant characters and privacy information in the original text information are removed, and duplicate content is also eliminated. The preprocessed text information enters the quality scoring and screening stage in the second step. During this process, an open-source large model scores the input text data in multiple dimensions according to the preset scoring criteria, and discards some data with a total score lower than the preset score threshold, so as to obtain high-quality text data after screening. In the third step, use an improved algorithm, such as an open-source large model, to calculate the embedding vector, apply the embedding vector to the process of calculating semantic similarity, extract at least one keyword from the text, and obtain the corresponding relationship between the keyword and a piece of text data. In the fourth step, input the at least one extracted keyword and its corresponding text data into the next step, classify the keywords using the clustering method, and finally output the keyword summary and classification result of the whole piece of data. For efficient and convenient use, the whole process can be encapsulated into a complete tool.

[0042] The following details each step in the method provided in this embodiment. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0043] In this embodiment, a keyword classification method based on text information is provided, which can be executed by computer devices such as servers, server clusters, or terminal devices. Figure 1 is a flowchart of the keyword classification method according to an embodiment of the present invention, as Figure 2 shown, and this process includes:

[0044] Step S101: Obtain the original text information, where the original text information includes at least one piece of text.

[0045] Among them, the original text information can be a piece of text or characters, or a text description composed of text and characters. The original text information can include one or more pieces of text. For example, one piece of text includes the following:

[0046] {

[0047] "instruction": "Three tips for maintaining health."

[0048] "output": " <think>\nTo answer this question, I need to think about three important tips for maintaining good health. These tips should be generally applicable and help people improve their health. I will consider aspects such as diet, exercise, and mental health.\n\n1. Diet: A healthy diet is the key to maintaining good health. It should include a variety of vegetables, fruits, whole grains, and protein sources.\n2. Exercise:......\n3. Pay attention to mental health: Keeping a positive mindset and learning to cope with stress are equally crucial for overall health.\n\n"。

[0049] }。

[0050] The above content is a piece of text. In addition, for example, the content of another piece of text includes:

[0051] {

[0052] "instruction": "What are the three primary colors?"

[0053] "output": " <think>The three primary colors refer to three basic colors in color theory that cannot be obtained by mixing other colors. In the field of pigments, the three primary colors are usually red, blue, and yellow, and they can be mixed to produce other colors. In the field of light, such as in televisions and computer monitors, the three primary colors are red, green, and blue, and the light of these three colors can be mixed to produce a wide range of colors. < / think> \n\n <answer>The three primary colors are red, blue, and yellow in the field of pigments and red, green, and blue in the field of light. < / answer> "。

[0054] }。

[0055] The above text can be obtained through software or captured in real time by the server. The acquisition objects include, but are not limited to, web pages, intermediate processing stages, and other methods. This embodiment does not limit this.

[0056] Step S102: Use a preset language model to score at least one piece of text, and perform preferential screening on at least one piece of text according to at least one scoring result to obtain one or more target texts.

[0057] Among them, the scoring criteria include at least one dimension: content quality, logic, and information volume; specifically, content quality is used to evaluate whether the required information content is included in the text, such as whether the text category meets the requirements and whether long or meaningless descriptions are avoided.

[0058] Logic is used to evaluate whether the content of the text has a clear logical structure and whether the sentences are smooth. Information volume is used to evaluate whether the number of information points included in the text reaches a preset number, that is, whether the text contains enough information points.

[0059] This step performs quality scoring and screening on at least one piece of text to ensure the data quality for subsequent analysis.

[0060] Optionally, the preset language model is an open-source language model, an advanced large language model (LLM), including but not limited to Qwen2.5-72B, DeepSeek-V3, etc. Among them, Qwen2.5-72B is a new generation of open-source large language model released by Tongyi Qianwen under Alibaba Cloud. The Qwen2.5-72B model has powerful multilingual capabilities and supports more than 29 languages such as Chinese, English, French, Spanish, etc. At the same time, it also supports multimodal models, such as the vision language model Qwen2-VL-72B, which can recognize pictures with different resolutions and aspect ratios and understand video content longer than 20 minutes.

[0061] DeepSeek-V3 is an LLM model released by an AI startup. It was officially released and open-sourced on December 26, 2024. DeepSeek-V3 has a wide range of application scenarios, including chat and coding scenarios, multilingual automatic translation, image generation, and AI painting. In multiple benchmark tests, DeepSeek-V3 performed excellently, especially showing outstanding performance in knowledge-based tasks, long text evaluation, and algorithmic code scenarios.

[0062] Step S103: Use the preset model to calculate at least one embedding vector corresponding to one or more target texts, and input the at least one embedding vector into a text processing algorithm for semantic similarity evaluation to obtain at least one keyword corresponding to the one or more target texts, and the corresponding relationship between each keyword and the target text where it is located.

[0063] Among them, the preset model can be an open-source model. For example, the preset model is a large language model LLM (such as Qwen2.5-72B, DeepSeek-V3, etc.). In this application, the preset model is used to calculate at least one embedding vector of the above one or more target texts. Then, the at least one embedding vector is input into a text processing algorithm for keyword extraction. Among them, the text processing algorithm can be an improved TextRank algorithm. The TextRank algorithm is an unsupervised keyword extraction algorithm, and its core idea comes from the PageRank algorithm. PageRank was originally used for web page ranking, and it evaluates the importance of nodes by calculating the link relationship between nodes. TextRank borrows this idea into natural language processing and realizes automatic keyword extraction by constructing a word relationship graph and calculating node weights.

[0064] It should be noted that the preset model in this step can be the same as the preset language model in the previous step S102, or it can also be different. This embodiment does not limit this.

[0065] To address the limitations of the traditional TextRank algorithm in practice, its main problem lies in that when constructing a word relationship graph, it often only considers the co-occurrence relationship of words, that is, nodes are connected based on the adjacent occurrences of words within a certain window range. This approach ignores the deep semantic relevance between words, resulting in the extracted keywords being potentially inaccurate and incomplete in some cases.

[0066] Based on the above problems, this embodiment proposes an improved method based on semantic similarity weights, introducing the ability of a language model to calculate the semantic association of words to optimize the construction process of the word relationship graph.

[0067] In this step, the introduction of embedding vectors enables the algorithm to utilize semantic information to depict the real connections between words. On this basis, using the improved TextRank algorithm, the cosine similarity between embedding vectors is used to calculate the semantic similarity between words and used as the weight of the edges in the graph.

[0068] This method can effectively enhance the expressive power of the word relationship graph, enabling it to not only reflect the surface relationships of words but also capture deep semantic associations. In addition, this improvement not only affects the construction process of the word relationship graph but also optimizes the subsequent calculation of node weights. In the traditional TextRank algorithm, the iterative calculation of node weights is mainly based on the co-occurrence frequency weights of edges, while in the implementation provided in this embodiment, the propagation and update process of node weights are also based on the semantic similarity of embedding vectors.

[0069] Step S104: Classify at least one keyword through a clustering algorithm according to at least one keyword and the corresponding relationship, generating one or more keyword summaries and classification results for one or more target texts.

[0070] After extracting keywords, combined with a clustering algorithm such as the K-Means algorithm, classify at least one target text, and finally output the keyword summary and classification results of the text data.

[0071] The keyword classification method provided in this embodiment can quickly identify and retain the most valuable or highly relevant texts by presetting a language model to score and preferentially screen the original text information, effectively reducing the amount of data for subsequent processing and improving the overall processing efficiency.

[0072] This method also calculates word embedding vectors using a preset model, calculates the semantic similarity between words through the embedding vectors, can more accurately capture the subtle semantic differences between texts, thereby extracting keywords highly relevant to the text content, and improving the accuracy and relevance of the keywords. The network structure of this method can capture the semantic relevance between words, rather than relying solely on word co-occurrence relationships, thus significantly improving the accuracy of keyword extraction.

[0073] In addition, classifying the extracted keywords through a clustering algorithm not only makes keyword management more orderly and systematic, but also helps users quickly understand and locate information on specific topics or categories, enhancing the accessibility and practicality of the information.

[0074] In a specific implementation manner of this embodiment, as Figure 3 shown, the above step S102, using a preset language model to score at least one piece of text, and preferentially screening at least one piece of text according to at least one scoring result to obtain one or more target texts, specifically includes:

[0075] Step S1021: Define a scoring criterion and a scoring threshold. The scoring criterion includes at least one dimension. For example, at least one dimension includes content quality, logic, and information quantity. The scoring threshold can be customized according to the actual situation. For example, the scoring threshold is set to 85 points, and the full score is a 100-point system.

[0076] Step S1022: According to the scoring criterion, use the preset language model to score at least one piece of text to obtain at least one scoring result, where each piece of text corresponds to one scoring result.

[0077] Among them, the preset language model can be the aforementioned large language model, such as Qwen2.5-72B, Deepseek-V3, etc.

[0078] Step S1023: Compare at least one scoring result with the scoring threshold respectively, and screen the texts whose scoring results are greater than or equal to the scoring threshold to obtain one or more target texts.

[0079] This embodiment proposes a text quality scoring method based on an open-source large language model, comprehensively scores the text from dimensions such as content quality, logic, and information density, and screens out high-quality texts by setting a scoring threshold to ensure the effectiveness of subsequent analysis.

[0080] This embodiment introduces a text preprocessing and quality scoring mechanism, effectively removes redundant characters, duplicate information, privacy data, and violent content, comprehensively evaluates and screens the text quality, and ensures that subsequent analysis is based on high-quality texts. This method can greatly reduce the interference of low-quality texts on the analysis results and improve the accuracy and reliability of the overall analysis.

[0081] In a specific implementation manner of this embodiment, as Figure 4 shown, the above step S103 specifically includes:

[0082] Step S1031: Encode the words in one or more target texts using a preset model to generate at least one context-related embedding vector, where each embedding vector includes a keyword and its corresponding vector value.

[0083] In this embodiment, the embedding vector includes a keyword and a vector value. One keyword can correspond to one or more vector values.

[0084] Step S1032: Input the keywords and vector values in at least one embedding vector into a text processing algorithm for semantic similarity evaluation to obtain a cosine similarity, and calculate the similarity between the words in the target text based on the cosine similarity.

[0085] Among them, the text processing algorithm can be an improved TextRank algorithm, which has two basic functions: 1. Keyword extraction, by analyzing the adjacent relationships between words to find keywords; 2. Automatic abstract generation, by calculating the similarity between sentences to generate an abstract of the text.

[0086] In this step, use the improved TextRank algorithm to calculate the semantic similarity between words using the cosine similarity between embedding vectors and use it as the weight of the edges in the graph.

[0087] Step S1033: Calculate the weight value corresponding to each keyword according to the similarity.

[0088] Step S1034: Sort the at least one weight value corresponding to at least one keyword in descending order, and extract the top N keywords to obtain the correspondence between the N keywords and the target text where they are located.

[0089] Among them, N≥1 and is a positive integer. For example, take the top 1 keyword, that is, the keyword with the largest weight value. N can be customized.

[0090] Furthermore, the text processing algorithm is a graph-based sorting algorithm, which is used to calculate the semantic similarity between words using the cosine similarity between embedding vectors and use it as the weight of the edges in the graph.

[0091] In another possible implementation manner of this embodiment, as Figure 5 shown, after obtaining the original text information in the above step S101, it further includes:

[0092] Step S101’: Perform at least one preprocessing on the original text information to generate processed text information; where the at least one preprocessing includes: removing redundant characters, deduplication, and privacy protection processing.

[0093] Specifically, removing redundant characters means: using regular expressions to clean irrelevant characters such as HTML tags and special symbols. Deduplication means: detecting duplicate content and deleting it. Privacy protection processing means: masking sensitive information such as ID numbers and phone numbers through regular expression matching.

[0094] Furthermore, at least one preprocessing is performed on the original text information to generate processed text information, including:

[0095] Using the Python programming language model to write regular expressions to match and delete tags and redundant characters for at least one piece of text; or, using the Python programming language model to detect and delete duplicate lines in the content of at least one piece of text; or, using the Python programming language model to write regular expressions to match and delete ID numbers, phone numbers, and sensitive information in a fixed format.

[0096] For the text data screened and scored, the improved TextRank algorithm is used to calculate and extract keywords therefrom. For the specific process, refer to step S103 of the foregoing embodiment, which will not be elaborated herein.

[0097] In this embodiment, the Python programming language model can be understood as a general term for the Python programming language and the methods and tools used in processing text information. As a powerful programming language, Python provides various functions and libraries when processing text information. Specifically, several preprocessing operations mentioned above - using regular expressions to match and delete tags and redundant characters, detecting and deleting duplicate lines, matching and deleting sensitive information such as ID numbers and phone numbers, etc., can all be implemented through the Python programming language and its related libraries (such as the re module for regular expressions, the set data structure for detecting duplicate lines, etc.).

[0098] Python programming language: As a tool for implementing text preprocessing tasks. It has the following functions:

[0099] Related libraries and modules: Such as the re module for regular expression matching, and other libraries that may be used for text processing (such as the string module for string operations, the os and io modules for file I / O, etc.).

[0100] Programming paradigms and algorithms: Including programming paradigms such as loops, conditional judgments, function calls, etc., and specific algorithms for processing text information (such as regular expression matching algorithms, deduplication algorithms, etc.).

[0101] In this embodiment, the Python programming language model includes but is not limited to the re module, the pandas module, etc.

[0102] The above step S102: using a preset language model to score at least one text, and selecting at least one text according to at least one scoring result to obtain one or more target texts, further includes:

[0103] Step S102 ′: score the processed text information using a preset language model, and select at least one text based on at least one scoring result to obtain one or more target texts.

[0104] The keyword classification method provided in this embodiment can quickly identify and retain the most valuable or relevant text by scoring and selecting the original text information through a preset language model, effectively reducing the amount of data to be processed later and improving the overall processing efficiency. In addition, by introducing semantic similarity calculation and optimizing the network structure of TextRank, the contextual relevance and semantic information between words can be fully considered.

[0105] This method also uses a preset model to calculate word embedding vectors. By calculating the semantic similarity between words through embedding vectors, it can more accurately capture the subtle semantic differences between texts, thereby extracting keywords that are highly relevant to the text content, improving the accuracy and relevance of keywords. The network structure of this method can capture the semantic relevance between words, rather than relying solely on word co-occurrence relationships, thereby significantly improving the accuracy of keyword extraction. Scoring screening based on preset models is suitable for the processing and analysis of large-scale text data. Unsupervised keyword extraction and automatic classification methods reduce manual intervention and significantly improve processing efficiency.

[0106] In addition, classifying the extracted keywords through clustering algorithms not only makes keyword management more orderly and systematic, but also helps users quickly understand and locate information on specific topics or categories, enhancing the accessibility and practicality of information.

[0107] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.

[0108] The method provided in this embodiment is described below with reference to a specific example:

[0109] Obtain two target texts. The above step S103 specifically includes: extracting keywords from the input sentences of the above two target texts, and using the improved TextRank to calculate the embedding vector, that is, the above step S1031, to obtain the embedding vector, for example, an embedding vector is a keyword and a vector value.

[0110] Step S1032: Calculate the semantic similarity between every two words using cosine similarity. For example, the calculated semantic similarities between two words are as follows: Vector A: [0.12, 0.34, 0.56,...]; Vector B: [0.23, 0.45, 0.67,...]; The cosine similarity between Vector A and Vector B is 0.92.

[0111] Step S1033: Determine the weight values corresponding to each keyword according to the above calculated cosine similarity. Specifically, assume that the initial weight of each keyword is 1, and use the cosine similarity calculated above to update the weights. The updated weight values of sentence 1 are: 1.25, 1.30, 1.28. The updated weight values of sentence 2 are: 1.22, 1.18, 1.15.

[0112] Step S1034: Extract at least one keyword. Sort the above weight values in descending order and extract the keyword corresponding to top1. Similarly, extract the top1 keyword for sentence 2, and so on.

[0113] Step S104: Perform text clustering processing on the at least one keyword and the corresponding sentence extracted above. In this embodiment, it is assumed that two keywords are extracted, so the clustering result may be a total of two clusters. If there is more input text data or more keywords, multiple clustering results can be output.

[0114] In this embodiment, a keyword classification device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0115] This embodiment provides a keyword classification device, as Figure 6 shown. This device includes: an acquisition module 610, a screening module 620, a processing module 630, and a clustering module 640. In addition, this device may also include other more or fewer modules, such as a storage module, a transceiver module, etc. This embodiment does not limit this.

[0116] Among them, the acquisition module 610 is used to acquire the original text information, and the original text information includes at least one piece of text.

[0117] The screening module 620 is used to score at least one piece of text using a preset language model, and perform optimal screening on at least one piece of text according to at least one scoring result to obtain one or more target texts.

[0118] A processing module 630 is configured to calculate at least one embedding vector corresponding to one or more target texts by using a preset model, and input the at least one embedding vector into a text processing algorithm for semantic similarity evaluation, so as to obtain at least one keyword corresponding to the one or more target texts, and the corresponding relationship between each keyword and the target text where it is located.

[0119] A clustering module 640 is configured to classify at least one keyword by using a clustering algorithm according to the at least one keyword and the corresponding relationship, and generate one or more keyword summaries and classification results of the one or more target texts.

[0120] Optionally, in a possible implementation manner of this embodiment, the screening module 620 is specifically configured to define a scoring criterion and a scoring threshold, where the scoring criterion includes at least one dimension; score at least one text by using a preset language model according to the scoring criterion to obtain at least one scoring result; compare the at least one scoring result with the scoring threshold respectively, and screen the texts whose scoring results are greater than or equal to the scoring threshold to obtain one or more target texts.

[0121] Among them, the scoring criterion includes at least one dimension: content quality, logic, and information quantity. Further, the content quality is used to evaluate whether the required information content is included in the text; the logic is used to evaluate whether the content of the text has a clear logical structure and whether the sentences are smooth; the information quantity is used to evaluate whether the number of information points included in the text reaches a preset quantity.

[0122] Optionally, in another possible implementation manner of this embodiment, the processing module 630 is specifically configured to encode the words in one or more target texts by using a preset model to generate at least one context-related embedding vector, where each embedding vector includes a keyword and a corresponding vector value; input the keywords and vector values in the at least one embedding vector into a text processing algorithm for semantic similarity evaluation to obtain a cosine similarity, and calculate the similarity between the words in the target text according to the cosine similarity; calculate the weight value corresponding to each keyword according to the similarity; sort the at least one weight value corresponding to the at least one keyword in descending order, and extract the top N keywords, where N≥1 and N is a positive integer, to obtain the corresponding relationship between the N keywords and the target text where they are located.

[0123] Among them, the text processing algorithm is a graph-based sorting algorithm, and the text processing algorithm is used to calculate the semantic similarity between words by using the cosine similarity between embedding vectors and use it as the weight of the edges in the graph.

[0124] Optionally, in another possible implementation of this embodiment, after obtaining the original text information, the processing module 630 is further configured to perform at least one preprocessing on the original text information to generate processed text information; use a preset language model to score the processed text information, and perform optimal selection and screening on at least one piece of text according to at least one scoring result to obtain one or more target texts.

[0125] Wherein, the at least one preprocessing includes: redundant character removal, duplicate removal, and privacy protection processing.

[0126] The processing module 630 is specifically configured to use a programming language python model to write a regular expression to match and delete tags and redundant characters for at least one piece of text; or, use a programming language python model to detect the content of at least one piece of text and delete duplicate lines; or, use a programming language python model to write a regular expression to match and delete ID card numbers, telephone numbers, and sensitive information in a fixed format.

[0127] For the description of the features in the corresponding embodiment of the keyword classification device provided in this embodiment, reference may be made to the relevant description of the corresponding embodiment of the method shown above, which will not be elaborated here one by one. Figures 1 to 5 shown in the relevant description of the corresponding embodiment of the method shown above, which will not be elaborated here one by one.

[0128] The embodiment of the present invention further provides a computer device having the above Figure 7 shown keyword classification device.

[0129] Please refer to Figure 7 , which is a schematic structural diagram of a computer device provided by an optional embodiment of the present application. The computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface).

[0130] In some optional implementations, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 7 Here, one processor 10 is taken as an example.

[0131] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.

[0132] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the keyword classification method shown in the above embodiments.

[0133] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can include a memory remotely arranged relative to the processor 10, and these remote memories can be connected to the computer device through a network.

[0134] The memory 20 can include a volatile memory, for example, a random access memory; the memory can also include a non-volatile memory, for example, a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.

[0135] The computer device further includes an input device and an output device. The processor 10, the memory 20, the input device, and the output device can be connected through a bus or other means. Among them, the input device can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device can include a display device, an auxiliary lighting device (for example, an LED), and a tactile feedback device (for example, a vibration motor), etc. The above-mentioned display device includes, but is not limited to, a liquid crystal display, a light-emitting diode, a display, and a plasma display.

[0136] In addition, the computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks. Among them, the number of communication interfaces 30 can be one or more, and this embodiment does not limit this.

[0137] Embodiments of the present application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described method embodiments of the keyword classification device based on text information when running.

[0138] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs that can store computer programs.

[0139] Embodiments of the present application also provide a computer program product, the computer program product including a computer program, and the steps in any of the above-described method embodiments of the keyword classification method based on text information are implemented when the computer program is executed by a processor.

[0140] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and the steps in any of the above-described method embodiments of the keyword classification method based on text information are implemented when the computer program is executed by a processor.

[0141] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0142] The above has introduced in detail a keyword classification method based on text information provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.< / think>

Claims

1. A keyword classification method based on text information, characterized in that: The method comprises: Acquire original text information, wherein the original text information includes at least one text; Scoring the at least one text using a preset language model, and selectively selecting the at least one text according to at least one scoring result to obtain one or more target texts; Calculate at least one embedding vector corresponding to the one or more target texts using a preset model, and input the at least one embedding vector into a text processing algorithm for semantic similarity evaluation to obtain at least one keyword corresponding to the one or more target texts, and a corresponding relationship between each keyword and the target text in which it is located; According to the at least one keyword and the corresponding relationship, the at least one keyword is classified by a clustering algorithm to generate one or more keyword summaries and classification results for the one or more target texts.

2. The keyword classification method according to claim 1, characterized in that: The step of scoring the at least one text by using a preset language model, and selectively selecting the at least one text according to at least one scoring result to obtain one or more target texts includes: defining a scoring criterion and a scoring threshold, wherein the scoring criterion includes at least one dimension; According to the scoring criteria, scoring the at least one text using the preset language model to obtain at least one scoring result, wherein each text corresponds to one scoring result; The at least one scoring result is compared with the scoring threshold respectively, and texts having scoring results greater than or equal to the scoring threshold are screened to obtain the one or more target texts.

3. The keyword classification method according to claim 2, characterized in that: The scoring criteria include at least one dimension: content quality, logic and information content; Among them, the content quality is used to evaluate whether the text contains the required information content; the logic is used to evaluate whether the content of the text has a clear logical structure and whether the sentences are fluent; the information quantity is used to evaluate whether the number of information points contained in the text reaches a preset number.

4. The keyword classification method according to claim 1, characterized in that: Using a preset model to calculate at least one embedding vector corresponding to the one or more target texts, and inputting the at least one embedding vector into a text processing algorithm for semantic similarity evaluation, to obtain at least one keyword corresponding to the one or more target texts, and a corresponding relationship between each keyword and the target text in which it is located, including: Encoding the words in the one or more target texts using the preset model to generate at least one context-related embedding vector, each of the embedding vectors including a keyword and a corresponding vector value; Inputting the keyword and the vector value in the at least one embedding vector into the text processing algorithm to perform semantic similarity evaluation to obtain cosine similarity, and calculating the similarity between words in the target text according to the cosine similarity; Calculate the weight value corresponding to each keyword according to the similarity; At least one weight value corresponding to at least one keyword is sorted in descending order, and the top N keywords are extracted to obtain the corresponding relationship between the N keywords and the target texts where they are located, where N≥1 and is a positive integer.

5. The keyword classification method according to claim 4, characterized in that: The text processing algorithm is a graph-based sorting algorithm, which is used to calculate the semantic similarity between words using the cosine similarity between embedded vectors and use it as the weight of the edge in the graph.

6. The keyword classification method according to any one of claims 1 to 5, characterized in that: After obtaining the original text information, it also includes: Performing at least one preprocessing on the original text information to generate processed text information; wherein the at least one preprocessing includes: redundant character removal, deduplication and privacy protection processing; The step of scoring the at least one text by using a preset language model, and selectively selecting the at least one text according to at least one scoring result to obtain one or more target texts includes: The processed text information is scored using the preset language model, and the at least one text is preferentially screened according to at least one scoring result to obtain one or more target texts.

7. The keyword classification method according to claim 6, characterized in that: The performing at least one preprocessing on the original text information to generate processed text information includes: Use the programming language Python model to write a regular expression to match the at least one text and delete tags and redundant characters; or, Use the programming language Python model to detect and delete duplicate lines from the content of the at least one text; or, Use the programming language Python model to write a regular expression to match and delete ID card numbers, phone numbers, and sensitive information in a fixed format.

8. A keyword classification device based on text information, characterized in that: The device comprises: An acquisition module, used to acquire original text information, wherein the original text information includes at least one text; A screening module, used to score the at least one text using a preset language model, and to preferentially screen the at least one text according to at least one scoring result to obtain one or more target texts; A processing module, configured to calculate at least one embedding vector corresponding to the one or more target texts using a preset model, and input the at least one embedding vector into a text processing algorithm for semantic similarity evaluation, thereby obtaining at least one keyword corresponding to the one or more target texts, and a corresponding relationship between each keyword and the target text in which it is located; The clustering module is used to classify the at least one keyword by a clustering algorithm according to the at least one keyword and the corresponding relationship, and generate one or more keyword summaries and classification results for the one or more target texts.

9. A computer device, characterized in that: comprising a memory and a processor, wherein: The memory is used to store computer programs; The processor is used to implement the steps of the keyword classification method based on text information as claimed in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer readable storage medium stores a computer program, wherein: When the computer program is executed by a processor, the steps of the keyword classification method based on text information are implemented as claimed in any one of claims 1 to 7.