Subject processing method and device, electronic equipment and storage medium
By using the term frequency-inverse document frequency algorithm and dynamic topic model in text processing, the problem of inaccurate text topic analysis results is solved, and dynamic evolution analysis of text topics over time is realized, thereby improving the accuracy of the analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AGRICULTURAL BANK OF CHINA
- Filing Date
- 2022-11-02
- Publication Date
- 2026-07-21
AI Technical Summary
Existing text topic analysis methods lack dynamic topic evolution analysis in the time dimension, resulting in poor accuracy of analysis results.
By acquiring multiple texts to be processed within a first preset time period, preprocessing them, extracting text keywords using a term frequency-inverse document frequency algorithm, and determining the text topic based on a dynamic topic model, the evolution trend of the text topic within a second preset time period is displayed.
It improves the accuracy of text topic analysis, enables dynamic evolution analysis of text topics over time, and reflects the dynamic evolution trend of topics.
Smart Images

Figure CN116090438B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer application technology, and in particular to a subject processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] Topic mining technology is often used in personalized recommendation, search and anomaly detection. At present, topic mining research is mainly focused on social networking sites and medical communities. For text topic extraction, statistical and content analysis methods are usually used, but there is a lack of research on dynamic topic evolution analysis of text topics from a time dimension.
[0003] Therefore, common methods for analyzing text topics are rather crude, resulting in poor accuracy in the analysis results. Summary of the Invention
[0004] This invention provides a topic processing method, apparatus, electronic device, and storage medium to solve the technical problem of poor accuracy in text topic analysis results.
[0005] According to one aspect of the present invention, a topic processing method is provided, wherein the method includes:
[0006] Obtain multiple first texts to be processed within a first preset time period, preprocess each first text to obtain a second text corresponding to the first text;
[0007] For each of the second texts, the text keywords of the second text are extracted using the term frequency-inverse document frequency algorithm;
[0008] The text topic corresponding to the second text is determined based on multiple text keywords of the second text and a pre-built dynamic topic model;
[0009] The text topics within a second preset time period are obtained and displayed in a preset manner, wherein the second preset time period is longer than the first preset time period.
[0010] According to another aspect of the present invention, a subject processing apparatus is provided, wherein the apparatus comprises:
[0011] The text acquisition module is used to acquire multiple first texts to be processed within a first preset time period, and to preprocess each first text to obtain a second text corresponding to the first text.
[0012] The keyword extraction module is used to extract text keywords from each piece of the second text using a term frequency-inverse document frequency algorithm.
[0013] The text topic determination module is used to determine the text topic corresponding to the second text based on multiple text keywords of the second text and a pre-built dynamic topic model;
[0014] The text topic display module is used to obtain text topics within a second preset time period and display the obtained text topics in a preset manner, wherein the second preset time period is longer than the first preset time period.
[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the subject processing method according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the subject processing method described in any embodiment of the present invention.
[0020] The technical solution of this invention involves acquiring multiple first texts to be processed within a first preset time period, preprocessing each first text to obtain a second text corresponding to the first text, and acquiring the data-cleaned second text. For each second text, a term frequency-inverse document frequency algorithm is used to extract text keywords, making the extracted text keywords more accurate. Based on the text keywords of multiple second texts and a pre-constructed dynamic topic model, the text topic corresponding to the second text is determined, improving the accuracy of the text topic analysis results. The text topic within a second preset time period is acquired, and the acquired text topic is displayed in a preset manner. The second preset time period is longer than the first preset time period, and dynamic topic evolution analysis is performed from a time dimension. By dynamically extracting text topics from different time periods, the dynamic evolution of text topics in the time dimension is reflected.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of a topic processing method provided in Embodiment 1 of the present invention;
[0024] Figure 2 This is a flowchart of a topic processing method provided according to Embodiment 2 of the present invention;
[0025] Figure 3 This is a flowchart of a topic processing method provided according to Embodiment 3 of the present invention;
[0026] Figure 4 This is a schematic diagram of the structure of a topic processing device according to Embodiment 4 of the present invention;
[0027] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the topic processing method of the present invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] Example 1
[0031] Figure 1The present invention provides a flowchart of a topic processing method according to Embodiment 1. This embodiment is applicable to data processing situations. The method can be executed by a topic processing device, which can be implemented in hardware and / or software and can be configured in a computer.
[0032] like Figure 1 As shown, the method includes:
[0033] S110. Obtain multiple first texts to be processed within a first preset time period, and preprocess each first text to obtain a second text corresponding to the first text.
[0034] The first preset time period can be understood as the time period used to acquire the first text. In this embodiment of the invention, the first preset time period can be preset according to scenario requirements, and is not specifically limited here. Optionally, the first preset time period can be a specific time period on the same day, or a specific time period spanning multiple days. For example, the first preset time period can be 4 hours, 12 hours, 24 hours, or 48 hours, etc.
[0035] Here, the first text can be understood as the text to be processed for a specific topic. Optionally, the first text may be text related to the financial field. In this embodiment of the invention, the first text can be obtained according to the needs of the scenario, and no specific limitation is made here. For example, the first text may be text obtained from the comment section of a financial community. Specifically, a combination of data scraping tools can be used, that is, a distributed crawler written based on Python's Requests + BeautifulSoup library can be used to scrape data from the comment section of the financial community as the first text.
[0036] The second text can be understood as the text obtained by preprocessing the first text. Optionally, the preprocessing includes, but is not limited to, at least one of the following: lowercase conversion, punctuation removal, stop word removal, Uniform Resource Locator (URL) removal, and spell correction. Specifically, each piece of the first text can be preprocessed with lowercase conversion, punctuation removal, stop word removal, URL removal, and spell correction to obtain the second text corresponding to the first text.
[0037] S120. For each piece of the second text, extract the text keywords of the second text using the term frequency-inverse document frequency algorithm.
[0038] The term frequency–inverse document frequency (TF-IDF) algorithm is an algorithm that can extract text keywords from the second text. The text keywords can be understood as words that represent the core ideas and content of the second text.
[0039] Specifically, the term frequency-inverse document frequency (IDF) algorithm is a keyword extraction technique based on weighted statistics. Term frequency (TF) can be understood as the number of times a word appears in the second text, representing its importance; inverse document frequency (IDF) eliminates unimportant, common words. In this embodiment of the invention, using the term frequency-inverse document frequency algorithm to extract text keywords from the second text can achieve the goal of reducing model iteration time and improving model accuracy.
[0040] S130. Determine the text topic corresponding to the second text based on multiple text keywords of the second text and a pre-built dynamic topic model.
[0041] The Dynamic Topic Models (DTM) can be understood as models that can automatically process corpora containing time attributes dynamically, and extract topics and mine the covariance network and evolution trend between topics and keywords based on dynamic topic recognition by introducing the influence of the previous time slice on the next time slice.
[0042] The text topic can be understood as representing key content in multiple second texts, i.e., topic information.
[0043] S140. Obtain the text topics within the second preset time period and display the obtained text topics in a preset manner.
[0044] The second preset time period can be longer than the first preset time period. The second preset time period may include one or more of the first preset time periods. In this embodiment of the invention, the second preset time period can be preset according to scenario requirements, and is not specifically limited here.
[0045] The preset method can be understood as a way of displaying the evolution trend of the text topic over time. In this embodiment of the invention, the preset method can be preset according to the needs of the scenario, and is not specifically limited here. Optionally, pyecharts can be used to display the evolution trend of the acquired text topic. pyecharts can be understood as a library for generating charts. Furthermore, dynamic graphs, change curves, or aesthetically pleasing online reports of the evolution trend of the text topic can be displayed based on pyecharts. This solves the problem that common methods for analyzing text topics lack the influence of previous time slices on subsequent time slices, and realizes the research and analysis of the evolution trend of text topics from a time dimension.
[0046] The technical solution of this invention involves acquiring multiple first texts to be processed within a first preset time period, preprocessing each first text to obtain a second text corresponding to the first text, and acquiring the data-cleaned second text. For each second text, a term frequency-inverse document frequency algorithm is used to extract text keywords, making the extracted text keywords more accurate. Based on the text keywords of multiple second texts and a pre-constructed dynamic topic model, the text topic corresponding to the second text is determined, improving the accuracy of the text topic analysis results. The text topic within a second preset time period is acquired, and the acquired text topic is displayed in a preset manner. The second preset time period is longer than the first preset time period, and dynamic topic evolution analysis is performed from a time dimension. By dynamically extracting text topics from different time periods, the dynamic evolution of text topics in the time dimension is reflected.
[0047] Example 2
[0048] Figure 2 This is a flowchart of a topic processing method provided in Embodiment 2 of the present invention. This embodiment appends text keywords extracted from the second text using the term frequency-inverse document frequency algorithm described in the above embodiments. Figure 2 As shown, the method includes:
[0049] S210. Obtain multiple first texts to be processed within a first preset time period, and preprocess each first text to obtain a second text corresponding to the first text.
[0050] S220. For each piece of the second text, determine the text length of the second text data, and determine the word frequency-reverse file frequency algorithm based on the text length.
[0051] The text length can be understood as the length of the second text data. It should be understood that the determination of the text length is related to the bus width of the specific system. For example, in a 32-bit system, the text length of one word can be 4 bytes.
[0052] Furthermore, it's important to understand that the traditional TF-IDF algorithm only considers the computational results of feature terms on the sample set. In this embodiment of the invention, the first text obtained from the comment section of a financial community is typically short, resulting in sparse TF calculations. Therefore, the TF-IDF algorithm is improved by reducing the weight of the TF calculation values and taking the logarithm, thus preventing IDF from excessively reducing the importance of word frequencies.
[0053] Optionally, the algorithm for determining the term frequency-inverse document frequency based on the text length includes:
[0054] If the text length does not exceed a preset length, the term frequency-inverse document frequency (TF-IDF) is calculated based on the following formula:
[0055]
[0056] Wherein, TF is the term frequency of the text keyword, and IDF is the inverse document frequency of the text keyword.
[0057] The preset length can be understood as the text length used to determine the term frequency-inverse document frequency (TF-IDF) algorithm for the second text. In this embodiment of the invention, the preset length can be preset according to scenario requirements and is not specifically limited here. Optionally, the preset length can be 8 bytes or 12 bytes, etc. Furthermore, the TF-IDF can be calculated based on different formulas for the second text of different lengths.
[0058] Optionally, the algorithm for determining the term frequency-inverse document frequency based on the text length includes:
[0059] If the text length exceeds a preset length, the term frequency-inverse document frequency (TF-IDF) is calculated based on the following formula:
[0060] TF-IDF = TF * log(IDF) + 1
[0061] Wherein, TF is the term frequency of the text keyword, and IDF is the inverse document frequency of the text keyword.
[0062] Specifically, for each piece of the second text, the text length of the second text data is determined; if the text length does not exceed a preset length, based on... The term frequency-inverse document frequency (TF-IDF) is calculated. If the text length exceeds a preset length, the TF-IDF is calculated based on TF-IDF = TF * log(IDF) + 1. This improves the accuracy of extracting text keywords from the second text using the TF-IDF.
[0063] S230. Use the term frequency-inverse document frequency algorithm to extract the text keywords of the second text.
[0064] S240. Determine the text topic corresponding to the second text based on multiple text keywords of the second text and a pre-built dynamic topic model.
[0065] S250. Obtain the text topic within a second preset time period and display the obtained text topic in a preset manner, wherein the second preset time period is longer than the first preset time period.
[0066] The technical solution of this invention determines the text length of the second text data, determines the term frequency-inverse document frequency (TF-IDF) algorithm based on the text length, and calculates the TF-IDF using different calculation formulas under different conditions. This improves the accuracy of extracting text keywords from the second text using the TF-IDF algorithm.
[0067] Example 3
[0068] Figure 3 This is a flowchart of a topic processing method provided in Embodiment 3 of the present invention. This embodiment refines the text topic determined based on multiple text keywords of the second text and a pre-constructed dynamic topic model, as described in the above embodiments. For example... Figure 3 As shown, the method includes:
[0069] S310. Obtain multiple first texts to be processed within a first preset time period, and preprocess each first text to obtain a second text corresponding to the first text.
[0070] S320. For each of the second texts, extract the text keywords of the second text using the term frequency-inverse document frequency algorithm.
[0071] S330. Construct a bag-of-words model based on multiple text keywords of the second text to obtain text vectors corresponding to the second text.
[0072] The bag-of-words model can be understood as a model that can vectorize the second text. In this embodiment of the invention, the bag-of-words model can perform word frequency statistics on all words in multiple pieces of the second text, transforming the second text into a set of word frequencies ordered according to word order.
[0073] Specifically, firstly, multiple second texts are scanned to generate a dictionary index containing words, and the frequency of each word appearing in multiple second texts is counted; further, each second text can be represented in the form of [[word frequency 1], [word frequency 2], [word frequency 3]]; further still, the second text is vectorized using the bag-of-words model package dic2bow in Python to obtain the text vector corresponding to the second text.
[0074] The text vector can be understood as the vector obtained by vectorizing the second text through the bag-of-words model.
[0075] S340. The text vector is processed by a pre-constructed dynamic topic model to obtain the text topic corresponding to the second text.
[0076] Optionally, before processing the text vector using a pre-built dynamic topic model, the method further includes: determining the optimal number of topics for the dynamic topic model based on the topic coherence of the text topics determined by the pre-built dynamic topic model.
[0077] The topic coherence can be understood as measuring the frequency of the text topic in the second text. Specifically, the topic coherence of the text topics can be determined by calculating the average distance between the first K text topics.
[0078] The optimal number of topics can be understood as the number of topics that allow the dynamic temporal topic model to reach its optimal state. In this embodiment of the invention, the dynamic temporal topic model can be evaluated based on the topic coherence.
[0079] In this embodiment of the invention, the calculation of topic coherence can take various forms. Optionally, the topic processing method further includes:
[0080] For each text topic determined by a pre-built dynamic topic model, a word pair is formed between the text topic and its preceding keyword in the second text. The topic coherence of the text topic is calculated based on the log-conditional probability of the word pair; or,
[0081] The text topics determined by the pre-constructed dynamic topic model are segmented using a sliding window to obtain target word pairs. The topic coherence of the text topics is determined based on the normalized point-state mutual information and cosine similarity of the target word pairs.
[0082] For example, specifically, for each text topic determined by a pre-built dynamic topic model, the text topic is paired with its preceding text keyword in the second text to form a word pair. The topic coherence of the text topic is calculated based on the log-conditional probability of the word pair, i.e., the topic coherence of the text topic is calculated using the Umass calculation method. The calculation formula may be:
[0083]
[0084] Where N represents the number of text topics, i represents the time slice identifier, and c Umass w represents the log-conditional probability. i and w j This represents the co-occurring words in the second text, and ε represents the parameter to avoid occurrences of 0.
[0085] For example, specifically, text topics determined by a pre-built dynamic topic model are segmented using a sliding window to obtain target word pairs. The topic coherence of the text topics is determined based on the normalized point-state mutual information and cosine similarity of the target word pairs; that is, the topic coherence of the text topics is calculated using a CV (cosine similarity) calculation method. The calculation formula can be:
[0086]
[0087] in, This represents the target word pair, where i represents the time slice identifier, and w represents the target word pair. i and w j Indicates co-occurring words in the second text.
[0088]
[0089] Among them, NPMI(w i , xj ) Y Represents normalized point-state mutual information, i represents the time slice identifier, and w represents the normalized point-state mutual information. i and w j This represents the co-occurring words in the second text, and ε represents the parameter to avoid occurrences of 0.
[0090]
[0091] in, Represents cosine similarity, w i This indicates co-occurring words in the second text, and 'i' represents the time slice identifier.
[0092] S350. Obtain the text topic within the second preset time period and display the obtained text topic in a preset manner, wherein the second preset time period is longer than the first preset time period.
[0093] The technical solution of this invention involves constructing a bag-of-words model based on multiple text keywords of the second text to obtain text vectors corresponding to the second text; then, processing the text vectors using a pre-constructed dynamic topic model to obtain text topics corresponding to the second text. Based on dynamic topic recognition, topics are extracted and the covariance network and evolution trend between topics and keywords are mined, enabling analysis of dynamic topic evolution from a time dimension.
[0094] Optionally, in this embodiment of the invention, the overall process of the information recognition method may be:
[0095] 1. Obtain the first text.
[0096] We used a combination of common data scraping tools, and wrote a distributed crawler based on Python's Requests + BeautifulSoup library to scrape multiple financial comment texts within a first preset time period. We then performed data cleaning on the scraped data.
[0097] 2. Word segmentation.
[0098] It is understandable that the results of word segmentation may affect the results of subsequent topic extraction and topic recognition. In this embodiment of the invention, the Chinese word segmentation toolkit Jieba is used to segment words in the corpus, and a custom dictionary and stop word list are built on this basis to improve the accuracy of word segmentation.
[0099] 3. Extract keywords.
[0100] It is important to understand that the traditional TF-IDF algorithm only considers the calculation results of feature terms on the sample set. TF can represent the importance of words, while IDF can remove unimportant common words. Since Chinese financial commentary texts are generally short, the results of TF calculation are often sparse. Therefore, this invention improves TF-IDF by reducing the weight of TF calculation values and by taking the logarithm to prevent IDF from excessively reducing the importance of word frequency.
[0101] The traditional TF-IDF algorithm's calculation formula is:
[0102]
[0103] The improved TF-IDF algorithm is calculated as follows:
[0104]
[0105] 4. Text topic identification and evolution based on dynamic topic model.
[0106] The dynamic temporal topic model is constructed using gensim in Python. The principle of the dynamic temporal topic model is as follows.
[0107] The model is essentially composed of a set of time-series independent linear discriminant analysis (LDA) topic models. Under different time windows, both the Dirichlet distribution and the Dirichlet distribution evolve over time.
[0108] Specifically, the algorithm can be divided into four main steps:
[0109] (1) Constructing a bag-of-words model. The bag-of-words model can be understood as the process of performing word frequency statistics on all words in the entire corpus, and packaging the documents into a set of word frequencies ordered by word order. Specifically, firstly, the corpus is scanned to generate a dictionary index containing words, and the frequency of each word in the entire corpus is counted. Further, each text is represented in the form of [[word frequency 1], [word frequency 2], [word frequency 3]]. The Chinese financial commentary text data of this invention is vectorized using the dic2bow bag-of-words model package in Python.
[0110] (2) Determine the topic distribution of the text. The topic distribution of the text is calculated using a static LDA model. The entire static LDA model modeling process is no different from that of common LDA models. The static LDA model parameter estimation is used to initialize the dynamic time series topic model. During the initialization process, the observed variance is used to approximate the true variance and the frequency variance, and the evolution over time is determined by the Gaussian parameters defined in the distribution.
[0111] (3) The Expectation Maximization Algorithm (EM) is used to iterate the model. The EM algorithm is an iterative optimization strategy. Each iteration consists of two steps: an expectation step (E-step), which is a clustering process, and a maximization step (M-step), which estimates the maximum likelihood probability. Specifically, in the E-step, the topic-word probability distribution in the LDA model is updated iteratively for each time slice. That is, the topic content of time slice i+1 is iteratively formed based on the LDA model of time slice i. Typically, true posterior inference is difficult to handle, so an appropriate lower bound for iteration needs to be set. In the algorithm, the KL probability is minimized by...
[0112] (Kullback-Liebler Divergence) is used to optimize this bound. After training the corpus with a dynamic topic model, the probability distribution of each document on each topic, as well as the topic-vocabulary probability distribution for each time period, can be obtained.
[0113] (4) Determining the Optimal Number of Topics. To obtain the optimal topic model, the dynamic temporal topic model needs to select the optimal number of topics. There are many methods for evaluating the number of topics in a topic model. This invention selects the consistency evaluation index—topic coherence—to evaluate the topic model. Topic coherence is mainly used to measure the coherence of words within a topic. It evaluates the quality of a model by measuring the frequency of topic words in the corpus and calculating the average distance between the first K words. There are many ways to calculate topic coherence. The main methods used are c_v and Umass. Umass mainly calculates topic coherence by pairing a word with its preceding word for a single word in the segmented document and calculating the log-conditional probability of the word pair. The c_v calculation method is based on a sliding window, performing one-set segmentation on topic words (any two words in a set form a word pair), calculating the normalized point mutual information (NPMI), and finally calculating the cosine similarity as the coherence evaluation index.
[0114] 5. Display the text theme.
[0115] By extracting the maximum probability of topic terms from the optimal model, we use Python-based Pyecharts to write code for visual analysis of the data results, visually presenting the hot topics in financial commentary and the evolution trend of these topics over time.
[0116] This invention proposes a topic processing method, filling a gap in the research on topic mining methods for Chinese financial communities and providing valuable reference for subsequent research on Chinese financial topic mining. By using a dynamic topic model, the influence of the previous time slice on topic extraction in the next time slice is introduced. Based on the time dimension, potential topics and evolution trends are identified, proposing an effective method to improve the accuracy and precision of topic identification and achieve real-time topic evolution.
[0117] Example 4
[0118] Figure 4 This is a schematic diagram of a subject processing device provided in Embodiment 4 of the present invention. Figure 4 As shown, the device includes: a text acquisition module 410, a keyword extraction module 420, a text topic determination module 430, and a text topic display module 440.
[0119] The system includes a text acquisition module 410, which acquires multiple first texts to be processed within a first preset time period, preprocesses each first text, and obtains a second text corresponding to the first text; a keyword extraction module 420, which extracts text keywords from each second text using a word frequency-inverse document frequency algorithm; a text topic determination module 430, which determines the text topic corresponding to the second text based on the text keywords of multiple second texts and a pre-built dynamic topic model; and a text topic display module 440, which acquires text topics within a second preset time period and displays the acquired text topics in a preset manner, wherein the second preset time period is longer than the first preset time period.
[0120] The technical solution of this invention involves acquiring multiple first texts to be processed within a first preset time period, preprocessing each first text to obtain a second text corresponding to the first text, and acquiring the cleaned text. For each second text, the word frequency-inverse document frequency algorithm is used to extract text keywords, reducing model iteration time and improving model accuracy. Based on the text keywords of multiple second texts and a pre-constructed dynamic topic model, the text topic corresponding to the second text is determined. Based on dynamic topic recognition, the topic is extracted, and the covariance network and evolution trend between the topic and keywords are mined. The text topics within a second preset time period are acquired and displayed in a preset manner. The second preset time period is longer than the first preset time period, allowing for dynamic topic evolution analysis from a time dimension. This improves the accuracy of the text topic analysis results.
[0121] Optionally, the topic processing device further includes an algorithm determination module.
[0122] The algorithm determination module is used to determine the text length of the second text data before extracting the text keywords of the second text data using the term frequency-inverse file frequency algorithm, and to determine the term frequency-inverse file frequency algorithm based on the text length.
[0123] Optionally, the algorithm determination module is used for:
[0124] If the text length does not exceed a preset length, the term frequency-inverse document frequency (TF-IDF) is calculated based on the following formula:
[0125]
[0126] Wherein, TF is the term frequency of the text keyword, and IDF is the inverse document frequency of the text keyword.
[0127] Optionally, the algorithm determination module is used for:
[0128] If the text length exceeds a preset length, the term frequency-inverse document frequency (TF-IDF) is calculated based on the following formula:
[0129] TF-IDF = TF * log(IDF) + 1
[0130] Wherein, TF is the term frequency of the text keyword, and IDF is the inverse document frequency of the text keyword.
[0131] Optionally, the text topic determination module 430 includes: a text vector determination submodule and a text vector processing submodule.
[0132] The text vector determination submodule is used to construct a bag-of-words model based on multiple text keywords of the second text to obtain the text vector corresponding to the second text.
[0133] The text vector processing submodule is used to process the text vector through a pre-built dynamic topic model to obtain the text topic corresponding to the second text.
[0134] Optionally, the text topic determination module 430 may also include a topic number determination module.
[0135] The topic number determination module is used to determine the optimal number of topics in the dynamic topic model by using the topic coherence of the text topics determined by the pre-built dynamic topic model before processing the text vector through the pre-built dynamic topic model.
[0136] Optionally, the topic number determination module is further configured to:
[0137] For each text topic determined by a pre-built dynamic topic model, a word pair is formed between the text topic and its preceding keyword in the second text. The topic coherence of the text topic is calculated based on the log-conditional probability of the word pair; or,
[0138] The text topics determined by the pre-constructed dynamic topic model are segmented using a sliding window to obtain target word pairs. The topic coherence of the text topics is determined based on the normalized point-state mutual information and cosine similarity of the target word pairs.
[0139] The topic processing apparatus provided in the embodiments of the present invention can execute the topic processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0140] Example 5
[0141] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0142] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0143] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0144] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as topic processing methods.
[0145] In some embodiments, the topic processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the topic processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the topic processing method by any other suitable means (e.g., by means of firmware).
[0146] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0147] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0148] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0150] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0151] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0152] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0153] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A topic processing method, characterized in that, include: Obtain multiple first texts to be processed within a first preset time period, preprocess each first text to obtain a second text corresponding to the first text; For each of the second texts, the text keywords of the second text are extracted using the term frequency-inverse document frequency algorithm; The text topic corresponding to the second text is determined based on multiple text keywords of the second text and a pre-built dynamic topic model; Obtain the text topics within a second preset time period and display the obtained text topics in a preset manner, wherein the second preset time period is longer than the first preset time period; Prior to extracting the text keywords of the second text using the term frequency-inverse document frequency algorithm, the method further includes: The text length of the second text is determined, and a term frequency-reverse file frequency algorithm is determined based on the text length; wherein, the text length of the second text is related to the bus width of the system to which it belongs; The algorithm for determining term frequency-inverse document frequency based on the text length includes: If the text length does not exceed a preset length, the term frequency-reverse document frequency is calculated based on the following formula. : ; in, The word frequency of the keywords in the text. The inverse file frequency of the text keywords; If the text length exceeds a preset length, the term frequency-reverse document frequency is calculated based on the following formula. : ; in, The word frequency of the keywords in the text. The inverse file frequency of the text keywords.
2. The method according to claim 1, characterized in that, The determination of the text topic corresponding to the second text based on multiple text keywords of the second text and a pre-built dynamic topic model includes: A bag-of-words model is constructed based on multiple text keywords of the second text to obtain text vectors corresponding to the second text; The text vector is processed by a pre-built dynamic topic model to obtain the text topic corresponding to the second text.
3. The method according to claim 2, characterized in that, Before processing the text vector using a pre-built dynamic topic model, the method further includes: The optimal number of topics for the dynamic topic model is determined by the topic coherence of the text topics identified through a pre-built dynamic topic model.
4. The method according to claim 3, characterized in that, Also includes: For each text topic determined by a pre-built dynamic topic model, a word pair is formed between the text topic and its preceding keyword in the second text. The topic coherence of the text topic is calculated based on the log-conditional probability of the word pair; or, The text topics determined by the pre-constructed dynamic topic model are segmented using a sliding window to obtain target word pairs. The topic coherence of the text topics is determined based on the normalized point-state mutual information and cosine similarity of the target word pairs.
5. A subject processing apparatus, characterized in that, include: The text acquisition module is used to acquire multiple first texts to be processed within a first preset time period, and to preprocess each first text to obtain a second text corresponding to the first text. The keyword extraction module is used to extract text keywords from each piece of the second text using a term frequency-inverse document frequency algorithm. The text topic determination module is used to determine the text topic corresponding to the second text based on multiple text keywords of the second text and a pre-built dynamic topic model; The text topic display module is used to obtain text topics within a second preset time period and display the obtained text topics in a preset manner, wherein the second preset time period is longer than the first preset time period; The topic processing device further includes an algorithm determination module, which is used to determine the text length of the second text before extracting the text keywords of the second text using the term frequency-reverse document frequency algorithm, and to determine the term frequency-reverse document frequency algorithm based on the text length; wherein the text length of the second text is related to the bus width of the system to which it belongs; The algorithm determination module is configured to: calculate the term frequency-reverse document frequency based on the following formula, provided that the text length does not exceed a preset length. : ; in, The word frequency of the keywords in the text. The inverse file frequency of the text keywords; If the text length exceeds a preset length, the term frequency-reverse document frequency is calculated based on the following formula. : ; in, The word frequency of the keywords in the text. The inverse file frequency of the text keywords.
6. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the subject processing method according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the subject processing method according to any one of claims 1-4.