Article popularity calculation method and device, terminal and storage medium

By introducing keyword weight and heat decay models into the article heat calculation method, the problem of inability to fully reflect the influence and heat changes in the existing technology is solved, and more detailed and dynamic heat calculation is achieved, which improves the efficiency and accuracy of hot article recognition.

CN120146044APending Publication Date: 2025-06-13AIJI MICRO CONSULTING (XIAMEN) CO LTD
View PDF -1 Cites 0 Cited by

Patent Information

Application Number
CN202510289374.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-22
Filing Date
2025-03-12
Publication Date
2025-06-13

Smart Images

  • Figure CN120146044A_ABST
    Figure CN120146044A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an article popularity calculation method and device, a terminal and a storage medium. According to the scheme, the target article can be obtained, data cleaning is carried out on the target article, word splitting processing is carried out on the cleaned target article through the preset model to obtain the keywords in the target article, the weight information of each keyword is calculated, and the instant popularity of the target article is calculated based on the weight information corresponding to all the keywords. And calculating the instant popularity of the target article at each moment according to the popularity attenuation in a preset period. According to the scheme provided by the embodiment of the invention, the keyword weight and the popularity attenuation model are introduced, so that the calculation of the popularity of the article is more detailed and dynamic, the recognition efficiency and the recognition accuracy of the hot article can be improved, and the attention degree of the article can be reflected more truly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a method, device, terminal and storage medium for calculating article popularity. Background Art

[0002] With the increasing development of the Internet industry, journalists need to timely discover and identify hot articles, so as to obtain the hot trends that the public is concerned about at present from the hot articles. At present, journalists generally identify the relatively hot articles at present according to the hot article click rankings on some large websites.

[0003] In the actual use process, the applicant found that: at present, most article popularity calculations only rely on intuitive data such as click-through rate and forwarding number, which cannot comprehensively reflect the actual influence of the article content, and the popularity evaluation based on social media metrics usually depends on the number of user interactions (such as likes, comments, etc.), without carefully considering the keyword weights of the article content, and even less able to reflect the change process of article popularity. Summary of the Invention

[0004] Embodiments of the present invention provide a method, device, terminal and storage medium for calculating article popularity. By introducing keyword weights and a popularity decay model, the calculation of article popularity is made more detailed and dynamic, which is beneficial to improving the recognition efficiency and accuracy of hot articles, and can more truly reflect the degree of attention of the article.

[0005] Embodiments of the present invention provide a method for calculating article popularity, including:

[0006] Obtain a target article and perform data cleaning on the target article;

[0007] Perform word splitting processing on the cleaned target article through a preset model to obtain the keywords in the target article;

[0008] Calculate the weight information of each keyword, and calculate the instant popularity of the target article based on the weight information corresponding to each keyword;

[0009] Calculate the instant popularity of the target article at each moment according to the popularity decay amount within a preset period.

[0010] Embodiments of the present invention further provide an article popularity calculation device, including:

[0011] An acquisition unit, configured to obtain a target article and perform data cleaning on the target article;

[0012] A word splitting unit, configured to perform word splitting processing on the cleaned target article through a preset model to obtain the keywords in the target article;

[0013] A first calculation unit, configured to calculate the weight information of each keyword, and calculate the instant popularity of the target article based on the weight information corresponding to each keyword.

[0014] A second calculation unit, configured to calculate the instant popularity of the target article at each moment according to the heat decay amount within a preset period.

[0015] An embodiment of the present invention further provides a terminal, including: a memory and a processor. Wherein, an application program processing program is stored on the memory, and when the application program processing program is executed by the processor, the steps of the article heat calculation method provided by any embodiment of the present invention are implemented.

[0016] An embodiment of the present invention further provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute any article heat calculation method provided by an embodiment of the present invention.

[0017] The article heat calculation method provided by the embodiment of the present invention can obtain a target article, perform data cleaning on the target article, perform word splitting processing on the cleaned target article through a preset model to obtain keywords in the target article, calculate the weight information of each keyword, and calculate the instant popularity of the target article based on the weight information corresponding to each keyword, and calculate the instant popularity of the target article at each moment according to the heat decay amount within a preset period. The solution provided by the embodiment of the present invention makes the article heat calculation more detailed and dynamic by introducing keyword weights and a heat decay model, which is beneficial to improving the recognition efficiency and accuracy of hot articles, and can more truly reflect the degree of attention of the article. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.

[0019] Figure 1 is the first flowchart of the article heat calculation method provided by the embodiment of the present invention;

[0020] Figure 2 is the second flowchart of the article heat calculation method provided by the embodiment of the present invention;

[0021] Figure 3 is a schematic structural diagram of an article heat calculation device provided by the embodiment of the present invention;

[0022] Figure 4It is a schematic structural diagram of a terminal provided by an embodiment of the present invention. Detailed implementation manners

[0023] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0024] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, components, features, and elements with the same name in different embodiments of the present invention may have the same meaning or different meanings, and their specific meanings need to be determined according to their explanations in the specific embodiments or further in combination with the context of the specific embodiments.

[0025] It should be understood that although the steps in the flowcharts in the embodiments of the present invention are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps is not strictly limited in order, and they can be executed in other orders. Moreover, at least some of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0026] It should be noted that in this article, step codes such as 101 and 102 are used. The purpose is to more clearly and briefly express the corresponding content and do not constitute a substantial limitation in order. Those skilled in the art may execute 102 first and then 101 during specific implementation, etc., but these should all be within the protection scope of the present invention.

[0027] References to "embodiments" in this specification mean that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0028] An embodiment of the present invention provides a method for calculating article popularity. The execution subject of this method for calculating article popularity can be the article popularity calculation device provided by the embodiments of the present invention, or an intelligent terminal and a server integrated with the article popularity calculation device, where the article popularity calculation device can be implemented in a hardware or software manner.

[0029] In the context of the existing Internet technology, traditional keyword extraction algorithms usually only consider the frequency of words appearing in the text or the distinctiveness from other documents, without combining the weight of words on a specific platform and the heat dynamic decay model. And the heat evaluation based on social media metrics usually relies on the number of user interactions (such as likes, comments, etc.), without carefully considering the keyword weights of the article content.

[0030] Based on this, an embodiment of the present invention provides a new method for calculating article popularity, which can calculate the popularity of an article without the support of business data (such as clicks, likes, shares, etc.), fully considers the weights of different sites, as well as the content related to heat decay, and makes the calculation of article popularity more detailed and dynamic, can more truly reflect the degree of attention of the article, and can perform heat comparison across sites, reflecting the popularity of the article from a new dimension. And through the pre-trained AI model for word splitting processing, it is convenient for system upgrade and expansion.

[0031] Specifically, please refer to Figure 1 , Figure 1 which is the first process schematic diagram of the method for calculating article popularity provided by the embodiments of the present invention. The specific process of this method for calculating article popularity can be as follows:

[0032] 101. Obtain a target article and perform data cleaning on the target article.

[0033] First of all, in this embodiment, it is necessary to screen out the target article for which the popularity value needs to be calculated from numerous articles. This process can be achieved through means such as search engines, professional databases, and social networks. After collecting the target article, perform data cleaning on the screened target article.

[0034] In one embodiment, during the process of data cleaning, it is necessary to first identify and eliminate the interfering data in the target article. These interfering data mainly include HTML tags and special symbols. HTML tags are usually used in web design to display elements such as text, pictures, and links, but they do not belong to the actual content of the target article, so they need to be removed from the target article. Special symbols include various punctuation marks, mathematical symbols, programming language symbols, etc., and they may also interfere with the normal reading and understanding of the target article.

[0035] To ensure the quality and readability of the target article, this embodiment needs to adopt a certain algorithm to detect and delete these interfering data. First, regular expressions can be used to match HTML tags and special symbols, so as to quickly locate their positions in the target article. Next, use the deletion function provided by the text editor or programming language to completely remove these interfering data from the target article.

[0036] When deleting HTML tags and special symbols, the following points need to be noted:

[0037] 1. Ensure that all irrelevant HTML tags are deleted, including headings, paragraphs, lists, etc., so as to make the structure of the target article clearer.

[0038] 2. When deleting special symbols, avoid accidentally deleting key information in the target article, such as mathematical formulas, programming codes, etc.

[0039] 3. After deleting HTML tags and special symbols, reformat the target article to keep it in good layout and beautiful.

[0040] After the above processing, the interfering data in the target article will be effectively eliminated, which helps to improve the accuracy and effect of subsequent processing. That is to say, the steps of data cleaning for the target article described above may include: detecting the interfering data in the target article, where the interfering data includes HTML tags and special symbols, and deleting the interfering data from the target article.

[0041] In other embodiments, the stop words in the target article can be further cleaned to eliminate the meaningless words in the target article. Common stop words include "de", "shi", "he", "zai", etc. These words do not have analytical value in article analysis. The methods for removing stop words include the dictionary-based method and the statistic-based method. The dictionary-based method is to match and delete according to the stop word list, while the statistic-based method is to judge whether a word is a stop word according to its appearance frequency in the corpus and delete the stop word.

[0042] 102. Perform word segmentation on the cleaned target article through a preset model to obtain the keywords in the target article.

[0043] Word segmentation processing is to divide the cleaned target article into individual words. In this embodiment, the above word segmentation methods include dictionary-based word segmentation and statistics-based word segmentation. The dictionary-based word segmentation method relies on a pre-constructed dictionary library and realizes word segmentation of the target article by matching and pruning words. The statistics-based word segmentation method calculates the occurrence frequency of words in a corpus and performs word segmentation according to certain rules. After word segmentation processing, the keywords in the target article can be obtained.

[0044] After completing the word segmentation processing, keywords with higher relevance can be screened according to the part-of-speech information of the segmented words. For example, only retain keywords with practical meanings such as nouns and verbs. Finally, sort the screened keywords to obtain the final keyword list. The sorting can be performed according to various factors such as the part-of-speech of the keywords and their occurrence positions. The more important the part-of-speech and the more prominent the occurrence position of the keyword, the higher its ranking.

[0045] 103. Calculate the weight information of each keyword and calculate the instant popularity of the target article based on the weight information corresponding to all keywords.

[0046] In one embodiment, the weight of the keyword can be further calculated to calculate the subsequent article popularity. The weight of the keyword can be evaluated according to the importance of the keyword in the cleaned target article. In this embodiment, various factors such as word frequency, semantic weight, and sentence position can be used to calculate the weight information of the keyword. For example: the weight of the keyword = word frequency weight + semantic weight + position weight, where the word frequency weight is determined according to the word frequency of the keyword, the semantic weight is based on the semantic importance of the keyword itself, and the position weight is determined according to the occurrence position of the keyword in the target article.

[0047] Specifically, the occurrence frequency (word frequency) of each keyword can be counted by performing text analysis on the cleaned target article. For example, in a target article about technology products, if the keyword "smartphone" appears 10 times and the keyword "tablet computer" appears 5 times, then it can be confirmed that the word frequency of "smartphone" is higher than that of "tablet computer". In Python, the collections.Counter class can be used to quickly count the word frequency. Correspondingly, the higher the word frequency, the relatively higher the importance of the keyword in the target article and the corresponding increase in its weight.

[0048] The semantic weight is determined based on the semantic importance of the keyword itself, and this determination process can be carried out through an external semantic knowledge base or corpus. For example, in the ICT field, the semantic weights of professional terms such as "semiconductor" and "communication" are higher than those of ordinary words. Specifically, the semantic weight of a keyword can be determined by querying a professional dictionary or an industry standard terminology database. For example, in the ICT industry database, corresponding semantic importance annotations are provided for various professional term names, and these annotations can be referred to when calculating the article popularity to determine the semantic weight. A keyword with a higher semantic weight may play a key role in the calculation of article popularity even if its frequency is relatively low, so its weight will also be higher than that of ordinary keywords with a higher occurrence frequency.

[0049] The sentence position mainly considers whether the position where the keyword appears in the target article is crucial. For example, the weights of keywords that appear in positions such as the title of the target article, the beginning paragraph, the ending paragraph, and the first and last sentences of paragraphs will be relatively high. The position weight can also be combined with the word frequency and semantic weight to calculate the weight information of the keyword. For example, a keyword appears multiple times in the middle paragraph of the target article but does not appear in the title and key positions, then its weight will be lower due to the position factor than that of a keyword with relatively fewer occurrences in key positions but with high semantic importance. By comprehensively considering these three factors, the weight information of the keyword can be calculated more comprehensively and accurately.

[0050] Then, based on the weight information of each keyword, calculate their respective popularity values, and average the calculated popularity values of each keyword to obtain the popularity value of the target article. Among them, the occurrence times of each keyword in the target article can be calculated through word frequency statistics. Word frequency statistics can help users understand which keywords appear more frequently in the target article, so as to mine the theme or key information of the target article.

[0051] 104. Calculate the instant popularity of the target article at each moment according to the popularity decay amount within a preset period.

[0052] In an embodiment, corresponding color identifiers can also be set for the instant popularity level of the target article at each moment to visually display the change in article popularity. The steps for calculating the instant popularity of the target article at each moment can specifically include:

[0053] 1. Determine the preset period: The preset period can be adjusted according to the actual situation, such as daily, weekly, or monthly, etc. The popularity decay within the preset period can be calculated based on the time stamp.

[0054] 2. Set the heat attenuation coefficient: According to the target article type and market demand, set an attenuation coefficient for the heat of each preset period. For example, on the first day, the article heat attenuation coefficient is 1; on the second day, the heat attenuation coefficient is 0.9; on the third day, the heat attenuation coefficient is 0.85, and so on.

[0055] 3. Calculate the heat at each moment: According to the formula "instantaneous heat = initial heat × heat attenuation coefficient", calculate the instantaneous heat of the target article at each moment. Among them, the initial heat can be the heat when the target article is published.

[0056] 4. Heat level classification: Divide the instantaneous heat into different levels, such as popular, relatively hot, average, cold, etc. Corresponding colors can be set for each level to facilitate visual display of the change in the heat of the target article.

[0057] 5. Update hotspots in real time: During the preset period, update the heat of the target article in real time according to user behavior data such as reading, liking, and commenting. At the same time, adjust the instantaneous heat of the target article at each moment according to the heat attenuation coefficient.

[0058] 6. Data visualization: Display the heat data of the target article in the form of a chart or dynamic curve, so that users can clearly understand the change trend of the heat of the target article.

[0059] Through the above methods, the instantaneous heat of the target article at each moment can be calculated more accurately, and valuable heat information can be provided to users. On this basis, other indicators such as article quality and user feedback can be combined to provide more personalized recommended content for users. Thereby improving the user experience and enhancing the platform activity.

[0060] Through the above steps, a method for calculating article heat based on word frequency can be realized. By cleaning, splitting words, extracting core words and calculating weights of the content of the target article, combined with site weights and time decay factors, the heat of the target article is evaluated. It mainly solves the problem that the determination of article heat in the existing technology overly relies on business data and the method is single and not accurate and comprehensive enough.

[0061] As described above, the article popularity calculation method proposed in the embodiments of the present invention can obtain a target article, perform data cleaning on the target article, perform word splitting on the cleaned target article through a preset model to obtain keywords in the target article, calculate the weight information of each keyword, and calculate the instant popularity of the target article based on the weight information corresponding to each keyword. The instant popularity of the target article at each moment is calculated according to the heat decay amount within a preset period. The solution provided by the embodiments of the present invention makes the article popularity calculation more detailed and dynamic by introducing keyword weights and a heat decay model, which is beneficial to improving the recognition efficiency and accuracy of hot articles and can more truly reflect the degree of attention of the article.

[0062] According to the method described in the previous embodiments, further detailed description will be given below.

[0063] Please refer to Figure 2 , Figure 2 which is the second process schematic diagram of the article popularity calculation method provided by the embodiments of the present invention. The method includes:

[0064] 201. Obtain a target article and perform data cleaning on the target article.

[0065] 202. Establish a hot word library according to the hot articles and hot words on the network within a preset time period.

[0066] In one embodiment, the process of establishing the hot word library may specifically include: 1. Determine the time period: First, a time range needs to be determined so as to collect hot articles within this time range. This time range can be adjusted according to actual situations, such as one week, one month, or one quarter, etc. 2. Select platforms: To ensure the comprehensiveness and representativeness of the collected hot articles, multiple network platforms can be selected for collecting hot articles, such as news websites, social media, blogs, etc. 3. Article screening: Conduct a preliminary screening of the collected hot articles, and eliminate the hot articles that are irrelevant to the hot words or have low quality. Preliminary screening can be carried out through manual screening or by using text mining techniques. 4. Hot word extraction: Extract the hot words in the screened hot articles. The hot words can be keywords, themes, or topics, etc., and can be selected according to actual situations, thereby completing the construction of the hot word library.

[0067] 203. Extract the split words in the target article, map the split words to the hot word library using a word set model, and obtain keywords according to the mapping results.

[0068] In one embodiment, a hot word thesaurus obtained according to the above steps is used, and a split word in the target article is mapped to the hot word thesaurus by using a word set model. The word set model is a simplified model for article analysis and information retrieval, mainly focusing on the set of words appearing in the article, without considering the frequency or order of the word appearance. In this word set model, each article is regarded as a set of words, and the focus is on determining which words are included in the article, rather than the specific situation of the word appearance. Specifically, each split word is compared with the words in the hot word thesaurus one by one through the word set model. If a split word exists in the hot word thesaurus, then the position recording the corresponding relationship between the split word and the hot word thesaurus is marked as 1, indicating a successful match; if the split word does not exist in the hot word thesaurus, it is marked as 0, indicating a mismatch. The word frequency or order is not considered during the processing of the word set model. After all the split words are mapped, a matching result list composed of 0s and 1s is obtained, and the split words corresponding to 1 are the keywords.

[0069] Further, after obtaining the keywords, the above method may further include: performing word frequency statistics on the obtained keywords through the TF-IDF algorithm, and selecting multiple target keywords with frequencies higher than a preset value according to the statistical results. Among them, TF-IDF (term frequency–inverse document frequency) is a commonly used weighting technique for information retrieval and text mining. It is used to evaluate the importance of a word for a document set or an article in a corpus (used to evaluate the importance of a word or term relative to other words in a document set or a corpus). The importance of a word increases in direct proportion to the number of times it appears in the article, but at the same time decreases in inverse proportion to the frequency of its appearance in the corpus.

[0070] 204. Calculate the weight value of each keyword according to the basic weight, platform weight, and appearance frequency of the keyword.

[0071] 205. Calculate the heat value of each keyword according to the weight value of the keyword, and calculate the instant heat of the target article based on the heat value of each keyword.

[0072] In one embodiment, the basic weight is assigned according to the commonness and importance of the keyword in the network. Generally speaking, the basic weight of a commonly used and important keyword will be higher. For example, if a keyword is a core concept in the industry, then its basic weight will be higher. The platform weight is calculated according to the performance of the keyword on major network platforms. For example, if a keyword appears frequently on platforms such as search engines, social media, and news websites, then its platform weight will be correspondingly increased.

[0073] In addition to the basic weight and platform weight, the frequency of occurrence of keywords in the target article also needs to be considered. Keywords with a high frequency of occurrence often represent the attention and activity of the content. Therefore, the frequency of occurrence is also an important factor in calculating the keyword weight. In one embodiment, the above three factors (basic weight, platform weight, and frequency of occurrence) can be combined to calculate the final weight of the keyword. Specifically, the importance of these three factors can be evaluated first. For example, for a certain application scenario, it can be confirmed that the basic weight is the most important, followed by the platform weight, and finally the frequency of occurrence. Then, they are added according to their respective importance ratios to finally obtain the weight value of the keyword. If it is considered that the basic weight, platform weight, and frequency of occurrence are equally important, then they can be added and averaged to be used as the weight value of the keyword.

[0074] After calculating the weight value of each keyword, the popularity value of each keyword can be further calculated. The popularity value is comprehensively calculated based on the weight value and the frequency of occurrence, and it can reflect the popularity of a keyword in the network.

[0075] Finally, according to the popularity value of each keyword, the instant popularity of the target article can be calculated. The instant popularity is a dynamic indicator that changes with the change of the keyword popularity. Thus, it provides a data basis for users to reflect the popularity of the article in real time.

[0076] 206. Calculate the instant popularity of the target article at each moment according to the heat decay amount within the preset period.

[0077] For example, the preset period can be set to one week. On the first day, the article heat decay coefficient is 1; on the second day, the heat decay coefficient is 0.9; on the third day, the heat decay coefficient is 0.85. And so on, the instant popularity of the target article can be calculated every day within one week.

[0078] 207. Stop calculating the instant popularity when the heat cycle is exceeded after the target article is published, and average all the instant popularities calculated within the heat cycle to obtain the total popularity of the target article within the heat cycle.

[0079] In one embodiment, after obtaining the total popularity of the target article within the heat cycle, the method may further include: calculating a calibration coefficient according to the interaction data of the target article, where the interaction data includes click-through rate, repost rate, and comment rate, and calibrating the total popularity of the target article based on the calibration coefficient. The above interaction data may further include the number of likes, the number of bullet screens, and the number of collections, etc., and the present invention will not elaborate further on this.

[0080] Specifically, the importance weights of the above-mentioned interaction data can be determined according to the scenario first. For example, when the target article is a news article, the click-through rate is relatively important because it represents the exposure of the target article. When the target article is content on social media, the repost rate and comment rate can better reflect the influence of the target article. Therefore, it is necessary to evaluate the importance of these interaction data according to the actual scenario.

[0081] Then, the calibration coefficient is obtained by comprehensively calculating each interaction data index. For example, if the repost rate is regarded as the most important index and given a weight of 0.4, the comment rate is 0.3, the like count is 0.2, and the favorite count is 0.1, then the calibration coefficient can be calculated as follows: Calibration coefficient = repost rate × 0.4 + comment rate × 0.3 + like count × 0.2 + favorite count × 0.1. This calibration coefficient can be used as a scaling factor to adjust the total popularity. If the calibration coefficient is greater than 1, it means that the interaction data of the target article indicates that the target article is more popular than previously thought, and its total popularity should be increased. If the calibration coefficient is less than 1, it indicates that the interaction data of the target article is not outstanding enough, and the total popularity needs to be appropriately reduced. For example, if the total popularity of the target article is 100 and the calibration coefficient is 1.2, the calibrated total popularity is 100 × 1.2 = 120. If the calibration coefficient is 0.8, the calibrated total popularity is 100 × 0.8 = 80.

[0082] In one embodiment, the calibration of the total popularity of the target article is not a one-time operation, but can be carried out dynamically according to different popularity cycles (such as daily popularity, weekly popularity, monthly popularity, etc.). As time goes by, new interaction data is continuously generated, and the calibration coefficient will be updated, so as to continuously calibrate the total popularity, enabling the popularity of the target article to reflect its popularity and influence among users in real time.

[0083] As described above, the article popularity calculation method proposed in the embodiments of the present invention can obtain a target article, perform data cleaning on the target article, establish a hot word library based on hot articles and hot words in the network within a preset time period, extract split words from the target article, and use a word set model to map the split words to the hot word library, obtain keywords according to the mapping results, calculate the weight value of each keyword according to the basic weight, platform weight and occurrence frequency of the keyword, calculate the popularity value of each keyword according to the weight value of the keyword, calculate the instant popularity of the target article based on the popularity value of each keyword, calculate the instant popularity of the target article at each moment according to the popularity decay amount within a preset period, stop calculating the instant popularity when the target article is published and exceeds the popularity period, and average all the instant popularities calculated within the popularity period to obtain the total popularity of the target article within the popularity period. The solution provided by the embodiments of the present invention makes the article popularity calculation more detailed and dynamic by introducing keyword weights and a popularity decay model, which is beneficial to improving the recognition efficiency and accuracy of hot articles and can more truly reflect the degree of attention of the article.

[0084] To implement the above method, an article popularity calculation device is further provided in the embodiments of the present invention. The article popularity calculation device can be specifically integrated in terminal devices such as mobile phones and tablet computers.

[0085] For example, as Figure 3 shown, is a schematic structural diagram of an article popularity calculation device provided by an embodiment of the present invention. The article popularity calculation device may include:

[0086] An acquisition unit 301, configured to acquire a target article and perform data cleaning on the target article;

[0087] A word splitting unit 302, configured to perform word splitting processing on the cleaned target article through a preset model to obtain keywords in the target article;

[0088] A first calculation unit 303, configured to calculate the weight information of each keyword and calculate the instant popularity of the target article based on the weight information corresponding to each keyword;

[0089] A second calculation unit 304, configured to calculate the instant popularity of the target article at each moment according to the popularity decay amount within a preset period.

[0090] The article heat calculation device provided by the embodiments of the present invention can obtain a target article, perform data cleaning on the target article, perform word splitting processing on the cleaned target article through a preset model to obtain keywords in the target article, calculate the weight information of each keyword, and calculate the instant heat of the target article based on the weight information corresponding to each keyword, and calculate the instant heat of the target article at each moment according to the heat attenuation amount within a preset period. The solution provided by the embodiments of the present invention makes the article heat calculation more detailed and dynamic by introducing keyword weights and a heat attenuation model, which is beneficial to improving the recognition efficiency and accuracy of hot articles and can more truly reflect the degree of attention of the article.

[0091] The embodiments of the present invention also provide a terminal, as Figure 4 shown. The terminal 400 may include a memory 401 and a processor 402. Among them, an application processing program is stored on the memory 401, and when the application processing program is executed by the processor 402, the steps of the above-mentioned article heat calculation method based on the present embodiment are implemented.

[0092] Specifically, in this embodiment, the processor 402 in the terminal will load the executable files corresponding to the processes of one or more applications into the memory 401 according to the following instructions, and the processor 402 will run the applications stored in the memory 401 to implement various functions:

[0093] Obtain a target article and perform data cleaning on the target article;

[0094] Perform word splitting processing on the cleaned target article through a preset model to obtain keywords in the target article;

[0095] Calculate the weight information of each keyword and calculate the instant heat of the target article based on the weight information corresponding to each keyword;

[0096] Calculate the instant heat of the target article at each moment according to the heat attenuation amount within a preset period.

[0097] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the detailed description of the article heat calculation method above, and details will not be repeated here.

[0098] As can be seen from the above, the terminal according to the embodiment of the present invention can obtain a target article, perform data cleaning on the target article, perform word splitting on the cleaned target article through a preset model to obtain keywords in the target article, calculate the weight information of each keyword, and calculate the instant popularity of the target article based on the weight information corresponding to each keyword, and calculate the instant popularity of the target article at each moment according to the heat decay amount within a preset period. The solution provided by the embodiment of the present invention makes the calculation of article popularity more detailed and dynamic by introducing keyword weights and a heat decay model, which is beneficial to improving the recognition efficiency and accuracy of hot articles and can more truly reflect the degree of attention of the article.

[0099] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0100] Therefore, the embodiment of the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions can be loaded by a processor to execute the steps in any of the article heat calculation methods provided by the embodiment of the present invention.

[0101] For the specific implementation of each of the above operations, reference can be made to the previous embodiments and will not be elaborated here.

[0102] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0103] Since the instructions stored in the storage medium can execute the steps in any of the article heat calculation methods provided by the embodiment of the present invention, the beneficial effects that can be achieved by any of the article heat calculation methods provided by the embodiment of the present invention can be realized. For details, refer to the previous embodiments and will not be elaborated here.

[0104] The above has introduced in detail an article heat calculation method, device, terminal and storage medium provided by the embodiment of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for calculating article popularity, characterized in that: include: Obtaining a target article and performing data cleaning on the target article; The cleaned target article is processed by word splitting using a preset model to obtain keywords in the target article; Calculate the weight information of each keyword, and calculate the instant popularity of the target article based on the weight information corresponding to all the keywords; The instantaneous popularity of the target article at each moment is calculated according to the heat decay amount within a preset period.

2. The article popularity calculation method according to claim 1, characterized in that: The data cleaning of the target article includes: Detecting interference data in the target article, wherein the interference data includes HTML tags and special symbols; The interfering data is deleted from the target article.

3. The article popularity calculation method according to claim 1, characterized in that: The process of performing word splitting processing on the cleaned target article by using a preset model to obtain keywords in the target article includes: Establish a hot word database based on hot articles and hot words on the Internet within a preset time period; The split words in the cleaned target article are extracted, and the split words are matched to the hot word library using a word set model, and the keywords are obtained according to the matching results.

4. The article popularity calculation method according to claim 3, characterized in that: After obtaining the keyword, the method further includes: Use the TF-IDF algorithm to perform word frequency statistics on the obtained keywords; According to the statistical results, multiple target keywords with frequencies higher than the preset values ​​are selected.

5. The article popularity calculation method according to claim 1, characterized in that: The calculating of the weight information of each keyword, and the calculating of the instant popularity of the target article based on the weight information corresponding to all the keywords, includes: Calculate the weight value of each keyword according to the basic weight, platform weight and occurrence frequency of the keyword; Calculate the popularity value of each keyword according to the weight value of the keyword; The instant popularity of the target article is calculated based on the popularity value of each keyword.

6. The article popularity calculation method according to any one of claims 1 to 5, characterized in that: The method further comprises: When the target article is published and the heat period has expired, the calculation of the instant heat is stopped, and all the instant heats calculated within the heat period are averaged to obtain the total heat of the target article within the heat period.

7. The article popularity calculation method according to claim 6, characterized in that: After obtaining the total popularity of the target article in the popularity period, the method further includes: Calculating a calibration coefficient based on the interactive data of the target article, wherein the interactive data includes clicks, forwardings, and comments; The total popularity of the target article is calibrated based on the calibration coefficient.

8. An article popularity calculation device, characterized in that: include: An acquisition unit, used for acquiring a target article and performing data cleaning on the target article; A word splitting unit is used to perform word splitting processing on the cleaned target article through a preset model to obtain keywords in the target article; A first calculation unit, used to calculate the weight information of each keyword, and calculate the instant popularity of the target article based on the weight information corresponding to all the keywords; The second calculation unit is used to calculate the instantaneous popularity of the target article at each moment according to the heat decay amount within a preset period.

9. A terminal, characterized in that: The terminal comprises: a memory and a processor, wherein the memory stores an application processing program, and when the application processing program is executed by the processor, the steps of the article heat calculation method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the article popularity calculation method described in any one of claims 1 to 7.