News data classification method and device and electronic equipment

Through the news data classification method of multi-level feature mining and dynamic feedback mechanism, the problem that existing technology is difficult to adapt to regional cultural diversity and policy sensitivity is solved, and the accuracy, regional adaptability and timeliness of news classification are improved.

CN120179823AActive Publication Date: 2025-06-20WEIFANG UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510661148.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The existing technology is difficult to adapt to the regional cultural diversity and policy sensitivity of regional content, and traditional news classification methods are prone to misjudgment of cultural events when dealing with local news, and cannot quantify the degree of response of news to local policies.

Method used

A news data classification method with a multi-level feature mining and dynamic feedback mechanism is adopted to construct regional correlation characteristics by extracting dialect matching characteristics, cultural matching characteristics and policy matching characteristics, and dynamically correct classification priorities based on user interaction data and regional correlation.

Benefits of technology

Significantly improve the accuracy, regional adaptability and timeliness of news classification, ensure that local events are not diluted by common vocabulary, high-communication news is automatically highlighted, and the robustness of core topic identification is enhanced, taking into account policy orientation and public attention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179823A_ABST
    Figure CN120179823A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a news data classification method and device and electronic equipment, and the method comprises the steps: extracting a dialect matching feature, a culture matching feature and a policy matching feature of a collected news text; constructing a regional association feature based on the dialect matching feature, the culture matching feature and the policy matching feature of the news text; matching the news text with four types of preset keywords of livelihood, capital construction, culture and ecology to determine a news text type and a news text type index; constructing a popularity index of the news text based on the forwarding, comment and like quantity of the news text and the regional association characteristics of the news text, and updating a news text category index according to the popularity index of the news text; and updating the news text category based on the popularity index of the news text, the news text category index and the policy matching feature. The classification efficiency of the news data is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a news data classification method, apparatus, and electronic device. Background Art

[0002] With the rapid spread of Internet news, traditional news classification methods mostly rely on basic keyword matching or general topic models (such as LDA), and it is difficult to adapt to the regional cultural diversity and policy sensitivity of regional content.

[0003] Existing technologies often adopt single-dimensional feature extraction. For example, only based on word frequency to analyze news topics, ignoring deep regional correlation information such as dialect expressions, local cultural customs, and policy orientations. Especially when dealing with local news, such methods are prone to misjudging the attribution of cultural events due to the dominance of common words, and cannot quantify the response degree of news to local policies. In addition, traditional popularity calculation models often simply superimpose indicators such as reposts, comments, etc., without dynamically adjusting the classification weights in combination with the regional attributes of news, resulting in news with high interaction but deviating from regional core issues occupying resources, affecting the accuracy of policy promotion or cultural dissemination. The static classification logic and single data source dependence of current technologies are difficult to meet the comprehensive requirements of accuracy, timeliness, and adaptability in scenarios such as personalized recommendation of local news. Summary of the Invention

[0004] The purpose of the present invention is to provide a news data classification method, apparatus, and electronic device to solve at least one of the problems existing in the prior art.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A news data classification method includes: extracting the dialect matching feature, cultural matching feature, and policy matching feature of the collected news text; constructing a regional association feature based on the dialect matching feature, cultural matching feature, and policy matching feature of the news text; matching the news text with four types of preset keywords of people's livelihood, infrastructure, culture, and ecology to determine the news text category and news text category index; constructing a popularity index of the news text based on the number of reposts, comments, likes of the news text and the regional association feature of the news text, and updating the news text category index according to the popularity index of the news text; updating the news text category based on the popularity index of the news text, the news text category index, and the policy matching feature.

[0006] Optionally, extract the dialect words in the news text, count the number of dialect words in the news text as F, count the total number of words in the news text as Fz, and set the dialect matching feature of the news text as Mf. The expression of Mf is: ; In the formula, dk is the edit distance between the k-th dialect word in the news text and the corresponding standard word, and Dmax is the edit distance threshold.

[0007] Optionally, count the number of sentences containing the j-th cultural keyword in the news text as Wj, count the total number of sentences in the news text as Wz, and set the cultural matching feature as Mh. The expression of Mh is ; In the formula, α1 is the first weight, α2 is the second weight, α1 + α2 = 1, J is the number of cultural keywords in the news text, and Jz is the preset number threshold.

[0008] Optionally, calculate the relevance between the news text and each policy document in the target area policy library, and set the relevance Qi between the news text and the i-th policy document in the target area policy library. The expression of Qi is: ; In the formula, Ti is the number of words in the i-th policy document in the target area policy library, tf(ti) is the number of times the ti-th word in the i-th policy document in the target area policy library appears in the news text, k1 is the saturation control factor, b is the length normalization factor, G is the average number of words in all policy documents in the target area policy library, N is the average number of words in all documents in the policy library, and df(ti) is the number of policy documents in the target area policy library that contain the ti-th word in the i-th policy document; Take the ratio of the average value of the relevance of each policy document in the target area policy library to the preset relevance threshold as the policy matching feature, denoted as zf.

[0009] Optionally, perform data fusion on the dialect matching feature Mf, cultural matching feature Mh, and policy matching feature zf of the news text to construct the regional association feature Dy. The expression of Dy is: Dy = x1×Mf + x2×Mh + x3×zf; in the formula, x1 is the dialect weight, x2 is the cultural weight, and x3 is the policy weight.

[0010] Optionally, match the news text with the preset keywords of various categories, calculate the category score of the news text according to the matching results, and denote the matching score of the news text with the preset keywords of category c as Sc. The expression of Sc is: ; Where Ca is the number of preset keywords of category c in the news text, and Wca is the position weight of the ca-th preset keyword of category c in the news text.

[0011] Optionally, a heat index H of the news text is constructed based on the number of forwards, comments, likes of the news text and the regional association characteristics of the news text. The expression of the heat index H of the news text is as follows: H = ln(β1×r1 / u1 + β2×r2 / u2 + β3×r3 / u3 + 1) × tanh(3×Dy); Where H is the heat index of the news text, β1 is the forward weight, β2 is the comment weight, β3 is the like weight, β1 + β2 + β3 = 1, r1 is the number of forwards, u1 is the forward threshold, r2 is the number of comments, u2 is the comment threshold, r3 is the number of likes, and u3 is the like threshold; Compare the heat index H of the news text with the heat threshold H0. When H is greater than or equal to H0, update the news text category index to Sg.

[0012] Optionally, the heat index of the news text, the news text category index and the policy matching characteristics are fused with data to construct a confidence index ZX of the news text. The expression of the confidence index ZX is: ZX = (2×news text category index×H) / {(news text category index + H)×[1 + exp(-5×zf)]}; Compare the confidence index ZX with the confidence threshold zx0. When ZX is less than zx0, update the news text category to no category. Otherwise, maintain the news text category.

[0013] According to another aspect of the present application, there is provided a news data classification device, including: A feature extraction unit for extracting the dialect matching feature, cultural matching feature and policy matching feature of the collected news text; An association feature construction unit for constructing a regional association feature based on the dialect matching feature, cultural matching feature and policy matching feature of the news text; A classification unit for matching the news text with four types of preset keywords of people's livelihood, infrastructure, culture, and ecology to determine the news text category and the news text category index; A heat construction unit for constructing a heat index of the news text based on the number of forwards, comments, likes of the news text and the regional association feature of the news text, and updating the news text category index according to the heat index of the news text; A category update unit for updating the news text category based on the heat index of the news text, the news text category index and the policy matching feature.

[0014] According to another aspect of the present application, there is provided an electronic device, which includes: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the news data classification method described above.

[0015] The beneficial effects of the present invention are as follows: Through multi-level feature mining and a dynamic feedback mechanism, the accuracy, regional adaptability, and timeliness of news classification are significantly improved. First, by integrating multi-dimensional analysis of dialect words, cultural keywords, and policy relevance, the limitation of traditional single-semantic classification is broken through, and the regional characteristics of news are accurately captured, especially suitable for regions with complex dialects and diverse cultures (such as southern Fujian and Cantonese-speaking areas), ensuring that local events are not diluted by common vocabulary. Second, based on the heat calculation of user interaction data and regional relevance, the classification priority is dynamically corrected, making highly transmissible news stand out automatically. Third, through keyword position weighting and confidence filtering mechanisms, the recognition robustness of core issues such as people's livelihood and infrastructure is strengthened, avoiding interference from misclassification. Finally, taking into account policy orientation and public attention, an efficient matching of "content - region - demand" is achieved. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of the news data classification method for this embodiment.

[0018] Figure 2 It is a flowchart of the feature extraction method described in this embodiment.

[0019] Figure 3 It is a flowchart of the update method of the news text category index for this embodiment.

[0020] Figure 4 It is a structural schematic diagram of the news data classification device for this embodiment.

[0021] Figure 5 It is a structural schematic diagram of the electronic device for this embodiment. Detailed Embodiments

[0022] To more clearly illustrate the present invention, the present invention will be further described below in conjunction with preferred embodiments and the accompanying drawings. Similar components in the drawings are denoted by the same reference numerals. Those skilled in the art should understand that the content specifically described below is illustrative rather than restrictive, and should not be used to limit the protection scope of the present invention.

[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so as to implement the embodiments of the present application described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0024] Specifically, the news data classification method described in this embodiment is applied to the classification of local news, providing news content that conforms to the dialect and cultural preferences for regional users (such as in southern Fujian), and increasing the user's stay time.

[0025] Please refer to Figure 1 as shown, which is a schematic flowchart of the news data classification method in this embodiment, including: Step S101, extract the dialect matching feature, cultural matching feature and policy matching feature of the collected news text.

[0026] Specifically, by mining dialect words, culture-related expressions and policy relevance in the news, the regional attributes and cultural background of the text are captured from multiple dimensions, enhancing the sensitivity to local characteristic content, and providing multi-dimensional feature support for subsequent classification.

[0027] Exemplarily, in this embodiment, news texts in the target area can be crawled through a crawler. The target area is the area for news text classification in this embodiment, and structured fields such as title, text, release time, etc. are extracted; in this embodiment, the collection method of news texts is not specifically limited, and those skilled in the art can freely set it according to needs.

[0028] Please refer to Figure 2 as shown, the feature extraction method includes: Step S201, extract the dialect matching feature of the collected news text.

[0029] Specifically, extract the dialect words in the news text, count the number of dialect words in the news text as F, count the total number of words in the news text as Fz, and set the dialect matching feature of the news text as Mf. The expression of Mf is: ; In the formula, dk is the edit distance between the k-th dialect word in the news text and the corresponding standard word, and Dmax is the edit distance threshold.

[0030] Specifically, in this embodiment, only the news texts with dialect words are analyzed.

[0031] Specifically, identify the unique dialect words in the news and the degree of difference from the standard words, accurately reflect the localization expression tendency of the text, avoid missing local characteristics due to the dominance of common words, and improve the pertinence of regional analysis.

[0032] Exemplarily, in this embodiment, the edit distance threshold can be set to 5; in this embodiment, no specific limitation is made on the setting of the edit distance threshold, and those skilled in the art can freely set it according to needs.

[0033] Specifically, the edit distance between the dialect word and the corresponding standard word in this embodiment is the minimum number of edit operations required to convert the dialect word into the corresponding standard writing in the standard word.

[0034] Exemplarily, when extracting the dialect words in the news text in this embodiment, a dialect word dictionary of the target area, such as "Dictionary of Minnan Dialect", etc., can be obtained. Use a pre-trained dialect NER model (such as CRF + BiLSTM) to scan the full text of the news. After text tokenization, judge whether each word is a dialect word through the model, and only retain the dialect words that match the target area label (for example, when processing Guangdong news, filter out the words belonging to the "Guangdong" label), and count the number of dialect words as F; for obtaining the total number of words in the news text, the news text can be cleaned to remove noise such as punctuation marks, HTML tags, non-character characters, etc., and use a Chinese word segmentation tool such as HanLP to segment the text, filter out stop words (such as "de", "le", "zai") and single-character words (such as "a", "o"), only retain the words with practical meanings, and count the total number Fz; the pre-trained dialect NER model can be called to identify the word boundaries, and the Python Levenshtein library can be used to obtain the edit distance between the dialect word and the corresponding standard word.

[0035] Please continue to refer to Figure 2 As shown, the feature extraction method further includes: Step S202, extract the cultural matching features of the collected news text.

[0036] Specifically, the number of sentences containing the j-th cultural keyword in the news text is counted as Wj, the total number of sentences in the news text is counted as Wz, and the cultural matching feature is set as Mh. The expression of Mh is ; In the formula, α1 is the first weight, α2 is the second weight, α1 + α2 = 1, J is the number of cultural keywords in the news text, and Jz is the preset quantity threshold.

[0037] Specifically, by statistically analyzing the distribution density of cultural keywords at the sentence level, the deeply related content of cultural events or traditions can be effectively identified, and the classification ability of news on cultural themes such as intangible cultural heritage and folk customs can be strengthened.

[0038] Exemplarily, in this embodiment, the first weight can be set to 0.7, the second weight can be set to 0.3, and the preset quantity threshold can be set to 50; in this embodiment, no specific limitations are imposed on the above values, and those skilled in the art can freely set them according to requirements.

[0039] Exemplarily, in this embodiment, a Chinese sentence segmentation tool (such as rule-based sentence splitting or a pre-trained model) can be used to split the news text into independent sentences, with full stops (.), question marks (?), exclamation marks (!), line breaks (\n), etc. as sentence boundary markers, and special cases are processed (such as ellipsis (...) being merged into a single sentence terminator). At the same time, invalid sentences with a length less than the threshold (such as sentences with a length ≤ 2, which may be noise or typesetting errors) are excluded, and the total number of sentences is counted as Wz. Each sentence is cleaned, and it is determined whether it contains a cultural keyword. If there is a match, the counter is incremented, and the total number of sentences in the news text is counted as Wz.

[0040] Please continue to refer to Figure 2 As shown, the feature extraction method further includes: Step S203, extracting the policy matching feature from the collected news text.

[0041] Specifically, the relevance between the news text and each policy document in the target area policy library is calculated, and the relevance Qi between the news text and the i-th policy document in the target area policy library is expressed as: ; Wherein, Ti is the number of words in the i-th policy document in the target area policy library, tf(ti) is the number of times the ti-th word in the i-th policy document in the target area policy library appears in the news text, k1 is the saturation control factor, b is the length normalization factor, G is the average number of words in all policy documents in the target area policy library, N is the average number of words in all documents in the policy library, and df(ti) is the number of policy documents in the target area policy library that contain the ti-th word in the i-th policy document; The ratio of the average of the relevance degrees of each policy document in the target area policy library to the preset relevance threshold is used as the policy matching feature, denoted as zf.

[0042] Specifically, quantify the lexical relevance between news and local policies, measure the degree of news dissemination or response to policies, and assist in evaluating the social impact and public attention of news, so as to improve the accuracy of constructing the geographical association features of news texts.

[0043] Exemplarily, in this embodiment, the saturation control factor can be set to 1.2, the length normalization factor can be set to 0.75, and the preset relevance threshold can be set to 50. In this embodiment, no specific limitations are imposed on the above values, and those skilled in the art can freely set them according to needs.

[0044] Exemplarily, in this embodiment, policy documents can be regularly crawled through web crawlers, the word frequency tf(ti) can be counted by using the built-in collections.Counter in Python, the number of words Ti can be directly calculated by len() after word segmentation, the average number of words in all documents can be calculated by using the Python numerical calculation library, and df(ti) can be obtained by calling APIs such as.termvectors(). In this embodiment, no specific limitations are imposed on the above data acquisition methods, and those skilled in the art can freely set them according to needs.

[0045] Please continue to refer to Figure 1 As shown, the news data classification method further includes: Step S102, constructing geographical association features based on the dialect matching feature, cultural matching feature, and policy matching feature of the news text.

[0046] Specifically, the dialect matching feature Mf, cultural matching feature Mh, and policy matching feature zf of the news text are fused to construct the geographical association feature Dy. The expression of Dy is: Dy = x1×Mf + x2×Mh + x3×zf; where x1 is the dialect weight, x2 is the cultural weight, x3 is the policy weight, and x1 + x2 + x3 = 1.

[0047] Specifically, by integrating the characteristics of dialect, culture, and policy in three dimensions, a comprehensive regional correlation index is generated to avoid the one-sidedness of single-feature analysis and improve the overall discrimination accuracy of region-related content.

[0048] Exemplarily, in this embodiment, the dialect weight can be set to 0.5, the culture weight can be set to 0.3, and the policy weight can be set to 0.2; the above values are not specifically limited in this embodiment, and those skilled in the art can freely set them according to requirements.

[0049] Please continue to refer to Figure 1 As shown, the news classification method further includes: Step S103: Match the news text with four types of preset keywords, namely people's livelihood, infrastructure, culture, and ecology, to determine the news text category and the news text category index.

[0050] Specifically, match the news text with the preset keywords of each category, calculate the category score of the news text according to the matching result, and record the matching score of the news text with the preset keywords of category c as Sc. The expression of Sc is: ; In the formula, Ca is the number of preset keywords of category c in the news text, and Wca is the position weight of the ca-th preset keyword of category c in the news text. If it is in the title position, Wca takes the value of 1; if it is in the first paragraph of the text, Wca takes the value of 2; if it is in the last paragraph of the text, Wca takes the value of 5; if it is in other paragraphs of the text, Wca takes the value of 3. Take the matching score of the news text with the preset keywords of each category as the category score of the news text, sort the category scores of the news text in descending order, take the category corresponding to the maximum category score as the news text category, and take its corresponding category score as the news text category index, denoted as Sh.

[0051] Specifically, the preset keywords of the culture category in this embodiment are the same as the above-mentioned culture keywords. c is the category number. When c = 1, it is the people's livelihood category; when c = 2, it is the infrastructure category; when c = 3, it is the culture category; when c = 4, it is the ecology category. Ca includes repeated preset keywords. That is, when calculating the matching score, if a repeated preset keyword appears in the news text, the position weight of each preset keyword needs to be counted.

[0052] Specifically, by weighted matching of keywords in fixed categories such as people's livelihood and infrastructure, the news theme attribution is clarified, and the content bias is reflected by combining keyword weights, helping to quickly locate the core field of the news.

[0053] Exemplarily, the setting of the position weight is not specifically limited in this embodiment, and those skilled in the art can freely set it according to requirements.

[0054] Exemplarily, in this embodiment, when a preset keyword appears at any position in the news text, it is regarded as a successful match. The position of the preset keyword (such as character index or paragraph number) is marked by a word segmentation tool and recorded. In this embodiment, the method for obtaining the position of the preset keyword in the news text is not specifically limited, and those skilled in the art can freely set it according to their needs.

[0055] Exemplarily, in this embodiment, the setting of preset keywords for various categories is not specifically limited, and those skilled in the art can freely set it according to their needs. For example, in the southern Fujian region, "medical insurance subsidy, affordable housing project, Taiwanese enterprise recruitment, fisherman subsidy, Jinjiang labor employment, village road hardening, tuition fee reduction..." can be used as preset keywords for people's livelihood categories, and "Xiamen-Zhangzhou Railway, Quanzhou Bridge, Zhangzhou Nuclear Power Plant, port expansion, charging pile coverage, old pipe network renovation..." can be used as preset keywords for infrastructure categories, and "Nanyin inheritance, Huian stone carving, Mid-Autumn Festival dice game, Mazu parade, Minnan language teaching materials, Tulou application for world heritage..." can be used as preset keywords for cultural categories, and "oyster carbon sequestration, mangrove restoration, water quality of the Jiulong River, sea drift garbage treatment, stone factory emissions prohibition, photovoltaic industrial park..." can be used as preset keywords for ecological categories.

[0056] Please continue to refer to Figure 1 as shown, the news data classification method further includes: Step S104, constructing a heat index of the news text based on the number of forwards, comments, likes of the news text and the geographical association characteristics of the news text, and updating the news text category index according to the heat index of the news text.

[0057] Specifically, dynamically correct the classification priority in combination with user interaction data to ensure that high-heat news is prominently displayed in classification and enhance the real-time response ability to the focus of public attention.

[0058] Please refer to Figure 3 as shown, the update method of the news text category index includes: Step S301, constructing a heat index of the news text based on the number of forwards, comments, likes of the news text and the geographical association characteristics of the news text.

[0059] Specifically, the expression of the heat index of the news text is as follows: H = ln(β1×r1 / u1 + β2×r2 / u2 + β3×r3 / u3 + 1) × tanh(3×Dy); In the formula, H is the heat index of the news text, β1 is the forward weight, β2 is the comment weight, β3 is the like weight, β1 + β2 + β3 = 1, r1 is the number of forwards, u1 is the forward threshold, r2 is the number of comments, u2 is the comment threshold, r3 is the number of likes, and u3 is the like threshold.

[0060] Specifically, by integrating interactive metrics such as reposts and comments with regional association features, a normalized popularity score is generated to balance the differences in the scale of interactions and highlight high-popularity content with local attributes.

[0061] Exemplarily, in this embodiment, the repost weight can be set to 0.3, the comment weight can be set to 0.5, and the like weight can be set to 0.2. The repost threshold can be set to 1000 times, the comment threshold can be set to 200, and the like threshold can be set to 5000 times; the specific values of the above data are not limited in this embodiment, and those skilled in the art can freely set them according to requirements.

[0062] Exemplarily, in this embodiment, the number of reposts, comments, and likes of news texts can be obtained through a crawler tool (such as the requests or Scrapy libraries in Python); the specific acquisition method is not limited in this embodiment, and those skilled in the art can freely set it according to requirements.

[0063] Please continue to refer to Figure 3 As shown, the method for updating the news text category index further includes: Step S302, updating the news text category index according to the popularity index of the news text.

[0064] Specifically, the popularity index H of the news text is compared with the popularity threshold H0. When H is greater than or equal to H0, the news text category index is updated to Sg, where Sg = Sh × γ; in the formula, γ is the update coefficient.

[0065] Specifically, the category index adjustment is triggered by the popularity threshold, so that the classification weight of highly transmissible news is dynamically increased, the exposure priority of high-quality content is improved, and the information acquisition efficiency of users or decision-makers is optimized.

[0066] Exemplarily, in this embodiment, the popularity threshold can be set to 0.6, and the update coefficient can be set to 1.2; the specific values of the above data are not limited in this embodiment, and those skilled in the art can freely set them according to requirements.

[0067] Please continue to refer to Figure 1 As shown, the news data classification method further includes: Step S105, updating the news text category based on the popularity index of the news text, the news text category index, and the policy matching feature.

[0068] Specifically, the popularity index of the news text, the news text category index, and the policy matching feature are fused to construct a confidence index for the news text. The expression of the confidence index is: ZX = (2 × News text category index × H) / {(News text category index + H) × [1 + exp(-5 × zf)]}; Compare the confidence index ZX with the confidence threshold zx0. When ZX is less than zx0, update the news text category to unclassified; otherwise, maintain the news text category.

[0069] Specifically, perform final filtering by comprehensively considering popularity, classification score, and policy relevance to eliminate misclassified results with low confidence or deviation from the focus, ensuring that the output results have accuracy, timeliness, and policy value orientation.

[0070] Exemplarily, in this embodiment, the confidence threshold can be set to 0.6. In this embodiment, no specific limitation is imposed on the value of the confidence threshold, and those skilled in the art can freely set it according to requirements.

[0071] Please refer to Figure 4 As shown, the news data classification device includes: A feature extraction unit for extracting the dialect matching feature, cultural matching feature, and policy matching feature of the collected news text; An associated feature construction unit for constructing a regional association feature based on the dialect matching feature, cultural matching feature, and policy matching feature of the news text; A classification unit for matching the news text with four preset keyword categories of people's livelihood, infrastructure, culture, and ecology to determine the news text category and the news text category index; A popularity construction unit for constructing a popularity index of the news text based on the number of forwards, comments, and likes of the news text and the regional association feature of the news text, and updating the news text category index according to the popularity index of the news text; A category update unit for updating the news text category based on the popularity index, news text category index, and policy matching feature of the news text.

[0072] The news data classification device provided by the embodiment of the present application can execute the news data classification method provided by any embodiment of the present application and has corresponding functional modules and beneficial effects for executing the method.

[0073] Please refer to Figure 5 As shown, it is a schematic structural diagram of an electronic device in this embodiment. The electronic device in the embodiment of the present invention may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, and smart home devices. Figure 5The electronic device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present invention.

[0074] As Figure 5 shown, the electronic device includes: a processor 501, a memory 502, a communication interface 503, and a system bus 504. The processor includes at least one of a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA), and is configured to call computer programs and data stored in the memory and generate control instructions; the memory includes a random access memory (RAM) and / or a non-volatile memory (NVM), and the NVM includes flash memory, a solid-state drive (SSD), or a combination thereof, and is used to store computer programs, intermediate processing data, and historical data sets; the communication interface includes a wired communication module and a wireless communication module, the wired communication module supports Ethernet or RS-485 protocols and is used to connect to a sensor network; the wireless communication module supports LoRa, 5G, or satellite communication protocols and is used to transmit processing results to a remote server; the system bus adopts a PCI Express or AXI bus architecture to achieve high-speed data interaction and clock synchronization between the processor, the memory, and the communication interface.

[0075] This embodiment further provides a computer-readable storage medium, which physically stores computer-executable instructions. When the instructions are transmitted to the processing unit via an integrated circuit substrate, they are encapsulated and processed through the data channel of the bus system and then solidified into the non-volatile storage area of the storage module. The executable instructions are configured to implement the complete technical solution of the news data classification method when executed by the processor.

[0076] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is impossible to list all the implementation manners here. Any obvious changes or variations derived from the technical solutions of the present invention still fall within the protection scope of the present invention.

Claims

1. A news data classification method, characterized in that, Including: Extracting the dialect matching features, cultural matching features, and policy matching features of the collected news text; Constructing regional association features based on the dialect matching features, cultural matching features, and policy matching features of the news text; Matching the news text with four types of preset keywords: people's livelihood, infrastructure, culture, and ecology to determine the news text category and the news text category index; Constructing the popularity index of the news text based on the number of forwards, comments, likes of the news text and the regional association features of the news text, and updating the news text category index according to the popularity index of the news text; Updating the news text category based on the popularity index of the news text, the news text category index, and the policy matching features; 2. The news data classification method according to claim 1, characterized in that, Extracting the dialect words in the news text, counting the number of dialect words in the news text as F, counting the total number of words in the news text as Fz, and setting the dialect matching feature of the news text as Mf. The expression of Mf is: ; In the formula, dk is the edit distance between the k-th dialect word in the news text and the corresponding standard word, and Dmax is the edit distance threshold.

3. The news data classification method according to claim 2, characterized in that, Counting the number of sentences containing the j-th cultural keyword in the news text as Wj, counting the total number of sentences in the news text as Wz, and setting the cultural matching feature as Mh. The expression of Mh is ; In the formula, α1 is the first weight, α2 is the second weight, α1 + α2 = 1, J is the number of cultural keywords in the news text, and Jz is the preset quantity threshold.

4. The news data classification method according to claim 3, characterized in that, Calculating the correlation between the news text and each policy document in the target area policy library, and setting the relevance Qi of the news text and the i-th policy document in the target area policy library. The expression of Qi is: ; In the formula, Ti is the number of words in the i-th policy document in the target area policy library, tf(ti) is the number of times the ti-th word in the i-th policy document in the target area policy library appears in the news text, k1 is the saturation control factor, b is the length normalization factor, G is the average number of words in all policy documents in the target area policy library, N is the average number of words in all documents in the policy library, and df(ti) is the number of policy documents in the target area policy library that contain the ti-th word in the i-th policy document; Taking the ratio of the average value of the relevance of each policy document in the target area policy library to the preset relevance threshold as the policy matching feature, denoted as zf.

5. The news data classification method according to claim 4, characterized in that, Fusing the dialect matching feature Mf, cultural matching feature Mh, and policy matching feature zf of the news text to construct the regional association feature Dy. The expression of Dy is: Dy = x1×Mf + x2×Mh + x3×zf; In the formula, x1 is the dialect weight, x2 is the cultural weight, and x3 is the policy weight.

6. The news data classification method according to claim 5, characterized in that, Matching the news text with the preset keywords of each category, calculating the category score of the news text according to the matching result, and denoting the matching score of the news text and the preset keywords of category c as Sc. The expression of Sc is: ; In the formula, Ca is the number of preset keywords of category c in the news text, and Wca is the position weight of the ca-th preset keyword of category c in the news text.

7. The news data classification method according to claim 6, characterized in that, Construct a heat index H of news text based on the number of forwards, comments, likes of the news text and the regional association characteristics of the news text. The expression of the heat index H of the news text is as follows: H = ln(β1×r1 / u1 + β2×r2 / u2 + β3×r3 / u3 + 1) × tanh(3×Dy); In the formula, H is the heat index of the news text, β1 is the forward weight, β2 is the comment weight, β3 is the like weight, β1 + β2 + β3 = 1, r1 is the number of forwards, u1 is the forward threshold, r2 is the number of comments, u2 is the comment threshold, r3 is the number of likes, and u3 is the like threshold; Compare the heat index H of the news text with the heat threshold H0. When H is greater than or equal to H0, update the news text category index to Sg.

8. The news data classification method according to claim 7, characterized in that, Fuse the heat index of the news text, the news text category index and the policy matching characteristics to construct a confidence index ZX of the news text. The expression of the confidence index ZX is: ZX = (2×news text category index×H) / {(news text category index + H)×[1 + exp(-5×zf)]}; Compare the confidence index ZX with the confidence threshold zx0. When ZX is less than zx0, update the news text category to no category. Otherwise, maintain the news text category.

9. A news data classification device, characterized in that, Including: A feature extraction unit for extracting the dialect matching feature, cultural matching feature and policy matching feature of the collected news text; An association feature construction unit for constructing regional association features based on the dialect matching feature, cultural matching feature and policy matching feature of the news text; A classification unit for matching the news text with four preset keywords of people's livelihood, infrastructure, culture, and ecology to determine the news text category and the news text category index; A heat construction unit for constructing a heat index of the news text based on the number of forwards, comments, likes of the news text and the regional association characteristics of the news text, and updating the news text category index according to the heat index of the news text; A category update unit for updating the news text category based on the heat index of the news text, the news text category index and the policy matching characteristics.

10. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the news data classification method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Policy influence analysis method and device, computer device and storage medium

    CN109635082A

  • Text content category acquisition method and device, computer equipment and storage medium

    CN111506727A

  • News recommendation method based on image-text combination

    CN120011634A

  • Speech dialect classification for automatic speech recognition

    US20120109649A1

  • Automated text-evaluation of user generated text

    US20170147682A1