News data acquisition and processing method

By constructing HTTP requests and parsing HTML content, automatically obtaining news data, storing it as a CSV file and performing word segmentation and word frequency statistics, the problem of cumbersome news data acquisition and processing process is solved, and efficient automated data processing and analysis is achieved.

CN120492616APending Publication Date: 2025-08-15李若语
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510566191.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, the process of obtaining and processing news data is cumbersome, inefficient, and complex data processing, making it difficult to achieve efficient information collection and analysis.

Method used

By constructing HTTP requests, simulating browser behavior, obtaining network news data, parsing HTML content, extracting news list and detail page data, storing it as a CSV file, and performing word segmentation and word frequency statistics, and drawing a line chart.

Benefits of technology

It realizes the automated acquisition and processing of news data, reduces manual operations, improves the efficiency of data acquisition, and supports efficient statistical analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492616A_ABST
    Figure CN120492616A_ABST
Patent Text Reader

Abstract

The invention provides a news data acquisition and processing method, which comprises the following steps of: constructing and sending an HTTP (Hyper Text Transport Protocol) request for crawling public network news data, and acquiring HTML (Hyper Text Markup Language) response content; analyzing the HTML response content, and extracting a news list according to a set extraction condition; for each piece of news, constructing a URL (Uniform Resource Locator) of a detail page, and obtaining the content of the detail page; extracting news titles, dates and text contents from the detail page contents, and writing the news titles, dates and text contents into a CSV file; and reading news data from the CSV file, performing word segmentation and word frequency statistics, and finally drawing a word frequency comparison broken line graph. According to the invention, news data can be automatically extracted and analyzed, and automatic processing of information data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of news data information technology, and in particular relates to a method for acquiring news data and processing the acquired news data. Background Art

[0002] In today's society, network technology is developing rapidly, and information technology is becoming increasingly integrated into people's lives and work. Especially for work such as academic research, information collection is extremely important, and using information technology to collect information has become an indispensable element of work.

[0003] News, as a carrier of information, is an important source of information data. However, the amount of news data online is growing rapidly. Retrieving the data you need from news is a cumbersome process, inefficient and time-consuming. Furthermore, further processing of the extracted data (such as statistical analysis) is even more complex. Therefore, achieving efficient news data acquisition and processing has become a pressing issue for information researchers. Summary of the Invention

[0004] The purpose of the present invention is to provide a news data acquisition and processing method, which can automatically extract news data and analyze it, thereby realizing automated processing of information data.

[0005] In order to achieve the above object, the technical solution of the present invention is as follows:

[0006] A news data acquisition and processing method, comprising:

[0007] S1. Build and send an HTTP request for crawling public online news data and obtain the HTML response content;

[0008] S2. Parse the HTML response content and extract the news list according to the set extraction conditions;

[0009] S3. For each news item, construct the URL of the details page and obtain the details page content;

[0010] S4. Extract the news title, date, and text from the details page and write them into a CSV file.

[0011] S5. Read news data from the CSV file, perform word segmentation and word frequency statistics, and finally draw a word frequency comparison line chart.

[0012] Furthermore, constructing the HTTP request in step S1 includes:

[0013] S101. Use the headers dictionary to define HTTP request headers to simulate browser behavior;

[0014] S102, defining the year and date for obtaining data;

[0015] S103: Construct the URL of the news page according to the year and date.

[0016] Furthermore, the sending request in step S1 includes:

[0017] Set the defined request headers and send an HTTP GET request to the constructed URL.

[0018] Furthermore, parsing the HTML response content in step S2 includes:

[0019] Use BeautifulSoup to parse the HTML response content.

[0020] Furthermore, the extraction conditions set in step S2 include:

[0021] A certain number of news items in the parsed HTML; or news items with certain keywords.

[0022] Furthermore, step S4 further includes:

[0023] Use a dictionary data structure to store the extracted news title, date, and body content.

[0024] Furthermore, step S5 further includes:

[0025] Set a stop word set, and load the stop word set for automatic filtering during word segmentation and word frequency statistics.

[0026] Furthermore, during word segmentation and word frequency statistics in step S5, the part-of-speech tagging function of Jieba word segmentation is used to retain only nouns and exclude words with a length of less than 2.

[0027] Furthermore, when drawing the word frequency comparison line graph in step S5, the N words with the highest frequency of occurrence in the specified years are counted and the line graph is drawn.

[0028] Another aspect of the present invention further provides a news data acquisition and processing device, comprising:

[0029] Request module: Builds and sends HTTP requests for crawling public online news data and obtains HTML response content;

[0030] Parsing module: parses the HTML response content and extracts the news list according to the set extraction conditions;

[0031] Details acquisition module: For each news item, construct the URL of the details page and obtain the details page content;

[0032] File module: extracts news titles, dates, and text from detail pages and writes them into CSV files;

[0033] Statistics module: reads news data from CSV files, performs word segmentation and word frequency statistics, and finally draws a word frequency comparison line chart.

[0034] Furthermore, the request module includes:

[0035] Request header unit: Use the headers dictionary to define HTTP request headers to simulate browser behavior;

[0036] Date unit: defines the year and date for obtaining data;

[0037] URL unit: Constructs the URL of a news page based on the year and date.

[0038] Furthermore, the request module also includes:

[0039] Sending unit: sets the defined request header and sends an HTTP GET request to the constructed URL.

[0040] Furthermore, the parsing module includes:

[0041] Use BeautifulSoup to parse the HTML response content.

[0042] Furthermore, the extraction conditions set in the parsing module include:

[0043] A certain number of news items in the parsed HTML; or news items with certain keywords.

[0044] Furthermore, the file module also includes:

[0045] Use a dictionary data structure to store the extracted news title, date, and body content.

[0046] Furthermore, the statistics module also includes:

[0047] Set a stop word set, and load the stop word set for automatic filtering during word segmentation and word frequency statistics.

[0048] Furthermore, when performing word segmentation and word frequency statistics in the statistics module, the part-of-speech tagging function of the Jieba word segmentation is used to retain only nouns and exclude words with a length of less than 2.

[0049] Furthermore, when drawing a word frequency comparison line graph in the statistics module, the N words with the highest frequency of occurrence within a specified period of time are counted and a line graph is drawn.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] 1. The present invention simulates browser behavior by setting an HTTP request header and sending an HTTP request, thereby achieving smooth automatic acquisition of news data.

[0052] 2. The present invention first obtains the news list and then constructs the URL of the news details page to obtain the news details, so that the obtained news data meets a clear goal and avoids data redundancy.

[0053] 3. The present invention forms the acquired news data into a CSV file for word segmentation and word frequency statistics, thereby realizing the processing process of automated statistical analysis of news data. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 2 is a flow chart of automatic acquisition of news data according to embodiment 1 of the present invention.

[0055] Figure 2 1 is a flow chart of automatic statistical analysis of news data according to embodiment 1 of the present invention.

[0056] Figure 3 It is a word frequency comparison line chart of Example 1 of the present invention. DETAILED DESCRIPTION

[0057] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0058] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0059] Example 1:

[0060] The news data acquisition and processing method proposed in this embodiment is mainly divided into two parts, one of which is a method for automatically acquiring news data, and the other is a method for automatically processing the acquired news data.

[0061] 1: Automatic acquisition method of news data.

[0062] Take the news data acquisition of a daily newspaper website as an example. Figure 1 The process of automatically acquiring news data includes:

[0063] 1. Send an HTTP request, specifically:

[0064] Set request headers: This embodiment uses the headers dictionary to define HTTP request headers, simulating browser behavior for better access through the website.

[0065] Define a year list: This embodiment uses yearList, which contains the years to be crawled, such as 2024.

[0066] Define a crawling function to construct the URL of a news page based on the year and date, for example, January 2024.

[0067] Send a request: Use requests.get to send an HTTP GET request to the constructed URL.

[0068] 2. Parse the HTML document and extract the required data, specifically:

[0069] Sending an HTTP GET request will get the HTML response content of the website response, and you can use BeautifulSoup to parse the HTML response content.

[0070] Extract news list: extract news list from the parsed HTML. The extraction can be performed based on set keywords or a set number. In this embodiment, only the first 5 news items are extracted based on the number.

[0071] 3. Get the details page content, specifically:

[0072] For the five extracted news items, construct the URL of the detail page for each piece of news and send a request to obtain the content of the detail page.

[0073] The required data content is extracted from the acquired details page content. In this embodiment, the title, date and text content of the news are extracted.

[0074] The extracted data information (title, date, and text content) is stored in a data dictionary and the yield keyword is used to generate this dictionary so that it can be iterated outside the function.

[0075] 4. Write to CSV file, specifically:

[0076] Open a CSV file through the script and use csv.DictWriter to write the obtained data to the file.

[0077] In the above process, the try-except structure can be used to handle possible exceptions, such as network request failure or parsing error.

[0078] In order to avoid frequent requests and data interaction affecting the target news website, time.sleep(M) is used in this embodiment to set a pause of M seconds after each request.

[0079] Through the above method, we can automatically obtain news data from the Internet and save the obtained content into a CSV file for subsequent use.

[0080] 2. Automatic processing method after news data acquisition.

[0081] This method is based on the CSV file generated above, such as Figure 2 As shown, specifically including:

[0082] 1. Set Matplotlib Configuration: Matplotlib configuration involves adjusting the style, size, color, and other properties of a plot by setting various parameters and options of the Matplotlib drawing library to meet different visualization requirements. In this example, we mainly configure Chinese fonts and solve the minus sign display problem.

[0083] 2. Define the function load_stopwords to load stop words: Stop words are common, high-frequency words that are automatically filtered out to save storage space and improve efficiency in information retrieval or natural language processing. This step reads stop words from a commonly used "stop word" file and generates a stop word set to return.

[0084] 3. Define word segmentation and count word frequency function extract_nouns: use the part-of-speech tagging function of Jieba word segmentation to retain only nouns and exclude stop words and words with a length of less than 2.

[0085] 4. Define the read_csv function to read CSV files: read news data from CSV files and store them by year.

[0086] 5. Define the function plot_word_frequency_lineplot to count word frequencies and draw a line graph: count the 20 most common words in January 2024 and use matplotlib to draw a line graph.

[0087] 6. Based on the above configuration and functions, build the main program main. Its executable functions include loading stop words, reading news data, segmenting the news data and counting word frequencies, and finally drawing a word frequency comparison line chart.

[0088] 7. Execute the built main program main.

[0089] Through the above method, the automatic processing of news data after acquisition can be achieved. Figure 3 Shown is a line chart of news keyword statistics for January 2024 drawn after executing the above method.

[0090] Example 2:

[0091] This embodiment provides a news data acquisition and processing device, including:

[0092] Request module: Builds and sends HTTP requests for crawling public online news data and obtains HTML response content;

[0093] Parsing module: parses the HTML response content and extracts the news list according to the set extraction conditions;

[0094] Details acquisition module: For each news item, construct the URL of the details page and obtain the details page content;

[0095] File module: extracts news titles, dates, and text from detail pages and writes them into CSV files;

[0096] Statistics module: reads news data from CSV files, performs word segmentation and word frequency statistics, and finally draws a word frequency comparison line chart.

[0097] The request module includes:

[0098] Request header unit: Use the headers dictionary to define HTTP request headers to simulate browser behavior;

[0099] Date unit: defines the year and date for obtaining data;

[0100] URL unit: Constructs the URL of a news page based on the year and date.

[0101] The request module also includes:

[0102] Sending unit: sets the defined request header and sends an HTTP GET request to the constructed URL.

[0103] The parsing module includes:

[0104] Use BeautifulSoup to parse the HTML response content.

[0105] The extraction conditions set in the parsing module include:

[0106] A certain number of news items in the parsed HTML; or news items with certain keywords.

[0107] The file module also includes:

[0108] Use a dictionary data structure to store the extracted news title, date, and body content.

[0109] The statistics module also includes:

[0110] Set a stop word set, and load the stop word set for automatic filtering during word segmentation and word frequency statistics.

[0111] When performing word segmentation and word frequency statistics in the statistics module, the part-of-speech tagging function of the Jieba word segmentation is used to retain only nouns and exclude words with a length of less than 2.

[0112] When drawing a word frequency comparison line graph in the statistics module, count the N words with the highest frequency in the specified years and draw a line graph.

[0113] The news data acquisition and processing device proposed in this embodiment can implement the news data acquisition and processing method proposed in Example 1, and has the same technical effects as Example 1.

[0114] The above-described embodiments are merely preferred implementations of the present invention and are intended to help understand the method and core concepts of the present application. The scope of protection of the present invention is not limited to the above-described embodiments. All technical solutions within the scope of protection of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A news data acquisition and processing method, characterized in that: include: S1. Build and send an HTTP request for crawling public online news data and obtain the HTML response content; S2. Parse the HTML response content and extract the news list according to the set extraction conditions; S3. For each news item, construct the URL of the details page and obtain the details page content; S4. Extract the news title, date, and text from the details page and write them into a CSV file. S5. Read news data from the CSV file, perform word segmentation and word frequency statistics, and finally draw a word frequency comparison line chart.

2. The news data acquisition and processing method according to claim 1, characterized in that: Constructing the HTTP request in step S1 includes: S101. Use the headers dictionary to define HTTP request headers to simulate browser behavior; S102, defining the year and date for obtaining data; S103: Construct the URL of the news page according to the year and date.

3. The news data acquisition and processing method according to claim 2, characterized in that: The sending request in step S1 includes: Set the defined request headers and send an HTTP GET request to the constructed URL.

4. The news data acquisition and processing method according to claim 1, characterized in that: The HTML response content parsed in step S2 includes: Use BeautifulSoup to parse the HTML response content.

5. The news data acquisition and processing method according to claim 1, characterized in that: The extraction conditions set in step S2 include: A certain number of news items in the parsed HTML; or news items with certain keywords.

6. The news data acquisition and processing method according to claim 1, characterized in that: Step S4 further includes: Use a dictionary data structure to store the extracted news title, date, and body content.

7. The news data acquisition and processing method according to claim 1, characterized in that: Step S5 also includes: Set a stop word set, and load the stop word set for automatic filtering during word segmentation and word frequency statistics.

8. The news data acquisition and processing method according to claim 1, characterized in that: During word segmentation and word frequency statistics in step S5, the part-of-speech tagging function of Jieba word segmentation is used to retain only nouns and exclude words with a length of less than 2.

9. The news data acquisition and processing method according to claim 1, characterized in that: When drawing the word frequency comparison line graph in step S5, the N words with the highest frequency of occurrence within the specified years are counted and the line graph is drawn.

10. A news data acquisition and processing device, characterized in that: include: Request module: Builds and sends HTTP requests for crawling public online news data and obtains HTML response content; Parsing module: parses the HTML response content and extracts the news list according to the set extraction conditions; Details acquisition module: For each news item, construct the URL of the details page and obtain the details page content; File module: extracts news titles, dates, and text from detail pages and writes them into CSV files; Statistics module: reads news data from CSV files, performs word segmentation and word frequency statistics, and finally draws a word frequency comparison line chart.