News page data general acquisition method, device, equipment, medium and product

By using Selenium for dynamic content capture in the news page data collection, combining text density and symbol density analysis, and using large language models for key information extraction and data cleaning, the problem of lack of flexibility and versatility in the existing technology is solved, and efficient and accurate news page data collection is achieved.

CN120045767APending Publication Date: 2025-05-27SICHUAN COVER MEDIA TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510209720.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing news page data collection solutions lack flexibility and versatility, and cannot efficiently handle dynamic content, advertising and multimedia elements, making it cumbersome to obtain data on different platforms.

Method used

A general collection method for news page data is adopted. By sending an HTTP request to the target news website, it determines whether there is dynamic loading content. If it exists, Selenium is used for data crawling. Then, the text density and symbol density in the DOM tree are calculated, the content complexity is judged, and the key information extraction and data cleaning are extracted using multi-dimensional feature analysis and large language models.

Benefits of technology

It achieves efficient and accurate extraction of key information such as news titles, texts, authors and release moments from various types of news pages, improving the efficiency and accuracy of data collection and adapting to a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045767A_ABST
    Figure CN120045767A_ABST
Patent Text Reader

Abstract

The invention discloses a general news page data acquisition method, device and equipment, a medium and a product, and relates to the technical field of news page data acquisition. The method comprises the following steps: sending an HTTP (Hyper Text Transport Protocol) request to a target news website to obtain return data of a news webpage, calling a browser automation tool Selenium to perform data capture after all elements of the news webpage are loaded when judging that dynamic loading contents exist, and taking a capture result as original data of the news webpage, then, for each node in the DOM tree, calculating to obtain corresponding text density and symbol density, judging whether the webpage content is complex content or not based on a calculation result, and if the webpage content is complex content, analyzing the webpage content through multi-dimensional feature analysis and a big language model used for news page analysis based on rules; and positioning to obtain a final extraction result aiming at the news page key information, and finally carrying out data cleaning and standardization processing on the extraction result to obtain news page data with a uniform format and outputting the news page data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of news page data collection, and particularly relates to a general news page data collection method, device, equipment, medium and product. Background Art

[0002] With the development of the Internet and the rapid increase in the amount of information, the display content of news pages has become increasingly rich, and news media are facing huge challenges in collecting a large amount of news content. Traditional news page data collection solutions often rely on manual page parsing, which is inefficient and error-prone.

[0003] Currently, although there are some automated collection tools on the market that can be used for news page data collection, they often lack the ability to adapt to complex page structures and are mostly targeted at specific websites or formats. Therefore, these automated collection tools lack flexibility and generality and cannot efficiently handle dynamic content, advertisements, and multimedia elements, etc., making it cumbersome to obtain data on different platforms.

[0004] Therefore, how to provide a new general news page data collection solution to efficiently and accurately extract key news page information such as news titles, news texts, news authors, and / or news release times from various types of news pages is an urgent research topic for those skilled in the art. Summary of the Invention

[0005] The object of the present invention is to provide a general news page data collection method, device, computer equipment, computer-readable storage medium, and computer program product to solve the problem that it is cumbersome to obtain data on different platforms due to the lack of flexibility and generality of existing news page data collection solutions.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, a general news page data collection method is provided, including:

[0008] Sending an HTTP request to a target news website to obtain news web page return data;

[0009] Judging whether there is dynamically loaded content on the news page according to the news web page return data;

[0010] If it is determined that there is dynamically loaded content, calling the browser automation tool Selenium to perform data scraping after all elements on the news web page are loaded, and using the scraping result as the original news web page data;

[0011] Performing noise filtering processing on the original news web page data to obtain a DOM tree;

[0012] For each node in the DOM tree, calculate the corresponding text density and symbol density;

[0013] Based on the text density and symbol density of each node, determine whether the web page content is complex content;

[0014] If it is determined to be complex content, first, based on the text density and symbol density of each node, locate the preliminary extraction result of the key information of the news page from the DOM tree through multi-dimensional feature analysis, and then import the preliminary extraction result into a large language model based on rules and used for news page parsing, so that the large language model can perform in-depth semantic understanding and analysis on the preliminary extraction result according to the pre-constructed prompt word rule set, output the CSS selector corresponding to the key information of the news page, and apply the CSS selector to locate the final extraction result of the key information of the news page from the DOM tree, where the key information of the news page includes news title, news text, news author, and / or news release time;

[0015] Perform data cleaning and standardization processing on the final extraction result to obtain news page data with a unified format and output it.

[0016] Based on the above invention content, a new general news page data collection scheme is provided, that is, first send an HTTP request to the target news website to obtain the news web page return data, then when it is determined that there is dynamically loaded content, call the browser automation tool Selenium to perform data scraping after all elements of the news web page are loaded, and use the scraping result as the original news web page data. Then, for each node in the DOM tree, calculate the corresponding text density and symbol density, and based on the calculation results, determine whether the web page content is complex content. If so, through multi-dimensional feature analysis and a large language model based on rules and used for news page parsing, locate the final extraction result of the key information of the news page. Finally, perform data cleaning and standardization processing on the extraction result to obtain news page data with a unified format and output it. In this way, it is possible to efficiently and accurately extract key information of news pages such as news titles, news texts, news authors, and / or news release times from various types of news pages, which is convenient for practical application and promotion.

[0017] In a possible design, sending an HTTP request to the target news website includes:

[0018] Real-time monitor and analyze the historical access data, current network status, and recent request success rate of the target news website, and select a suitable user agent from the user agent library according to the analysis results to send an HTTP request to the target news website;

[0019] And / or, during the request process, select a secure IP address from the proxy IP pool according to the IP ban monitoring data of the target news website and the risk assessment result of the current HTTP request to replace the source address in the current HTTP request;

[0020] And / or, during the request process, generate an appropriate delay time according to the analysis results of the request response time distribution of the target news website, the current request traffic, and the importance of the target page, and send two adjacent HTTP requests to the target news website at intervals.

[0021] In a possible design, for each node in the DOM tree, calculate the corresponding text density, including:

[0022] For each node in the DOM tree, calculate the corresponding text density according to the following formula:

[0023]

[0024] In the formula, i represents a positive integer, TD i represents the text density of the i-th node in the DOM tree, T i represents the number of strings of the i-th node, LT i represents the number of strings of the links carried by the i-th node, TG i represents the number of tags of the i-th node, LTG i represents the number of tags of the links carried by the i-th node.

[0025] In a possible design, for each node in the DOM tree, calculate the corresponding text density and symbol density, including:

[0026] According to the overall layout style type, section division characteristic type of the target page, and historical parsing data, use a machine learning model pre-trained based on an artificial intelligence algorithm to dynamically adjust the required parameters and / or required thresholds for calculating text density and / or symbol density;

[0027] For each node in the DOM tree, apply the required parameters and / or the required thresholds to calculate the corresponding text density and symbol density.

[0028] In a possible design, perform data cleaning and standardization processing on the final extraction result, including:

[0029] Comprehensively analyze multi-dimensional information to identify redundant information in the final extraction result, and remove the redundant information from the final extraction result, where the multi-dimensional information includes text content, HTML structure, data semantic association features, and / or visual presentation features;

[0030] Perform unified format conversion processing on the dates, times, and / or numbers in the final extraction result to obtain dates, times, and / or numbers with a unified format.

[0031] In a possible design, after obtaining the DOM tree, the method further includes:

[0032] Adopt a multimedia element link recognition method based on multimedia format and / or multimedia encoding method to locate and extract the embedded links of multimedia elements from the DOM tree;

[0033] Classify, store, and manage the embedded links according to the type, source, importance of the multimedia elements, and / or the relevance of the multimedia elements to the news page data.

[0034] In a second aspect, a general news page data acquisition device is provided, including a data request unit, a first judgment unit, a data scraping unit, a filtering and processing unit, a density calculation unit, a second judgment unit, a page parsing unit, and a data processing unit that are sequentially communicatively connected;

[0035] The data request unit is used to send an HTTP request to a target news website to obtain news web page return data;

[0036] The first judgment unit is used to judge whether there is dynamically loaded content on the news page according to the news web page return data;

[0037] The data scraping unit is used to, when it is determined that there is dynamically loaded content, call the browser automation tool Selenium to scrape data after all elements on the news web page are loaded, and use the scraping result as the original news web page data;

[0038] The filtering and processing unit is used to perform noise filtering processing on the original news web page data to obtain a DOM tree;

[0039] The density calculation unit is used to calculate the corresponding text density and symbol density for each node in the DOM tree;

[0040] The second judgment unit is used to judge whether the web page content is complex content according to the text density and symbol density of each node;

[0041] The page parsing unit is configured to, when determining that the content is complex, first locate a preliminary extraction result of the key information of the news page from the DOM tree through multi-dimensional feature analysis based on the text density and symbol density of each node, and then import the preliminary extraction result into a large language model based on rules and used for news page parsing, so that the large language model performs in-depth semantic understanding and analysis on the preliminary extraction result according to a pre-constructed set of prompt rules, outputs a CSS selector corresponding to the key information of the news page, and applies the CSS selector to locate a final extraction result of the key information of the news page from the DOM tree, where the key information of the news page includes a news title, a news body, a news author, and / or a news release time;

[0042] The data processing unit is configured to perform data cleaning and standardization processing on the final extraction result to obtain news page data with a unified format and output it.

[0043] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a transceiver that are communicatively connected in sequence, where the memory is configured to store a computer program, the transceiver is configured to send and receive messages, and the processor is configured to read the computer program and execute the general news page data acquisition method as described in the first aspect or any possible design in the first aspect.

[0044] In a fourth aspect, the present invention provides a computer-readable storage medium, on which instructions are stored, and when the instructions are run on a computer, the general news page data acquisition method as described in the first aspect or any possible design in the first aspect is executed.

[0045] In a fifth aspect, the present invention provides a computer program product, including a computer program or instructions, and when the computer program or the instructions are executed by a computer, the general news page data acquisition method as described in the first aspect or any possible design in the first aspect is implemented.

[0046] Beneficial effects of the above solution:

[0047] (1) The present invention provides a new general solution for collecting news page data. First, an HTTP request is sent to the target news website to obtain the news web page return data. Then, when it is determined that there is dynamically loaded content, the browser automation tool Selenium is called to perform data scraping after all elements on the news web page are loaded, and the scraping result is used as the original data of the news web page. Then, for each node in the DOM tree, the corresponding text density and symbol density are calculated, and based on the calculation results, it is determined whether the web page content is complex content. If so, through multi-dimensional feature analysis and a large language model based on rules and used for news page parsing, the final extraction result of the key information for the news page is located. Finally, data cleaning and standardization processing are performed on the extraction result to obtain news page data with a unified format and output it. In this way, key information of news pages such as news titles, news texts, news authors, and / or news release times can be efficiently and accurately extracted from various types of news pages, facilitating practical applications and promotions;

[0048] (2) By classifying and storing multimedia links efficiently and managing them, seamless docking with external multimedia processing tools and platforms can be achieved, facilitating further processing, analysis, and application of multimedia elements in the future;

[0049] (3) It can quickly and accurately collect and process news manuscripts from the original web page: by optimizing the text and symbol density algorithms and combining the precise parsing ability of the large language model, it can effectively identify the key information in news pages with different structures, and has a powerful data cleaning and standardization ability to ensure the consistency and high quality of the output data, adapting to various application scenarios, providing an efficient and reliable solution for news media and related fields, and significantly improving the collection efficiency and accuracy of news data;

[0050] (4) A full-link efficient collection system from web page request to data cleaning is constructed: by integrating advanced anti-crawling technology, precise page parsing technology, efficient data cleaning and standardization technology, and intelligent multimedia processing technology, it can quickly and accurately collect news pages and output structured content, achieving the convenient goal of "one-key collection";

[0051] (5) Innovatively combines the text and symbol density algorithms with the large language model to form a precise data parsing mode: giving full play to the efficiency of the algorithm in basic structure parsing and the advantages of the model in complex logic processing, through a carefully designed guiding strategy, it can achieve in-depth parsing and precise information extraction of different types of news pages, effectively breaking through the limitations of traditional parsing methods in dealing with complex page structures and dynamic content;

[0052] (6) An adaptive data cleaning and standardization system and an intelligent multimedia content processing system have been developed, which can dynamically adjust data processing strategies and output formats according to different news sources, application scenarios, and user requirements, ensuring that the collected data is accurate, standardized, and can meet diverse usage and analysis needs, greatly enhancing the practicality and adaptability.

[0053] (7) The full-link intelligent collection of news pages has been successfully realized: The optimized text and symbol density algorithm and the accurate parsing ability of the large language model effectively identify the key information in news pages with different structures; The intelligent data cleaning and standardization system and the multimedia content processing system ensure the high efficiency, accuracy, and high quality of the output data, enabling it to widely adapt to various application scenarios, providing an extremely innovative and reliable solution for news media and related fields, significantly improving the collection efficiency and accuracy of news data, and strongly promoting the progress and development of news information processing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0055] Figure 1 It is a schematic flowchart of the general news page data collection method provided by the embodiments of the present application.

[0056] Figure 2 It is a schematic flowchart of data cleaning and standardization processing in the general news page data collection method provided by the embodiments of the present application.

[0057] Figure 3 It is a schematic flowchart of multimedia element link recognition and storage in the general news page data collection method provided by the embodiments of the present application.

[0058] Figure 4 It is a schematic structural diagram of the general news page data collection device provided by the embodiments of the present application.

[0059] Figure 5 It is a schematic structural diagram of the computer device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and the description of the embodiments or the prior art. Obviously, the following description of the structures of the drawings is only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these embodiments. It should be noted here that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation to the present invention.

[0061] It should be understood that although terms such as first and second may be used herein to describe various objects, these objects should not be limited by these terms. These terms are only used to distinguish one object from another. For example, the first object may be referred to as the second object, and similarly, the second object may be referred to as the first object, without departing from the scope of the exemplary embodiments of the present invention.

[0062] It should be understood that for the term "and / or" that may appear in this text, it is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, B exists alone, or A and B exist simultaneously; another example, A, B and / or C can represent any one of A, B and C or any combination of them; for the term " / and" that may appear in this text, it is a description of another association object relationship, indicating that there can be two relationships. For example, A / and B can represent: A exists alone or A and B exist simultaneously; in addition, for the character " / " that may appear in this text, generally it represents that the associated objects before and after are an "or" relationship.

[0063] Embodiment

[0064] As Figures 1-2 shown, the general news page data collection method provided in the first aspect of this embodiment can be, but is not limited to, executed by a computer device with certain computing resources, such as a cloud server, a personal computer (Personal Computer, PC, referring to a multi-purpose computer suitable for personal use in terms of size, price and performance; desktop computers, laptops, small laptops, tablet computers and ultrabooks, etc. all belong to personal computers), a smart phone, a personal digital assistant (Personal Digital Assistant, PDA) or a wearable device and other electronic devices. As Figure 1 shown, the general news page data collection method can be, but is not limited to, including the following steps S1 to S8.

[0065] S1. Send an HTTP (Hypertext Transfer Protocol) request to the target news website to obtain the news web page return data.

[0066] In the step S1, the HTTP request is a commonly used existing request method for accessing a website and obtaining web page content. The obtained news web page return data specifically includes, but is not limited to, content such as HTML (Hyper Text Markup Language) code. Considering that the target news website may run an anti-crawler mechanism, in order to significantly reduce the probability of being identified as a crawler, preferably, sending an HTTP request to the target news website includes, but is not limited to: real-time monitoring and analysis of the historical access data of the target news website, the current network status, and the recent request success rate, and selecting a suitable user agent from the user agent library according to the analysis results to send an HTTP request to the target news website. The user agent library has a rich and diverse user agent pool (that is, it contains multiple User-Agent headers to simulate different device and browser types), which needs to be constructed by deeply analyzing the user agent characteristics of many devices and browsers in advance, so that before each HTTP request is sent to the target news website, based on multiple factors such as the website's historical access data, the current network status, and the recent request success rate for comprehensive judgment (the specific comprehensive judgment and analysis process can be routinely deduced based on existing technical means), randomly select a user agent with a high degree of adaptability from the library (that is, randomly select a User-Agent header) to reduce the probability of being identified as a crawler. For example, if the target news website has restrictions on accessing a specific browser type, other unrestricted user agents with a good access record will be preferentially selected for the request.

[0067] In step S1, considering that the target news website may also implement IP (Internet Protocol Address) blocking and access frequency limitation policies, in order to effectively avoid IP blocking and access frequency limitation, preferably, an HTTP request is sent to the target news website, including but not limited to: during the request process, based on the IP blocking monitoring data of the target news website and the risk assessment result of the current HTTP request, a secure IP address is selected from the proxy IP pool to replace the source address in the current HTTP request. The proxy IP pool has a number of proxy IPs obtained through conventional intelligent algorithms and with high quality, so that by continuously replacing the IP address, each request appears to come from different devices (i.e., ensuring the diversity and security of the request source), effectively avoiding IP blocking and access frequency limitation. For example, if it is found that a certain IP address has abnormal requests to the target news website recently, the standby secure IP address can be quickly switched for subsequent requests.

[0068] In step S1, in order to prevent the anti-crawling mechanism from being triggered due to excessive request frequency, preferably, an HTTP request is sent to the target news website, including but not limited to: during the request process, based on the analysis results of the request response time distribution of the target news website, the current request traffic, and the importance of the target page, a suitable delay time is generated to send two adjacent HTTP requests to the target news website at intervals. In this way, by adding a suitable time delay between two consecutive requests (the specific delay duration can be randomly generated), it is possible to simulate human browsing behavior and prevent the anti-crawling mechanism from being triggered due to excessive request frequency. For example, for the hot news page of a popular news website, the delay time is appropriately extended during peak traffic periods to simulate natural human browsing behavior and avoid triggering the anti-crawling mechanism.

[0069] S2. Determine whether there is dynamically loaded content on the news page according to the data returned by the news web page.

[0070] In step S2, considering that some websites may load data content in the way of dynamic Js (JavaScript, a dynamic programming language widely used in web development, mainly used to enhance the interactivity and dynamics of web pages), resulting in incomplete data in the data returned by the news web page, so in order to ensure the integrity of the obtained web page content, it is necessary to make the above judgment. In addition, the specific process of determining whether there is dynamically loaded content on the news page can be derived conventionally based on existing technical means.

[0071] S3. If it is determined that there is dynamically loaded content, call the browser automation tool Selenium to perform data scraping after all elements on the news web page are loaded, and use the scraping result as the original data of the news web page.

[0072] In the step S3, the browser automation tool Selenium is an existing tool. Specifically, by setting implicit waiting (i.e., implicitly_wait()), it can be ensured that data scraping (i.e., reading the HTML code) is performed after all elements on the page are loaded, so as to still extract the required information when facing asynchronous loading and content generated by JavaScript. In addition, if it is determined that there is no dynamically loaded content, the returned data of the news web page can be directly used as the original data of the news web page.

[0073] S4. Perform noise filtering processing on the original data of the news web page to obtain a DOM tree.

[0074] In the step S4, the DOM tree (Document Object Model Tree) is an important data structure used by the browser to represent and render web pages; the DOM tree represents the content of an HTML document as a node tree, and each node represents a part of the document, such as a tag, an attribute, or text; the structure of the DOM tree includes a root node, element nodes, attribute nodes, and text nodes, etc. In addition, the specific process of the noise filtering processing is as follows: clear the data noise, redundant JavaScript code, and comment content on the page.

[0075] S5. For each node in the DOM tree, calculate the corresponding text density and symbol density.

[0076] In the step S5, since different key information on news web pages, such as news text, news titles, news authors, and news release times, has different text densities and symbol densities, calculating the text density and symbol density of each node can be used for subsequent positioning and extraction of key information on news web pages. Specifically, for each node in the DOM tree, the corresponding text density is calculated, including but not limited to: for each node in the DOM tree, the corresponding text density is calculated according to the following formula:

[0077]

[0078] In the formula, i represents a positive integer, TD i represents the text density of the i-th node in the DOM tree, T i represents the number of strings of the i-th node, LT i represents the number of strings of the links carried by the i-th node, TGi represents the number of tags of the i-th node, LTG i represents the number of tags of the links carried by the i-th node. The aforementioned TD i As a value for measuring the text density of a node in a web page, if this value is large, it means that the unformatted and unlinked text of this node must be more than the formatted and linked text. Then it proves that the possibility of this node belonging to the body content is relatively large, so it is conducive to clearly judging which node is useful body content. In addition, when calculating the text density, in addition to considering the ratio of the number of pure text words to the number of linked text words, visual features such as the font style, color, and / or size of the text can also be additionally analyzed for subsequent positioning and extraction of key information on the news page. In addition, the specific calculation method of the symbol density is similar to the aforementioned text density calculation method, and a comprehensive evaluation can be carried out according to the weight difference of the symbols in different contexts.

[0079] In the step S5, in order to adapt to the page structure, preferably, for each node in the DOM tree, the corresponding text density and symbol density are calculated, including but not limited to: first, according to the overall layout style type of the target page, the block division characteristic type, and the historical parsing data, the machine learning model pre-trained by using the artificial intelligence algorithm is used to dynamically adjust the required parameters and / or required thresholds for calculating the text density and / or symbol density; then, for each node in the DOM tree, the required parameters and / or the required thresholds are applied to calculate the corresponding text density and symbol density. The aforementioned overall layout style type is specifically but not limited to single-column or multi-column layout, etc., and the aforementioned block division characteristic type is specifically but not limited to layouts such as the header navigation, the main text area, or the sidebar. The artificial intelligence algorithm is a core artificial intelligence algorithm that specifically studies how a computer simulates or implements human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance, and is the fundamental way to make a computer intelligent; specifically, the artificial intelligence algorithm preferably adopts the linear regression algorithm based on the Python sklearn library to quickly and accurately find the rules in the data. The machine learning model needs to be pre-trained through a conventional calibration and verification modeling process based on a certain amount of sample data (the model input items are the overall layout style type, the block division characteristic type, and the historical parsing data, and the model output items are the empirical values of the required parameters and / or required thresholds) (the specific process includes the calibration process and the verification process of the model, that is, first, by comparing the model simulation results with the measured data, and then adjusting the model parameters according to the comparison results to make the simulation results coincide with the actual situation). Specifically, in the calibration and verification modeling process of the machine learning model, the Bayesian optimization algorithm based on the tree structure is used to optimize the model parameters. The aforementioned dynamic adjustment of the required parameters and / or required thresholds for calculating the text density and / or symbol density can be exemplified as follows: for a news page with a multi-column layout and many sidebar advertisements, the symbol density calculation parameters can be appropriately adjusted to more accurately locate the main text area. In addition, the artificial intelligence algorithm can also but not limited to adopt machine learning algorithms based on support vector machines, K-nearest neighbor methods, stochastic gradient descent methods, multivariate linear regression, multi-layer perceptrons, decision trees, backpropagation neural networks, or radial basis function networks, etc.

[0080] S6. Determine whether the web page content is complex content according to the text density and symbol density of each node.

[0081] In the step S6, specifically, based on the high and low levels of the text density and symbol density, the web page content can be determined whether it is complex content through a conventional threshold comparison method. For example, when the density sum of the text density and symbol density of each node exceeds the preset density threshold, it is determined that the web page content is complex content, otherwise it is not.

[0082] S7. If it is determined to be complex content, first, based on the text density and symbol density of each node, the preliminary extraction result of the key information of the news page is located from the DOM tree through multi-dimensional feature analysis. Then, the preliminary extraction result is imported into a large language model based on rules and used for news page parsing, so that the large language model can perform in-depth semantic understanding and analysis on the preliminary extraction result according to the pre-constructed prompt word rule set, output the CSS selector corresponding to the key information of the news page, and apply the CSS selector to locate the final extraction result of the key information of the news page from the DOM tree. Among them, the key information of the news page includes but is not limited to news title, news text, news author, and / or news release time, etc.

[0083] In the step S7, since different key information of news pages, such as news text, news titles, news authors, and news release times, has different text densities and symbol densities (specifically, when the number of pure text characters in a certain node is significantly more than the number of text characters containing links, then this node has a higher text density; at the same time, the punctuation symbol density of the text content is usually significantly higher than that in areas such as links, advertisement information, and navigation bars; based on this characteristic, considering the symbol density can effectively locate the news text content on the page and improve the accuracy of extraction), therefore, based on the text density and symbol density of each of the nodes, a preliminary extraction result for the key information of the news page can be located from the DOM tree through conventional multi-dimensional feature analysis. The large language model (LLM) refers to a deep learning model trained using a large amount of text data, enabling the model to generate natural language text or understand the meaning of language text; these models can provide in-depth knowledge and language production on various topics through training on a large dataset, and its core idea is to learn the patterns and structures of natural language through large-scale unsupervised training and to simulate the human language cognition and generation process to a certain extent. Therefore, the large language model can be used as a news page parsing framework obtained by deep customization and optimization for the news page parsing task after conventional pre-training. This framework can take the preliminary extraction result as input and conduct in-depth semantic understanding and analysis by combining multi-source data such as the HTML structure features, semantic information, and historical parsing cases of the page, so as to utilize the powerful language understanding and reasoning ability of the model to accurately identify titles, text, and other key information, and output highly accurate and optimized CSS selectors, effectively improving the accuracy and efficiency of the page structure parsing result, that is, in the case where the multi-dimensional feature analysis cannot cover, a pre-trained large model can be used as a fallback solution. The prompt rule set can be pre-constructed according to the common structural features, keyword distribution rules, and semantic logic of news pages to ensure that the CSS selectors output by the model have high accuracy and pertinence, effectively improving the parsing accuracy, that is, by guiding the large model to return the CSS selectors of the corresponding areas through prompts, unnecessary generated content can be reduced, thereby reducing the session Token consumption and network waiting time, especially applicable to scenarios with complex page structures and dynamic content loading, and can further improve the parsing accuracy of the news text content.The CSS selector is a pattern used to match elements in an HTML document. Through these selectors, it is possible to specify which elements should apply specific style rules. Furthermore, the CSS selector can be applied to locate the final extraction result of the key information of the news page from the DOM tree. For example, when encountering a complex nested table structure for presenting news content, the large language model can identify information such as titles and main texts in the table according to the prompt word rule set and output accurate CSS selectors to extract news titles and news texts, etc. In addition, if it is determined that the content is not complex, the final extraction result of the key information of the news page can be directly located from the DOM tree through multi-dimensional feature analysis based on the text density and symbol density of each node.

[0084] S8. Perform data cleaning and standardization processing on the final extraction result to obtain news page data with a unified format and output it.

[0085] In step S8, as Figure 2 shown, specifically, performing data cleaning and standardization processing on the final extraction result includes, but is not limited to, the following steps S81 - S82.

[0086] S81. Comprehensively analyze multi-dimensional information to identify redundant information in the final extraction result and remove the redundant information from the final extraction result. Among them, the multi-dimensional information includes, but is not limited to, text content, HTML structure, data semantic association features, and / or visual presentation features, etc.

[0087] In step S81, the redundant information specifically includes, but is not limited to, noise, advertisements, navigation bars, and javascripts in the page, etc., which can be conventionally identified by existing technical means and removed. For example, for advertisement information, it can be identified through advertisement keywords in the text and accurately judged and removed according to its visual position in the page (such as the sidebar or pop-up window, etc.).

[0088] S82. Perform unified format conversion processing on the dates, times, and / or numbers in the final extraction result to obtain dates, times, and / or numbers with a unified format.

[0089] In step S82, for example, for the date format, various formats such as "YYYY - MM - DD" and / or "MM / DD / YYYY" can be recognized and converted to the unified "YYYY - MM - DD" format according to predefined standards. In addition, the field names and content structure rules of the output data can be adaptively adjusted using field rules according to factors such as the application field of the target data, user demand preferences, and data association relationships.

[0090] Based on the general news page data acquisition method described in the foregoing steps S1 to S8, a new general news page data acquisition solution is provided, that is, first send an HTTP request to the target news website to obtain the news web page return data, and then when it is determined that there is dynamically loaded content, call the browser automation tool Selenium to perform data scraping after all elements on the news web page are loaded, and use the scraping result as the original data of the news web page. Then, for each node in the DOM tree, calculate the corresponding text density and symbol density, and based on the calculation results, determine whether the web page content is complex content. If so, through multi-dimensional feature analysis and a large language model based on rules and used for news page parsing, locate the final extraction result of the key information of the news page. Finally, perform data cleaning and standardization processing on the extraction result to obtain news page data with a unified format and output it. In this way, key information of news pages such as news titles, news texts, news authors, and / or news release times can be efficiently and accurately extracted from various types of news pages, which is convenient for practical application and promotion.

[0091] Based on the technical solution of the foregoing first aspect, this embodiment also provides a possible design for how to identify and store multimedia element links, that is, as Figure 3 shown, after obtaining the DOM tree, the method further includes but is not limited to the following steps S401 to S402.

[0092] S401. Adopt a multimedia element link identification method based on multimedia format and / or multimedia coding method to locate and extract the embedded links of multimedia elements from the DOM tree.

[0093] In the step S401, the multimedia elements specifically include but are not limited to pictures, videos, and / or audios, etc. Specifically, regular expression features can be used for identification, that is, through samples of a large number of different types of multimedia elements, multiple sets of rules are built, and a multimedia element link identification method that can quickly and accurately identify various multimedia formats and / or coding methods and can quickly locate and accurately extract the embedded links of multimedia elements in a complex HTML page structure is established. In addition, the key attributes of the multimedia elements (such as picture size and / or video duration, etc.) can also be preliminarily analyzed and evaluated.

[0094] S402. Classify, store, and manage the embedded links according to the type, source, importance, and / or the relevance of the multimedia elements to the news page data.

[0095] In the step S402, through the classified storage and efficient management of multimedia links, seamless docking with external multimedia processing tools and platforms can be achieved, facilitating subsequent further processing, analysis, and application of multimedia elements.

[0096] Based on the above possible design one, intelligent multimedia content processing can also be performed to achieve seamless docking with external multimedia processing tools and platforms, facilitating subsequent further processing, analysis, and application of multimedia elements.

[0097] As Figure 4 shown, the second aspect of this embodiment provides a virtual device for implementing the general news page data acquisition method described in the first aspect or possible design one, including a data request unit, a first judgment unit, a data scraping unit, a filtering processing unit, a density calculation unit, a second judgment unit, a page parsing unit, and a data processing unit that are sequentially communicatively connected;

[0098] The data request unit is used to send an HTTP request to a target news website to obtain news web page return data;

[0099] The first judgment unit is used to judge whether there is dynamically loaded content on the news page according to the news web page return data;

[0100] The data scraping unit is used to, when it is determined that there is dynamically loaded content, call the browser automation tool Selenium to scrape data after all elements on the news web page are loaded, and use the scraping result as the original news web page data;

[0101] The filtering processing unit is used to perform noise filtering processing on the original news web page data to obtain a DOM tree;

[0102] The density calculation unit is used to calculate the corresponding text density and symbol density for each node in the DOM tree;

[0103] The second judgment unit is used to judge whether the web page content is complex content according to the text density and symbol density of each node;

[0104] The page parsing unit is configured to, when determining that the content is complex, first locate a preliminary extraction result of the key information of the news page from the DOM tree through multi-dimensional feature analysis based on the text density and symbol density of each node, and then import the preliminary extraction result into a large language model based on rules and used for news page parsing, so that the large language model performs in-depth semantic understanding and analysis on the preliminary extraction result according to a pre-constructed set of prompt rules, outputs a CSS selector corresponding to the key information of the news page, and applies the CSS selector to locate a final extraction result of the key information of the news page from the DOM tree, where the key information of the news page includes a news title, news text, news author, and / or news release time;

[0105] The data processing unit is configured to perform data cleaning and standardization processing on the final extraction result to obtain news page data with a unified format and output it.

[0106] For the working process, working details and technical effects of the foregoing device provided in the second aspect of this embodiment, reference may be made to the news page data general acquisition method described in the first aspect or possibly designed in one, which will not be elaborated herein.

[0107] As Figure 5 shown, a computer device provided in the third aspect of this embodiment executes the news page data general acquisition method described in the first aspect or possibly designed in one, and includes a memory, a processor, and a transceiver that are sequentially communicatively connected, where the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the news page data general acquisition method described in the first aspect or possibly designed in one. Specifically, for example, the memory may include, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a flash memory, a first input first output (FIFO), and / or a first input last output (FILO), etc.; the processor may include, but is not limited to, a microprocessor of the STM32F105 series. In addition, the computer device may further include, but is not limited to, a power module, a display screen, and other necessary components.

[0108] For the working process, working details and technical effects of the foregoing computer device provided in the third aspect of this embodiment, reference may be made to the news page data general acquisition method described in the first aspect or possibly designed in one, which will not be elaborated herein.

[0109] In the fourth aspect of this embodiment, a computer-readable storage medium storing instructions for the general news page data acquisition method as described in the first aspect or any possible design is provided. That is, instructions are stored on the computer-readable storage medium, and when the instructions are run on a computer, the general news page data acquisition method as described in the first aspect or any possible design is executed. Among them, the computer-readable storage medium refers to a carrier for storing data, and may include, but is not limited to, computer-readable storage media such as floppy disks, optical discs, hard disks, flash memories, USB flash drives, and / or Memory Sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0110] For the working process, working details, and technical effects of the aforementioned computer-readable storage medium provided in the fourth aspect of this embodiment, reference may be made to the general news page data acquisition method as described in the first aspect or any possible design, which will not be elaborated herein.

[0111] In the fifth aspect of this embodiment, a computer program product is provided, including a computer program or instructions, and when the computer program or the instructions are executed by a computer, the general news page data acquisition method as described in the first aspect or any possible design is implemented. Among them, the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0112] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A general method for collecting news page data, characterized in that: include: Send HTTP request to the target news website to obtain the news webpage return data; Determine whether there is any dynamically loaded content on the news page according to the returned data of the news page; If it is determined that there is dynamically loaded content, the browser automation tool Selenium is called to crawl data after all elements of the news webpage are loaded, and the crawling results are used as the original data of the news webpage; Performing noise filtering on the original data of the news webpage to obtain a DOM tree; For each node in the DOM tree, the corresponding text density and symbol density are calculated; Determining whether the webpage content is complex content according to the text density and symbol density of each node; If it is determined to be complex content, firstly, based on the text density and symbol density of each node, a preliminary extraction result for the key information of the news page is located from the DOM tree through multi-dimensional feature analysis, and then the preliminary extraction result is imported into a rule-based large language model for news page parsing, so that the large language model performs deep semantic understanding and analysis on the preliminary extraction result according to a pre-built prompt word rule set, outputs a CSS selector corresponding to the key information of the news page, and uses the CSS selector to locate the final extraction result for the key information of the news page from the DOM tree, wherein the key information of the news page includes a news title, a news body, a news author and / or a news release time; The final extraction result is subjected to data cleaning and standardization processing to obtain news page data in a unified format and output it.

2. The general method for collecting news page data according to claim 1, characterized in that: Send an HTTP request to the target news website, including: Perform real-time monitoring and analysis on the historical access data, current network status and recent request success rate of the target news website, and select a suitable user agent from the user agent library based on the analysis results to send HTTP requests to the target news website; and / or, during the request process, selecting a safe IP address from the proxy IP pool to replace the source address in the current HTTP request based on the IP blocking monitoring data of the target news website and the risk assessment result of the current HTTP request; And / or, during the request process, based on the analysis results of the request response time distribution of the target news website, the current request traffic and the importance of the target page, a suitable delay time is generated to send two adjacent HTTP requests to the target news website at intervals.

3. The universal method for collecting news page data according to claim 1, characterized in that: For each node in the DOM tree, the corresponding text density is calculated, including: For each node in the DOM tree, the corresponding text density is calculated according to the following formula: In the formula, i represents a positive integer, TD i represents the text density of the ith node in the DOM tree, T i Indicates the number of strings in the i-th node, LT i Indicates the number of strings linked to the i-th node, TG i Indicates the number of labels of the i-th node, LTG i Indicates the number of labels of the links of the i-th node.

4. The universal method for collecting news page data according to claim 1, characterized in that: For each node in the DOM tree, the corresponding text density and symbol density are calculated, including: According to the overall layout style type, section division characteristics type and historical analysis data of the target page, the pre-trained machine learning model based on the artificial intelligence algorithm is used to dynamically adjust the required parameters and / or required thresholds for calculating the text density and / or symbol density; For each node in the DOM tree, the required parameters and / or the required thresholds are applied to calculate and obtain the corresponding text density and symbol density.

5. The universal method for collecting news page data according to claim 1, characterized in that: The final extraction result is subjected to data cleaning and standardization processing, including: Comprehensively analyzing multi-dimensional information to identify redundant information in the final extraction result, and removing the redundant information from the final extraction result, wherein the multi-dimensional information includes text content, HTML structure, data semantic association features and / or visual presentation features; The date, time and / or number in the final extraction result is converted into a unified format to obtain a date, time and / or number in a unified format.

6. The universal method for collecting news page data according to claim 1, characterized in that: After obtaining the DOM tree, the method further includes: Using a multimedia element link identification method based on a multimedia format and / or a multimedia encoding method, locating and extracting an embedded link of a multimedia element from the DOM tree; The embedded links are classified, stored and managed according to the type, source, importance and / or association between the multimedia elements and news page data.

7. A general device for collecting news page data, characterized in that: It includes a data request unit, a first judgment unit, a data grabbing unit, a filtering processing unit, a density calculation unit, a second judgment unit, a page parsing unit and a data processing unit which are sequentially connected in communication; The data request unit is used to send an HTTP request to a target news website to obtain news webpage return data; The first determination unit is used to determine whether there is any dynamically loaded content on the news page according to the returned data of the news page; The data capture unit is used to call the browser automation tool Selenium to capture data after all elements of the news webpage are loaded when it is determined that there is dynamically loaded content, and use the capture result as the original data of the news webpage; The filtering processing unit is used to perform noise filtering on the original data of the news web page to obtain a DOM tree; The density calculation unit is used to calculate the corresponding text density and symbol density for each node in the DOM tree; The second judgment unit is used to judge whether the webpage content is complex content according to the text density and symbol density of each node; The page parsing unit is used to locate and obtain a preliminary extraction result for key information of the news page from the DOM tree through multi-dimensional feature analysis based on the text density and symbol density of each node when determining the content to be complex, and then import the preliminary extraction result into a rule-based large language model for news page parsing, so that the large language model performs deep semantic understanding and analysis on the preliminary extraction result according to a pre-built prompt word rule set, outputs a CSS selector corresponding to the key information of the news page, and uses the CSS selector to locate and obtain a final extraction result for the key information of the news page from the DOM tree, wherein the key information of the news page includes a news title, a news text, a news author and / or a news release time; The data processing unit is used to perform data cleaning and standardization processing on the final extraction result, obtain news page data with a unified format and output it.

8. A computer device, characterized in that: The invention comprises a memory, a processor and a transceiver which are communicatively connected in sequence, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program to execute the general news page data collection method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed on the computer, the general news page data collection method as described in any one of claims 1 to 6 is executed.

10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or the instruction is executed by a computer, the universal news page data collection method as claimed in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Bond product issuing document auditing method and system based on large model agent

    CN121301577A

  • Apparatus and Method for Automated Scraping Code Generation Using Large Language Model

    KR102987526B1