Page Rendering Type Automatic Detection Method and System for General Web Crawlers
By classifying and deciding the pages, the problem of difficulty in taking into account the efficiency and accuracy of general crawlers during the acquisition process is solved, and automated detection and more efficient acquisition effects are achieved.
Patent Information
- Application Number
- CN202111290283.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-02
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-11-02
AI Technical Summary
General-purpose crawlers cannot effectively distinguish between static rendering and dynamic rendering pages during the acquisition process, making it difficult to take into account both the acquisition efficiency and accuracy.
By classifying pages, setting the rendering type identification, and making rendering type judgments for each type of page, including different judgment methods for public pages, index pages and content pages, the final backfill of the judgment results to guide the collection of subsequent pages.
It realizes automatic detection of page rendering types, improves acquisition efficiency and accuracy, and reduces bandwidth usage and CPU and memory consumption.
Smart Images

Figure CN113987319B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of web crawlers, and particularly relates to a method and system for automatically detecting page rendering types for general-purpose crawlers. Background Art
[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] A crawler, also known as a web spider or web robot, is a program or script that automatically fetches information from the World Wide Web according to certain rules; the rendering type is the way a crawler obtains the content in a WEB page, which can be divided into two types: static rendering, directly obtaining the original HTML content through an HTTP request; dynamic rendering, requesting HTML and necessary resource files through a browser engine, executing the scripts in the page, and generating the final page content.
[0004] The inventors found that for most pages with good search engine optimization, the real content in the page can be obtained directly by using static rendering. However, for some websites, JavaScript is used to dynamically load page content, footers, paging buttons, etc. In this case, dynamic rendering must be used to obtain the complete page content. Static rendering is efficient but prone to incomplete page parsing. Dynamic rendering can accurately restore the page structure but has low efficiency and requires more bandwidth and CPU resources.
[0005] If it is necessary to balance the acquisition efficiency and accuracy, it is necessary to determine the rendering types of different types of pages in the site during the first acquisition, so as to correctly and quickly acquire subsequent pages. Since general-purpose crawlers need to acquire a batch of sites in a specific field, they cannot write scripts specifically, and there is no one on duty during the acquisition process. Therefore, how to provide a method that can intelligently and automatically complete the rendering type detection is an urgent problem for existing general-purpose crawlers. Summary of the Invention
[0006] To solve the above problems, the present disclosure provides a method and system for automatically detecting page rendering types for general-purpose crawlers. The solution classifies pages, sets rendering type identifiers, determines the rendering types of each type of page, and finally fills back the determination results to guide the acquisition of subsequent pages.
[0007] According to the first aspect of the embodiments of the present disclosure, a method for automatically detecting page rendering types for general-purpose crawlers is provided, including:
[0008] Determine the page type of the web page;
[0009] Based on the page type of the web page, respectively determine the page rendering types; wherein,
[0010] When the page type is a public page / index page, the rendering type determination is realized based on the number of links obtained from static requests and dynamic requests respectively.
[0011] When the page type is a content page, the rendering type determination is realized based on the similarity of the plain text in the page HTML between static requests and dynamic requests.
[0012] Furthermore, when the page type is a public page / index page, a static request and a dynamic request are respectively initiated to obtain all the links within the page. If the number of static request links is not less than the number of dynamic request links, the current page is statically rendered; otherwise, it is dynamically rendered.
[0013] Furthermore, when the page type is an index page, a static request is initiated to obtain all the links within the page and the paging buttons. If the number of links in the static request is greater than a preset first threshold and includes paging buttons, it is directly determined as a static rendering page.
[0014] Furthermore, if the static request does not meet the condition that the number of links is greater than the preset first threshold and includes paging buttons, a dynamic request is initiated to obtain all the links within the page and the paging buttons; then the conditions for determining dynamic rendering are:
[0015] There is a next page in the dynamic request and no next page in the static request;
[0016] Or
[0017] There is a next page in the dynamic request and the hypertext reference is not a URL.
[0018] Furthermore, when the page type is a content page, a static request and a dynamic request are respectively initiated to obtain the page HTML text, and the plain text is extracted from the HTML text; the similarity of the text is compared. If the text similarity is greater than a preset second threshold, it is determined as static rendering; otherwise, it is dynamically rendered.
[0019] According to the second aspect of the embodiments of the present disclosure, a page rendering type automatic detection system for a general crawler is provided, including:
[0020] A page type determination unit for determining the page type of a web page;
[0021] A rendering type determination unit for respectively performing page rendering type determination based on the page type of the web page; wherein, when the page type is a public page / index page, the rendering type determination is realized based on the number of links obtained from static requests and dynamic requests respectively; when the page type is a content page, the rendering type determination is realized based on the similarity of the plain text in the page HTML between static requests and dynamic requests.
[0022] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program running on the memory. When the processor executes the program, a method for automatically detecting page rendering types for a general-purpose crawler is implemented.
[0023] According to a fourth aspect of the embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, a method for automatically detecting page rendering types for a general-purpose crawler is implemented.
[0024] Compared with the prior art, the beneficial effects of the present disclosure are as follows:
[0025] (1) The present disclosure provides a method and system for automatically detecting page rendering types for a general-purpose crawler. The solution classifies pages, sets rendering type identifiers, determines the rendering types of each type of page, and finally fills back the determination results to guide the subsequent collection of pages.
[0026] (2) In actual scenarios, the proportion of statically rendered sites is the vast majority. Through automatic detection of page rendering types, the present disclosure can effectively reduce the usage rate of bandwidth and reduce the consumption of CPU and memory caused by dynamic renderers.
[0027] (3) In the solution, when the rendering type is DETECT or CONFIG, after the crawler ends normally, the rendering type field in the configuration table will be filled back. When running the crawler next time and the default rendering type is CONFIG, the already estimated rendering type can be directly read without running the rendering type detection again.
[0028] Advantages of additional aspects of the present disclosure will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments and descriptions thereof of the present disclosure are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure.
[0030] Figure 1 It is a flowchart for detecting the rendering type of the public page described in the first embodiment of the present disclosure;
[0031] Figure 2 It is a flowchart for detecting the rendering type of the index page described in the first embodiment of the present disclosure;
[0032] Figure 3 It is a flowchart for detecting the rendering type of the content page described in the first embodiment of the present disclosure. Detailed Implementation Modes
[0033] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0034] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.
[0035] It should be noted that the terms used herein are only for describing specific implementation modes and are not intended to limit the exemplary implementation modes according to the present disclosure. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0036] In the case of no conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other.
[0037] Embodiment 1:
[0038] The purpose of this embodiment is to provide an automated detection method for page rendering types for general-purpose crawlers.
[0039] An automated detection method for page rendering types for general-purpose crawlers includes:
[0040] Determine the page type of the web page;
[0041] Based on the page type of the web page, determine the page rendering type respectively; wherein,
[0042] When the page type is a public page / index page, based on the number of links obtained respectively by static requests and dynamic requests, implement the rendering type determination;
[0043] When the page type is a content page, based on the similarity of the plain text in the page HTML of the static request and the dynamic request, implement the rendering type determination.
[0044] Furthermore, when the page type is a public page / index page, initiate a static request and a dynamic request respectively to obtain all the links in the page. If the number of static request links is not less than the number of dynamic request links, the current page is statically rendered; otherwise, it is dynamically rendered.
[0045] Furthermore, when the page type is an index page, initiate a static request to obtain all the links and paging buttons in the page. If the number of links in the static request is greater than a preset first threshold and includes paging buttons, it is directly determined as a statically rendered page.
[0046] Further, if the static request does not satisfy that the number of links is greater than a preset first threshold and contains a paging button, a dynamic request is initiated to obtain all the links and the paging button within the page; then the conditions for determining dynamic rendering are as follows:
[0047] There is a next page in the dynamic request, and there is no next page in the static request;
[0048] Or
[0049] There is a next page in the dynamic request, and the hypertext reference is not a URL.
[0050] Further, when the page type is a content page, a static request and a dynamic request are respectively initiated to obtain the page HTML text, and the plain text is extracted from the HTML text; the similarity of the texts is compared. If the text similarity is greater than a preset second threshold, it is determined as static rendering, otherwise it is dynamic rendering.
[0051] Further, the page types include public pages, index pages, and content pages.
[0052] Further, for the detected rendering type, after the crawler ends normally, it is filled back into the rendering type field in the configuration table. When the crawler runs next time, the rendering type can be directly read from the configuration table.
[0053] Specifically, for the sake of easy understanding, the method described in this disclosure will be described in detail below in combination with specific applications:
[0054] In this embodiment, tests are carried out with an 8-core CPU and a 100 Mbps bandwidth to compare the resource consumption of static rendering and dynamic rendering:
[0055] During static rendering, a single-machine crawler can consume 90 Mbps of bandwidth and the QPS is 260;
[0056] During dynamic rendering, the renderer consumes a large amount of CPU. The CPU core number allocation between the crawler and the renderer should be more than 1:4;
[0057] During dynamic rendering, because additional resource files need to be downloaded, when the QPS is halved, the bandwidth is 1.65 times the original. That is, the data volume of each page is about 3.3 times that of static rendering.
[0058] This disclosure designs an automatic detection method for the page rendering type of a general crawler. By classifying the pages, setting the rendering type identifier, and determining the rendering type for each type of page, the determination result is finally filled back to guide the subsequent page collection. The method includes the following steps:
[0059] Step 1: Page type classification
[0060] First, the page types of a website are divided into three categories:
[0061] 1. Public page: Specifically refers to the home page of the site. The rendering type of the entire website is estimated by determining the home page (not accurate);
[0062] 2. Index page: Directory or column page, containing links to content pages;
[0063] 3. Content page: Page containing a title and article content;
[0064] Each of the public page, index page, and content page needs to have its own independent rendering type identifier. Among them, for the determination of the index page and content page, existing algorithms can be used, which will not be elaborated in detail in this article.
[0065] Step 2: Set the rendering type identifier
[0066] Set an independent rendering type identifier for each page type. Among them, the rendering type is controlled by 4 parameters:
[0067] 1. render_type_default: Default rendering type;
[0068] 2. render_type_site: Public page rendering type;
[0069] 3. render_type_index: Index page rendering type;
[0070] 4. render_type_content: Content page rendering type.
[0071] Among them, render_type_default is determined during crawler initialization and can be passed 4 types of values:
[0072] · STATIC - Global static: Efficiency first;
[0073] · DYNAMIC - Global dynamic: Accuracy first;
[0074] · DETECT - Determine during operation: When this value is used, when the crawler reads the first page, it will initiate 2 requests for static / dynamic in a blocking manner and determine the page type through a specific algorithm (comparing the number of links or comparing text similarity);
[0075] · CONFIG - Obtain from configuration: CONFIG is similar to DETECT. The difference is that CONFIG will directly read the stored rendering type in the configuration table. If it does not exist, it will run a type detection like DETECT.
[0076] Step 3: Determine the rendering type of the common page. The specific steps are as follows:
[0077] 1. Initiate a static request;
[0078] 2. Obtain all the links within the page;
[0079] 3. Initiate a dynamic request;
[0080] 4. Obtain all the links within the page;
[0081] 5. Compare the number of ordinary links in the two requests;
[0082] 6. If the number of static request links is greater than or equal to the number of dynamic request links, it is determined as static rendering; otherwise, it is dynamic rendering.
[0083] Step 4: Determine the rendering type of the index page. The specific steps are as follows:
[0084] 1. Initiate a static request;
[0085] 2. Obtain all the links and paging buttons within the page;
[0086] 3. If the number of links in the static request is greater than the threshold and includes paging buttons, it is directly determined as a static rendering page, and no dynamic request is initiated;
[0087] 4. Initiate a dynamic request;
[0088] 5. Obtain all the links and paging buttons within the page;
[0089] 6. If there is a next page in the dynamic rendering and no next page in the static rendering, it is determined as dynamic rendering;
[0090] 7. If there is a next page in the dynamic rendering and the href is not a URL, it may be dynamic paging and is determined as dynamic rendering;
[0091] 8. Compare the number of ordinary links in the two requests;
[0092] 9. If the number of static request links is greater than or equal to the number of dynamic request links, it is determined as static rendering,
[0093] otherwise, it is dynamic rendering.
[0094] Step 5: Determine the rendering type of the content page. The specific steps are as follows:
[0095] 1. Initiate a static request;
[0096] 2. Obtain the page HTML text;
[0097] 3. Obtain the plain text in the HTML;
[0098] 4. Initiate a dynamic request;
[0099] 5. Obtain the page HTML text;
[0100] 6. Obtain the plain text in the HTML;
[0101] 7. Compare the text similarity;
[0102] 8. If the text similarity is greater than the threshold, it is determined as static rendering, otherwise it is dynamic rendering.
[0103] Step Six: Rendering type backfill:
[0104] When the rendering type is DETECT or CONFIG, after the crawler ends normally, the rendering type field in the configuration table will be backfilled;
[0105] When the crawler runs next time, when the default rendering type is CONFIG, the already estimated rendering type can be directly read without running the rendering type detection again.
[0106] Furthermore, in the actual scenario, the proportion of statically rendered sites is the vast majority. Through automated page rendering type detection, the bandwidth usage can be effectively reduced, and the CPU and memory consumption caused by the dynamic renderer can be reduced.
[0107] Embodiment Two:
[0108] The purpose of this embodiment is to provide an automated detection system for page rendering types for a general-purpose crawler.
[0109] An automated detection system for page rendering types for a general-purpose crawler includes:
[0110] A page type determination unit for determining the page type of a web page;
[0111] A rendering type determination unit for respectively determining the page rendering type based on the page type of the web page; wherein, when the page type is a public page / index page, the rendering type determination is achieved based on the number of links respectively obtained through static requests and dynamic requests; when the page type is a content page, the rendering type determination is achieved based on the similarity of the plain text in the page HTML in static requests and dynamic requests.
[0112] In more embodiments, there is also provided:
[0113] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment One is completed. For the sake of brevity, it will not be elaborated here.
[0114] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0115] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0116] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in the first embodiment is completed.
[0117] The method in the first embodiment can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0118] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0119] The above-described method and system for automatically detecting the page rendering type for a general-purpose crawler provided by the above embodiment can be implemented and have broad application prospects.
[0120] The above are only the preferred embodiments of the present disclosure and are not used to limit the present disclosure. For those skilled in the art, the present disclosure can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. An automated detection method for page rendering types of general crawlers, characterized in that, it includes: Determine the page type of the web page; Based on the page type of the web page, determine the page rendering type respectively; among them, When the page type is a public page / index page, based on the number of links obtained by static requests and dynamic requests respectively, realize the rendering type determination; When the page type is a content page, based on the similarity of the plain text in the page HTML of static requests and dynamic requests, realize the rendering type determination; When the page type is a public page / index page, initiate a static request and a dynamic request respectively to obtain all the links in the page. If the number of static request links is not less than the number of dynamic request links, the current page is statically rendered, otherwise it is dynamically rendered; When the page type is an index page, initiate a static request to obtain all the links and paging buttons in the page. If the number of links in the static request is greater than a preset first threshold and includes paging buttons, it is directly determined as a static rendering page; If the static request does not meet the condition that the number of links is greater than the preset first threshold and includes paging buttons, initiate a dynamic request to obtain all the links and paging buttons in the page; then the conditions for determining dynamic rendering are: There is a next page in the dynamic request and there is no next page in the static request; Or There is a next page in the dynamic request, and the hypertext reference is not a URL.
2. An automated detection method for page rendering types of general crawlers according to claim 1, characterized in that, When the page type is a content page, initiate a static request and a dynamic request respectively to obtain the page HTML text, and extract the plain text from the HTML text; compare the similarity of the text. If the text similarity is greater than a preset second threshold, it is determined as static rendering, otherwise it is dynamically rendered.
3. An automated detection method for page rendering types of general crawlers according to claim 1, characterized in that, The page types include public pages, index pages and content pages.
4. An automated detection method for page rendering types of general crawlers according to claim 1, characterized in that, For the detected rendering type, after the crawler ends normally, fill it back into the rendering type field in the configuration table. When the crawler runs next time, you can choose to directly read the rendering type from the configuration table.
5. An automated detection system for page rendering types of general crawlers, characterized in that, it includes: A page type determination unit, which is used to determine the page type of the web page; A rendering type determination unit, which is used to determine the page rendering type respectively based on the page type of the web page; among them, when the page type is a public page / index page, based on the number of links obtained by static requests and dynamic requests respectively, realize the rendering type determination; when the page type is a content page, based on the similarity of the plain text in the page HTML of static requests and dynamic requests, realize the rendering type determination; When the page type is a content page, based on the similarity of the plain text in the page HTML of static requests and dynamic requests, realize the rendering type determination; When the page type is a public page / index page, a static request and a dynamic request are initiated respectively to obtain all the links within the page. If the number of links in the static request is not less than the number of links in the dynamic request, the current page is statically rendered; otherwise, it is dynamically rendered. When the page type is an index page, a static request is initiated to obtain all the links and paging buttons within the page. If the number of links in the static request is greater than a preset first threshold and includes paging buttons, it is directly determined as a static rendering page. If the static request does not meet the condition that the number of links is greater than the preset first threshold and includes paging buttons, a dynamic request is initiated to obtain all the links and paging buttons within the page. Then the conditions for determining dynamic rendering are: There is a next page in the dynamic request, and there is no next page in the static request. Or There is a next page in the dynamic request, and the hypertext reference is not a URL.
6. An electronic device, including a memory, a processor, and a computer program running on the memory, characterized in that, when the processor executes the program, it implements a method for automatically detecting the page rendering type of a general-purpose crawler according to any one of claims 1-4.
7. A non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by the processor, it implements a method for automatically detecting the page rendering type of a general-purpose crawler according to any one of claims 1-4.
Citation Information
Patent Citations
Page loading method and device, computer equipment and readable storage medium
CN112395533A
Systems and methods for pre-rendering HTML code of dynamically-generated webpages using a bot
US20200250259A1