Paging Logic Acquisition, and Website Paging Control Method and Device

By analyzing the source code of the target web page, the candidate URL and determining the next page URL with the highest similarity, the automatic acquisition of the website page turn logic is achieved, solving the problem of inefficiency of network crawlers when obtaining continuous web page information, and improving the degree of automation and operation efficiency.

CN114943023BActive Publication Date: 2025-06-27BEIJING JINTI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210387891.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-06-27
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

When obtaining public information on the Internet through web crawlers, the information faced is often distributed on multiple consecutive web pages, resulting in the need to manually control the web page to achieve automatic page turn, which is inefficient.

Method used

Provides a method of obtaining page turn logic. By obtaining the URL of the destination web page, parsing its source code to obtain the candidate URL, and selecting the URL with the highest similarity to the destination web page URL as the URL of the next page, thereby determining the page turn logic of the website to achieve automatic page turn.

Benefits of technology

It has achieved the improvement of automation in the scenario of obtaining massive website information, simplified the operation process of the network crawler system, and improved efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114943023B_ABST
    Figure CN114943023B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for obtaining paging logic and controlling website paging, which relate to the fields of computer technology and big data technology. The specific implementation solution is as follows: obtain the URL of the target web page; obtain the source code of the target web page by making a web page information request for the target web page, and parse the source code to obtain at least one candidate URL, and select the URL with the highest similarity to the URL of the target web page from the at least one candidate URL obtained by parsing the source code as the URL of the next page of the target web page; and determine and return the paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page, so as to be used to implement automatic paging when obtaining information from the website.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of computer technology and big data technology, and particularly relates to a method and device for obtaining paging logic, a method and device for controlling website paging, an electronic device, and a non-transitory computer-readable storage medium storing computer instructions. Background Art

[0002] When obtaining public information on the Internet through a web crawler, it often happens that the required information is displayed on multiple consecutive web pages. At this time, it is necessary to control the web page to achieve automatic paging to facilitate the web crawler to obtain the information on these consecutive web pages. Summary of the Invention

[0003] The present disclosure provides a method, device, equipment, and storage medium for obtaining paging logic and controlling website paging.

[0004] According to one aspect of the present disclosure, a method for obtaining paging logic is provided, including: obtaining the URL of a target web page; by making a web page information request for the target web page, obtaining the source code of the target web page, and parsing the source code to obtain at least one candidate URL, and selecting the URL with the highest similarity to the URL of the target web page from the at least one candidate URL obtained by parsing the source code as the URL of the next page of the target web page; and based on the difference between the URL of the target web page and the URL of the next page of the target web page, determining and returning the paging logic of the website where the target web page is located for realizing automatic paging when obtaining information from the website.

[0005] According to one aspect of the present disclosure, a method for controlling website paging is provided, including: obtaining the website paging logic of a website, where the website paging logic is determined according to the method for obtaining the website paging logic described in the embodiments of the present disclosure; and during the process of accessing the website, based on the website paging logic, controlling the web page of the website to automatically turn from the current page to the next page of the current page.

[0006] According to another aspect of the present disclosure, there is provided a paging logic acquisition device, including: a first acquisition module for acquiring the URL of a target web page; a second acquisition module for acquiring the source code of the target web page by making a web page information request to the target web page; a third acquisition module for parsing the source code to acquire at least one candidate URL; a selection module for selecting, from the at least one candidate URL obtained by parsing the source code, the URL with the highest similarity to the URL of the target web page as the URL of the next page of the target web page; and a fourth acquisition module for determining and returning the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page, so as to be used for implementing automatic paging of web pages when acquiring information from the website.

[0007] According to one aspect of the present disclosure, there is provided a website paging control device, including: a fifth acquisition module for acquiring the website paging logic of a website, where the website paging logic is determined according to the website paging logic acquisition method described in the embodiments of the present disclosure; and a control module for controlling, during the process of accessing the website, the web page of the website to automatically flip from the current page to the next page of the current page based on the website paging logic.

[0008] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the embodiments of the present disclosure.

[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the embodiments of the present disclosure.

[0010] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method described in the embodiments of the present disclosure.

[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0013] Figure 1AExemplarily shown is a system architecture diagram suitable for embodiments of the present disclosure;

[0014] Figure 1B Exemplarily shown is a scenario diagram that can implement embodiments of the present disclosure;

[0015] Figure 2 Exemplarily shown is a flowchart of a method for obtaining paging logic according to embodiments of the present disclosure;

[0016] Figure 3 Exemplarily shown is a flowchart of a method for controlling website paging according to embodiments of the present disclosure;

[0017] Figure 4 Exemplarily shown is a block diagram of an apparatus for obtaining paging logic according to embodiments of the present disclosure;

[0018] Figure 5 Exemplarily shown is a block diagram of an apparatus for controlling website paging according to embodiments of the present disclosure; and

[0019] Figure 6 Exemplarily shown is a block diagram of an electronic device for implementing embodiments of the present disclosure. Detailed Embodiments

[0020] The following describes exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0021] In an alternative embodiment, for an application scenario of obtaining information on multiple consecutive web pages through a web crawler, the URL of the next web page can be constructed based on the URL of the current web page (which can be any web page in the current website) and the website paging logic of the website where the current web page is located, so as to crawl relevant information page by page. However, the website paging logics of different websites may be different. Therefore, when facing a new website, it is always necessary to manually determine the website paging logic of that website. However, in the face of a scenario of obtaining a large amount of website information, the method of relying on manual determination of the website paging logic is cumbersome and inefficient.

[0022] In view of the problem that manual determination of the website paging logic is cumbersome and inefficient, in another alternative embodiment, a general web page pager can be constructed. When facing a scenario of obtaining a large amount of website information, the web page pager can be used to automatically construct or select the corresponding website paging logic for each website, so as to overcome the above-mentioned defects and improve the automation degree of the web crawler system at the same time.

[0023] The present disclosure will be elaborated in detail below in conjunction with the accompanying drawings and specific embodiments.

[0024] The system architecture of the paging logic acquisition method suitable for the embodiments of the present disclosure, as well as the website paging control method and device, is introduced as follows.

[0025] Figure 1A An exemplary system architecture suitable for the embodiments of the present disclosure is shown. It should be noted that Figure 1A The shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other environments or scenarios.

[0026] As Figure 1A shown, the system architecture 100 suitable for the embodiments of the present disclosure may include: a client 101, a server 102, and a server 103.

[0027] In the embodiments of the present disclosure, the client 101 can be used as the execution entity of the website paging logic acquisition method and the website paging control method, and the server 102 and the server 103 can cooperate with the client 101 to implement at least one of the above methods.

[0028] Exemplarily, when the client 101 acquires the website paging logic of a website, it can obtain the source code of the current web page and the source code of the next page of the current web page from the server 102. After the client 101 acquires the website paging logic of the website, it can update it to the website paging logic database of the server 103.

[0029] Or, exemplarily, when the client 101 acquires the website paging logic of a website, it can also directly obtain the existing website paging logic that conforms to the actual situation of the current website from the server 103, and update the occurrence frequency of the website paging logic in the website paging logic database of the server 103 when it can obtain the website paging logic that conforms to the actual situation of the current website.

[0030] The following will elaborate on this system architecture in detail in conjunction with specific embodiments.

[0031] It should be understood that Figure 1A the numbers of the client 101, the server 102, and the server 103 in

[0032] are merely illustrative. According to the implementation requirements, there can be any number of clients 101, servers 102, and servers 103.

[0033] As Figure 1BAs shown in the figure, when querying the public information of an enterprise, it is found that its public information is displayed on multiple consecutive web pages of the same website. Therefore, when using a web crawler to obtain this public information, it is necessary to implement automatic paging based on the website paging logic of this website. In such a scenario, the method for obtaining the website paging logic provided in the embodiments of the present disclosure can be used to construct or select a website paging logic that conforms to the actual situation of this website. Further, after obtaining a website paging logic that conforms to the actual situation of this website, when facing the scenario of obtaining a large amount of website information, the website paging control method provided in the embodiments of the present disclosure can also be used to control automatic paging of web pages during the process of obtaining public information on this website through a web crawler.

[0034] According to an embodiment of the present disclosure, the present disclosure provides a method for obtaining a paging logic.

[0035] Figure 2 The flowchart of the method for obtaining a paging logic according to an embodiment of the present disclosure is exemplarily shown.

[0036] As Figure 2 shown, method 200 may include operations S210 to S250.

[0037] In operation S210, obtain the URL of the target web page.

[0038] In operation S220, obtain the source code of the target web page by making a web information request to the target web page.

[0039] In operation S230, parse the source code obtained in operation S220 to obtain at least one candidate URL.

[0040] In operation S240, select the URL with the highest similarity to the URL of the target web page from at least one candidate URL obtained by parsing the source code as the URL of the next page of the target web page.

[0041] In operation S250, determine and return the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page, so as to be used to implement automatic paging of web pages when obtaining information about the website.

[0042] Exemplarily, if it is specified to query and obtain information on the Nth to N + 1th pages of a certain website, the Nth page among them can be used as the target web page in this embodiment.

[0043] After obtaining the URL of the target web page, a corresponding web information request can be generated so as to send it to the corresponding server for making a web information request to obtain the source code of the target web page.

[0044] It should be understood that if an HTTP GET request is initiated for a specified web page, the web page information request may include the request method GET and the URL of the web page at this time; if an HTTP POST request is initiated for a specified web page, the web page information request may include not only the request method POST and the URL of the web page, but also the form data (form) or payload transmitted by the POST request.

[0045] It should also be understood that the web page information request is not limited to the request method, the URL of the web page, form data, and payload, but may also include other content defined by the network request protocol that is helpful for judging the web page paging logic, such as request Headers, etc. The embodiments of the present disclosure do not limit this.

[0046] After obtaining the source code of the target web page, the source code can be parsed to obtain the URL of the next page of the target web page. It should be understood that if the website where the target web page is located consists of multiple web pages, there may be page number identifiers indicating page numbers or paging buttons such as "Next Page" on these web pages. Therefore, the URL of the next page of the target web page can be extracted from the source code of the target web page using at least one of regular expressions, CSS Selectors, and Xpaths.

[0047] It should be noted that when extracting the corresponding URL from the source code of a web page, regular expressions are more general and can be used for any type of web page; for Xpath, different types of web pages may require different Xpaths. Therefore, in some embodiments of the present disclosure, a combination of regular expressions and Xpath can be used to parse the URL of the next page.

[0048] Exemplarily, some regular expression and Xpath matching patterns and their priorities can be determined based on experience, and attempts to match can be made in order from high to low according to the priority. If the match is successful and the URL of the next page obtained by the match has a high similarity to the URL of the target web page, the loop ends.

[0049] It should be understood that since using regular expressions may match multiple URLs, but if the similarity between the matched URL and the URL of the target web page is low, it is likely that either the wrong place has been matched or the source code of the target web page does not explicitly provide the URL of the next page. For example, if the web page generates the URL of the next page using JavaScript, the source code will not explicitly provide the URL of the next page. Therefore, the URL with the highest similarity to the target web page can be calculated from the matched URLs as the URL of the next page of the target web page to improve the accuracy of the matching result and thus improve the accuracy of the finally constructed web page paging logic.

[0050] After obtaining the URL of the next page of the target web page, the differences between this URL and the URL of the target web page can be compared, and based on these differences, the website paging logic of the website where the target web page is located can be determined and returned.

[0051] As an optional embodiment, the website paging logic can be characterized by at least one of the following information (or variables): paging identifier (page_mark), connector between the paging identifier and the page number of the web page (contact_char), starting page number (start_page_no), whether the home page number is omitted (has_home_page_no), and page number increment between adjacent pages (step).

[0052] In the embodiments of the present disclosure, by comparing the differences between the URL of the target web page and the URL of the next page of the target web page, the corresponding website paging logic can be estimated with a certain strategy. Exemplarily, assume that the home page URL is http: / / xxxx / list.html and the URL of the next page of this home page is http: / / xxxx / list_2.html. By comparing the differences between the two URLs, it can be known that there is an additional "_2" in the URL of the next page, and before "_2" is "list". Thus, "page_mark" can be estimated as ["list"], "contact_char" can be estimated as ["_"], "start_page_no" can be estimated as ["1"], and since there is no page number identifier on the home page, so has_home_page_no can be estimated as [False]. The page number identifier of the second page (i.e., the next page of the home page) is ["2"], thus "step" can be estimated as ["1"].

[0053] It should be understood that the information or variables characterizing the website paging logic can be added, deleted, and modified according to the actual situation, and the embodiments of the present disclosure do not make limitations here.

[0054] Compared with the problem of cumbersome and inefficient processes in determining the website paging logic through manual methods, through the embodiments of the present disclosure, when obtaining network information, especially in the scenario of obtaining a large amount of website information, the website paging logic that conforms to the actual situation of the website where the target web page is located can be automatically constructed based on the URL of the target web page and the URL of the next page of the target web page, or the website paging logic that conforms to the actual situation can be selected from the existing website paging logics. Therefore, the operations and processes are more concise and efficient, and at the same time, the automation degree of the web crawler system can be improved.

[0055] It should be understood that from the source code of a web page, not only can the URL of its next page be parsed, but also other content helpful for judging the paging logic of the website can be parsed. For example, the URL of the page after the next page can also be parsed. Therefore, in other embodiments, for example, the paging logic of the website where the web page is located can be determined and returned according to the difference between the URL of a web page and the URL of the page after the next page of it.

[0056] As an optional embodiment, the method may further include: before determining and returning the paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page, making a web page information request based on the URL of the next page of the target web page. Wherein, in response to the successful web page information request, perform the operation of determining and returning the paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page.

[0057] Since the URL of the next page of the web page obtained by parsing and matching the source code of a web page may be correct or incorrect, in order to ensure that the paging logic of the website where the target web page is located can be accurately determined and returned based on the difference between the URL of the target web page and the URL of the next page of the target web page, before determining and returning the paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page, a web page information request can be made first based on the URL of the next page of the target web page to verify whether the obtained URL of the next page can be successfully requested. Wherein, if the request is successful, it means that the obtained URL of the next page is correct, so the operation of determining and returning the paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page can be performed.

[0058] Further, as an optional embodiment, the method further includes the following operations.

[0059] In response to the failure of the web page information request, predict the paging identifier in the paging logic of the website based on the URL of the target web page and / or the URL of the next page of the target web page.

[0060] Find at least one candidate website paging logic in the website paging logic database whose paging identifier is the same as or similar to the predicted paging identifier.

[0061] Find the paging logic that conforms to the website from at least one candidate website paging logic.

[0062] It should be understood that if the web page information request fails, it means that the URL of the next page currently obtained is incorrect. Therefore, even if the operation of determining and returning the website paging logic of the website where the target web page is located is performed based on the difference between the URL of the target web page and the URL of the next page of the target web page, the correct website paging logic cannot be obtained. In this regard, the website paging logic that conforms to the actual situation of this website can be directly selected from the existing website paging logics for this website.

[0063] Exemplarily, in the case of a request failure, for example, if the URL of the target web page is http: / / xxxx / list.html, the paging identifier "page_mark" in the website paging logic of this website (i.e., the website where the target web page is located) can be predicted to be ["list"]. Thus, at least one candidate website paging logic with the paging identifier ["list"] can be directly searched for in the website paging logic database.

[0064] Furthermore, in the process of implementing the embodiments of the present disclosure, the inventors found that there is a certain regularity in the website paging logic, and many websites use the same website paging logic. Therefore, when searching for at least one candidate website paging logic with the paging identifier ["list"] in the website paging logic database, N candidate website paging logics with the paging identifier ["list"] and the occurrence frequency of top N (the top N, N is greater than or equal to 1) can be selected, thereby improving the matching efficiency.

[0065] Furthermore, in the process of searching for the website paging logic that conforms to this website from the above N candidate website paging logics, it can be tried in turn according to the order of the occurrence frequencies of these website paging logics from high to low until the website paging logic that conforms to the actual situation of this website is found or the traversal ends. If found, the field value representing the occurrence frequency of the currently found web page paging logic is incremented by 1, and at the same time, the entire process ends and the found website paging logic is returned; otherwise, a failure warning is issued.

[0066] Exemplarily, in the process of searching for the website paging logic that conforms to this website from the candidate website paging logics, it can be traversed according to the order of the occurrence frequencies of each paging logic from high to low. For each candidate website paging logic traversed, the URL of the next page of the target web page and / or the URL of the next page of the target web page can be constructed based on this paging logic and the URL of the target web page, and a web page information request is made based on the constructed URL to verify whether the current candidate website paging logic is the website paging logic that conforms to the actual situation of this website.

[0067] Through the embodiments of the present disclosure, not only can the website paging logic of the website where the target web page is located be determined and returned based on the difference between the URL of the target web page and the URL of the next page of the target web page, but also when the strategy for determining the paging logic fails, other strategies can be used to continue to determine the website paging logic of the website where the target web page is located. For example, the paging identifier that may be included in the website paging logic of the website where the target web page is located can also be predicted based on the relevant information in the URL of the target web page and / or the URL of the next page of the target web page, and then the website paging logic of the website where the target web page is located can be estimated based on the paging identifier. Therefore, the means for obtaining the website paging logic are more diverse.

[0068] This embodiment proposes a method for automatically obtaining the website paging logic by combining the idea of rules (such as many websites using the same or similar website paging logics) and statistics (such as the occurrence frequency of each website paging logic). Among them, the rules are mainly reflected in the quantification of the abstract concept of the website paging logic. Statistics are mainly reflected in: when looking for the website paging logic, starting from the website paging logic with the highest occurrence probability (i.e., the highest occurrence frequency) can improve the efficiency of finding the website paging logic that conforms to the actual situation of the current website. However, since it is impossible to exhaust all websites in actual operation, it is difficult to accurately calculate the probability of each website paging logic occurring, but the frequency (i.e., the occurrence frequency) of each website paging logic can be used to estimate the probability of each website paging logic occurring, and the estimation error will gradually decrease as the sample size increases. And, as this solution is continuously applied to new websites, the estimation of the probability of various web page paging logics occurring by this solution will be closer and closer to the actual situation.

[0069] Or, as an alternative embodiment, the method further includes: before determining and returning the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page, perform the following operations.

[0070] Perform a web page information request based on the URL of the next page of the target web page.

[0071] In response to the successful web page information request, obtain the source code of the next page of the target web page.

[0072] Determine whether the similarity between the page structure described by the source code of the target web page and the page structure described by the source code of the next page of the target web page is greater than or equal to the corresponding similarity threshold.

[0073] Wherein, in response to the similarity being greater than or equal to the corresponding similarity threshold, perform the operation of determining and returning the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page.

[0074] Since the URL of the next page of a web page obtained by parsing and matching the source code of the web page may be correct or incorrect, although it is possible to verify whether the URL of the next page of the web page obtained by matching is correct or incorrect by attempting to request web page information according to the above embodiments, but through the above embodiments, when it is verified that the URL of the next page of the web page is correct, it is impossible to further verify whether the URL of the next page of the web page obtained by matching conforms to the actual situation of the website where the target web page is located, such as verifying whether the URL of the next page of the web page obtained by matching is the URL of a relevant web page of this website or the URL of a relevant web page of another website. Therefore, through the above embodiments, sometimes it is also impossible to ensure that based on the difference between the URL of the target web page and the URL of the next page of the target web page, the website paging logic of the website where the target web page is located can be accurately determined and returned.

[0075] In this regard, the embodiments of the present disclosure provide a further solution, that is, when the web page information request based on the URL of the next page of the target web page is successful, the source code of the next page can be further obtained, and the similarity of the page structures of the two web pages can be determined based on the source code of the target web page and the source code of its next page. If the similarity of the page structures of the two web pages is greater than or equal to the corresponding similarity threshold, it is verified that the URL of the next page of the web page obtained by the above matching conforms to the actual situation of the website where the target web page is located; otherwise, it is verified that the URL of the next page of the web page obtained by the above matching does not conform to the actual situation of the website where the target web page is located.

[0076] In the embodiments of the present disclosure, when it is determined that the above similarity is greater than or equal to the corresponding similarity threshold, an operation of determining and returning the website paging logic of the website where the target web page is located based on the URL of the target web page and the URL of the next page of the target web page is performed. Therefore, it can be ensured that based on the difference between the URL of the target web page and the URL of the next page of the target web page, the website paging logic of the website where the target web page is located can be accurately determined and returned.

[0077] Further, as an optional embodiment, the method may further include the following operations.

[0078] In response to the failure of the web page information request or the similarity being less than the corresponding similarity threshold, predict the paging identifier in the website paging logic of the website based on the URL of the target web page and the URL of the next page of the target web page.

[0079] Search for at least one candidate website paging logic in the website paging logic database whose paging identifier is the same as or similar to the predicted paging identifier.

[0080] Search for the website paging logic that conforms to the website in at least one candidate website paging logic.

[0081] Through the embodiments of the present disclosure, not only can the website paging logic of the website where the target web page is located be determined and returned according to the difference between the URL of the target web page and the URL of the next page of the target web page, but also when the above strategy fails to determine the paging logic, other strategies can be used to continue to determine the website paging logic of the website where the target web page is located. For example, the paging identifiers that may be included in the website paging logic of the website where the target web page is located can also be predicted based on the relevant information in the URL of the target web page and / or the URL of the next page of the target web page, and then the website paging logic of the website where the target web page is located can be estimated based on the paging identifiers. Therefore, the means of obtaining the website paging logic are more diversified.

[0082] It should be understood that in the embodiments of the present disclosure, the method for predicting the paging identifier in the website paging logic of the website where the target web page is located based on the URL of the target web page and the URL of the next page of the target web page and the method for selecting the website paging logic that conforms to the actual situation of the current website based on the paging identifier are the same or similar to the methods described in the foregoing embodiments, and the embodiments of the present disclosure will not be elaborated herein.

[0083] Further, as an optional embodiment, the method may further include: when a website paging logic that conforms to the website can be found from the above at least one candidate website paging logic, use the currently found website paging logic for paging when obtaining information about the website, and increment by 1 the field value representing the occurrence frequency of the currently found website paging logic in the website paging logic database.

[0084] Alternatively, as an optional embodiment, the method may further include: when a website paging logic that conforms to the website cannot be found from the above at least one candidate website paging logic, issue an alarm so as to determine the website paging logic of the website manually.

[0085] Through the embodiments of the present disclosure, not only can the website paging logic of the website where the target web page is located be determined and returned according to the difference between the URL of the target web page and the URL of the next page of the target web page, but also when the above strategy fails to determine the paging logic, other strategies can be used to continue to determine the website paging logic of the website where the target web page is located. For example, the paging identifiers that may be included in the website paging logic of the website where the target web page is located can also be predicted based on the relevant information in the URL of the target web page and / or the URL of the next page of the target web page, and then the website paging logic that conforms to the actual situation of the current website can be found from the existing website paging logics based on the paging identifiers. Further, when a website paging logic that conforms to the actual situation of the current website still cannot be found based on the paging identifier, an alarm is also used to notify the user to determine the website paging logic of the website manually. Therefore, the means of obtaining the website paging logic are more diversified.

[0086] As an alternative embodiment, the method may further include: updating the corresponding website paging logic database based on the determined and returned website paging logic of the website.

[0087] Exemplarily, it may be determined whether the website paging logic already exists in the website paging logic database. If it already exists, the field value representing the occurrence frequency of the website paging logic is incremented by 1; otherwise, the website paging logic is inserted into the database, and the field value representing its occurrence frequency is set to 1.

[0088] Through the embodiments of the present disclosure, on the one hand, a solution can be provided for web crawler developers to automatically obtain the website paging logic and thus quickly construct the link address of the next page, thereby eliminating the need for manual determination of the website paging logic and significantly improving the development efficiency when developing web crawlers for a large number of websites; on the other hand, by continuously applying this solution to new websites, some unconsidered website paging logics will be continuously added to the website paging logic database, and the estimation of the occurrence probabilities of various website paging logics will also become closer and closer to the actual situation, and this solution will therefore become more and more general and intelligent.

[0089] As an alternative embodiment, the method may further include: in the case where no candidate URL can be obtained after parsing the source code of the target web page, predicting and returning the website paging logic of the website where the target web page is located based on the URL of the target web page.

[0090] It should be noted that due to reasons such as the website's own reasons or incomplete coverage of the matching pattern, it cannot be guaranteed that the URL of the next page of the target web page can be obtained in operation S230. In this case, based on the information provided in the URL of the target web page, such as the paging identifier therein, the website paging logic with the same paging identifier as that in the URL of the target web page can be searched in the website paging logic database, and the found website paging logic can be used to attempt to construct the URL of the next page of the target web page or the URL of the page after the next page, and then it can be verified whether the constructed URL conforms to the actual situation of this website.

[0091] According to the embodiments of the present disclosure, the present disclosure provides a website paging control method.

[0092] Figure 3 Exemplarily shown is a flowchart of the website paging control method according to the embodiments of the present disclosure.

[0093] As Figure 3 shown, method 300 may include operations S310 to S320.

[0094] In operation S310, obtain the website paging logic of a website, where the website paging logic is determined according to the method for obtaining the website paging logic of the embodiments of the present disclosure, and details thereof are not described herein again in this embodiment.

[0095] In operation S320, during the process of accessing the website, based on the website paging logic, control the web page of the website to automatically flip from the current page to the next page of the current page.

[0096] Among them, the website automatic paging method described in operation S320 is the same as or similar to the website automatic paging method described in the foregoing embodiments, and details thereof are not described herein again in this embodiment.

[0097] Through the embodiments of the present disclosure, the automation degree of the web crawler system can be improved.

[0098] According to an embodiment of the present disclosure, the present disclosure also provides a device for obtaining website paging logic.

[0099] Figure 4 Exemplarily shows a block diagram of a paging logic obtaining device according to an embodiment of the present disclosure.

[0100] As Figure 4 shown, the device 400 may include: a first obtaining module 410, a second obtaining module 420, a third obtaining module 430, a selecting module 440, and a fourth obtaining module 450.

[0101] The first obtaining module 410 is configured to obtain the URL of a target web page.

[0102] The second obtaining module 420 is configured to obtain the source code of the target web page by making a web page information request to the target web page.

[0103] The third obtaining module 430 is configured to parse the source code obtained by the second obtaining module 420 to obtain at least one candidate URL.

[0104] The selecting module 440 is configured to select the URL with the highest similarity to the URL of the target web page from the at least one candidate URL obtained by parsing the source code as the URL of the next page of the target web page.

[0105] The fourth obtaining module 450 is configured to determine and return the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page, so as to be used to implement automatic paging of web pages when obtaining information about the website.

[0106] As an alternative embodiment, the device further includes: a first request module, configured to perform a web page information request based on the URL of the next page of the target web page before determining and returning the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page. Among them, the fourth acquisition module is further configured to, in response to a successful web page information request, perform an operation of determining and returning the website paging logic of the website where the target web page is located based on the URL of the target web page and the URL of the next page of the target web page.

[0107] Further, as an alternative embodiment, the device further includes: a first prediction module, configured to, in response to a failed web page information request, predict a paging identifier in the website paging logic of the website based on the URL of the target web page and / or the URL of the next page of the target web page; a first search module, configured to search for at least one candidate website paging logic in the website paging logic database whose paging identifier is the same as or similar to the predicted paging identifier; and a second search module, configured to search for a website paging logic that conforms to the website from the at least one candidate website paging logic.

[0108] Alternatively, as an alternative embodiment, the device may further include: a second request module, configured to perform a web page information request based on the URL of the next page of the target web page before determining and returning the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page; a fifth acquisition module, configured to, in response to a successful web page information request, acquire the source code of the next page of the target web page; a determination module, configured to determine whether the similarity between the page structure described by the source code of the target web page and the page structure described by the source code of the next page of the target web page is greater than or equal to a corresponding similarity threshold. Among them, the fourth acquisition module is further configured to, in response to the similarity being greater than or equal to the corresponding similarity threshold, perform an operation of determining and returning the website paging logic of the website where the target web page is located based on the URL of the target web page and the URL of the next page of the target web page.

[0109] Further, as an alternative embodiment, the device may further include: a second prediction module, configured to, in response to a failed web page information request or the similarity being less than the corresponding similarity threshold, predict a paging identifier in the website paging logic of the website based on the URL of the target web page and the URL of the next page of the target web page; a third request module, configured to search for at least one candidate website paging logic in the website paging logic database whose paging identifier is the same as or similar to the predicted paging identifier; and a fourth request module, configured to search for a website paging logic that conforms to the website from the at least one candidate website paging logic.

[0110] As an alternative embodiment, the apparatus may further include: a processing module, configured to, when a page turning logic of the website that conforms to the website can be found from the at least one candidate website page turning logic, use the currently found website page turning logic to perform page turning when obtaining information of the website, and increment by 1 a field value representing the occurrence frequency of the currently found website page turning logic.

[0111] Alternatively, as an alternative embodiment, the apparatus may further include: an alarm module, configured to issue an alarm when a page turning logic of the website that conforms to the website cannot be found from the at least one candidate website page turning logic, so as to determine the page turning logic of the website manually.

[0112] As an alternative embodiment, the apparatus may further include: an alarm module, configured to update a corresponding website page turning logic database based on the determined and returned page turning logic of the website.

[0113] As an alternative embodiment, the apparatus may further include: a third prediction module, configured to, when no candidate URL can be obtained after parsing the source code, predict and return the page turning logic of the website where the target web page is located based on the URL of the target web page.

[0114] As an alternative embodiment, the page turning logic of the website is characterized by at least one of the following information: a page turning identifier, a connector between the page turning identifier and the page number of the web page, a starting page number, whether the home page number is omitted, and a page number increment between adjacent pages.

[0115] It should be understood that the apparatus embodiments of the present disclosure correspond to the method embodiments of the present disclosure and are the same or similar, and the embodiments of the present disclosure will not be elaborated herein again.

[0116] According to an embodiment of the present disclosure, the present disclosure further provides a website page turning control apparatus.

[0117] Figure 5 The block diagram of the website page turning control apparatus according to an embodiment of the present disclosure is exemplarily shown.

[0118] As Figure 5 shown, the apparatus 500 may include: a fifth acquisition module 510 and a control module 520.

[0119] The fifth acquisition module 510 is configured to acquire the page turning logic of the website, where the page turning logic of the website is determined according to the method for acquiring the page turning logic of the website in the embodiment of the present disclosure.

[0120] The control module 520 is configured to, during the process of accessing the website, control the web page of the website to automatically turn from the current page to the next page of the current page based on the page turning logic of the website.

[0121] It should be understood that the device embodiments of the present disclosure correspond to the method embodiments of the present disclosure and are the same or similar, and the embodiments of the present disclosure will not be elaborated herein again.

[0122] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0123] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0124] As Figure 6 shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0125] A plurality of components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0126] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the method for obtaining website paging logic (or the website paging control method). For example, in some embodiments, the method for obtaining website paging logic (or the website paging control method) can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method for obtaining website paging logic (or the website paging control method) described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the method for obtaining website paging logic (or the website paging control method) in any other suitable manner (e.g., by means of firmware).

[0127] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0128] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program codes cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program codes may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0129] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0130] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, speech input, or tactile input).

[0131] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0132] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client - server relationship is created by computer programs running on the respective computers and having a client - server relationship with each other.

[0133] It should be understood that various forms of the flow shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0134] The above - mentioned specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for obtaining paging logic, comprising: Obtaining the URL of a target web page; By making a web page information request to the target web page, obtaining the source code of the target web page, and parsing the source code to obtain at least one candidate URL, and selecting from the at least one candidate URL obtained by parsing the source code the URL with the highest similarity to the URL of the target web page as the URL of the next page of the target web page; And Based on the difference between the URL of the target web page and the URL of the next page of the target web page, determining and returning the paging logic of the website where the target web page is located, so as to be used for automatic paging when obtaining information from the website; Before determining and returning the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page, making a web page information request based on the URL of the next page of the target web page, wherein, in response to the success of the web page information request, performing the operation of determining and returning the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page; In response to the failure of the web page information request, predicting the paging identifier in the website paging logic of the website based on the URL of the target web page and / or the URL of the next page of the target web page; Searching in the website paging logic database for at least one candidate website paging logic whose paging identifier is the same as or similar to the predicted paging identifier; and Searching for the website paging logic that conforms to the website from the at least one candidate website paging logic.

2. The method according to claim 1, further comprising: Before determining and returning the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page, making a web page information request based on the URL of the next page of the target web page; In response to the success of the web page information request, obtaining the source code of the next page of the target web page; Determining whether the similarity between the page structure described by the source code of the target web page and the page structure described by the source code of the next page of the target web page is greater than or equal to the corresponding similarity threshold, wherein, in response to the similarity being greater than or equal to the corresponding similarity threshold, performing the operation of determining and returning the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page.

3. The method according to claim 2, further comprising: In response to the failure of the web page information request or the similarity being less than the corresponding similarity threshold, predicting the paging identifier in the website paging logic of the website based on the URL of the target web page and the URL of the next page of the target web page; Searching in the website paging logic database for at least one candidate website paging logic whose paging identifier is the same as or similar to the predicted paging identifier; And Searching for the website paging logic that conforms to the website from the at least one candidate website paging logic.

4. The method according to claim 1 or 3, further comprising: When it is possible to find a website paging logic that matches the website from the at least one candidate website paging logic, use the currently found website paging logic to perform paging when obtaining information for the website, and increment by 1 the field value representing the occurrence frequency of the currently found website paging logic.

5. The method according to claim 1 or 3, further comprising: When it is not possible to find a website paging logic that matches the website from the at least one candidate website paging logic, issue an alarm so as to determine the website paging logic of the website manually.

6. The method according to claim 1, further comprising: Update the corresponding website paging logic database based on the determined and returned website paging logic of the website.

7. The method according to claim 1, further comprising: When no candidate URL can be obtained after parsing the source code, predict and return the website paging logic of the website where the target web page is located based on the URL of the target web page.

8. The method according to claim 1, wherein, The website paging logic is characterized by at least one of the following information: a paging identifier, a connector between the paging identifier and the page number, a starting page number, whether the home page number is omitted, and the page number increment between adjacent pages.

9. A website paging control method, comprising: Obtain the website paging logic of a website, where the website paging logic is determined according to the method described in any one of claims 1-8; and During the process of accessing the website, based on the website paging logic, control the web page of the website to automatically flip from the current page to the next page of the current page.

10. A paging logic acquisition device, comprising: A first acquisition module for acquiring the URL of a target web page; A second acquisition module for acquiring the source code of the target web page by making a web page information request for the target web page; A third acquisition module for parsing the source code to obtain at least one candidate URL; A selection module for selecting the URL with the highest similarity to the URL of the target web page from the at least one candidate URL obtained by parsing the source code as the URL of the next page of the target web page; And A fourth acquisition module for determining and returning the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page, so as to be used for realizing automatic web page paging when obtaining information for the website; Before determining and returning the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page, make a web page information request based on the URL of the next page of the target web page, where, in response to a successful web page information request, perform the operation of determining and returning the website paging logic of the website where the target web page is located based on the difference between the URL of the target web page and the URL of the next page of the target web page; In response to a failed web page information request, predict the paging identifier in the website paging logic of the website based on the URL of the target web page and / or the URL of the next page of the target web page; Search for at least one candidate website paging logic in the website paging logic database whose paging identifier is the same as or similar to the predicted paging identifier; and Search for the website paging logic that conforms to the website from the at least one candidate website paging logic.

11. A website paging control device, comprising: A fifth acquisition module, configured to acquire the website paging logic of a website, wherein the website paging logic is determined according to the method described in any one of claims 1-8; and A control module, configured to control the web page of the website to automatically flip from the current page to the next page of the current page based on the website paging logic during the process of accessing the website.

12. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any one of claims 1-9.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method described in any one of claims 1-9.

Citation Information

Patent Citations

  • Method and device for calculating relevant webpage URL pattern

    CN103617228A

  • Method and device for recognizing page number identification in webpage URL

    CN103631906A