Methods, devices, media, and programs for collecting web page content from public networks.
By coordinating the interaction between private and public networks through a private network manager, web page content is automatically collected, solving the problems of uncertainty and cumbersomeness in obtaining public network content from private networks, and realizing a secure, automated content acquisition and optimized collection mode.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- METEOROLOGICAL DEV & PLANNING INST OF CHINA METEOROLOGICAL ADMINISTRATION
- Filing Date
- 2025-03-25
- Publication Date
- 2026-05-05
AI Technical Summary
People in private networks find it difficult to securely and conveniently access and obtain web content from public networks. Existing technical solutions result in uncertainty, cumbersome processes, and time consumption, and cannot automatically adjust the collection mode according to network conditions.
Using a private network manager as an intermediary, the interaction between the private network and the public network is coordinated. The web page collector is invoked to automatically collect web page content, and the collection mode is set according to the network download speed, load, security level and importance of the column, and stored in the private network.
It enables secure and automated acquisition of public network content within a private network, ensuring network isolation and security, while optimizing acquisition frequency and modes to improve efficiency and reliability.
Smart Images

Figure CN120547165B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure relate to the field of computers, and more specifically to methods for collecting web page content from public networks, apparatus for collecting web page content from public networks, non-transitory computer-readable storage media, and computer program products. Background Technology
[0002] A private network typically refers to a network used internally by a company or organization, isolated from the public internet. Such networks are designed to provide enhanced security and privacy. In a private network, data and information transmission are strictly controlled and regulated to ensure that only authorized users can access it. A public network, also known as the public internet, is an open network environment that allows any user to access it using appropriate devices. Data transmission on a public network is public and is potentially vulnerable to various cyberattacks and eavesdropping.
[0003] Many businesses and organizations have explicit policies prohibiting or restricting direct connections between private and public networks. However, employees within these organizations may need to access public networks to search for web pages, such as regularly checking the latest content on certain news websites. In such cases, it becomes difficult to access and search public websites from within a private network. Summary of the Invention
[0004] According to one aspect of this disclosure, at least one embodiment provides a method for collecting web page content from a public network, comprising:
[0005] The private network manager obtains instructions for collecting web page content from the private network, wherein the instructions include at least the URL of the web page on the public network to be collected;
[0006] The private network manager invokes web page crawlers on the public network according to instructions, where the private network and the public network do not communicate directly.
[0007] The web scraping mode is set according to instructions by a private network manager or a web scraper, so that the web scraper can obtain web page content from the URL according to the scraping mode;
[0008] The private network manager stores web page content on a private network.
[0009] According to another aspect of this disclosure, at least one embodiment provides an apparatus for collecting web page content from a public network, comprising: a private network manager and a web page collector, wherein,
[0010] The private network manager obtains instructions for collecting web page content from the private network, wherein the instructions include at least the URL of the web page on the public network to be collected;
[0011] The private network manager invokes web page crawlers on the public network according to instructions, where the private network and the public network do not communicate directly.
[0012] The web scraping mode is set according to instructions by a private network manager or a web scraper, so that the web scraper can obtain web page content from the URL according to the scraping mode;
[0013] The private network manager stores web page content on a private network.
[0014] According to another aspect of this disclosure, at least one embodiment provides an apparatus for collecting web page content from a public network, comprising: a memory for storing computer instructions; and a processor for reading the computer instructions from the memory and executing a method according to at least one embodiment of this disclosure.
[0015] According to another aspect of this disclosure, at least one embodiment provides a non-transitory computer-readable storage medium having computer instructions stored thereon, wherein, when executed by a processor, the computer instructions cause the processor to perform a method according to at least one embodiment of this disclosure.
[0016] According to another aspect of this disclosure, at least one embodiment provides a computer program product having computer instructions stored thereon, wherein, when executed by a processor, the computer instructions cause the processor to perform a method according to at least one embodiment of this disclosure. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A scenario diagram illustrating the collection of web page content from a public network according to at least one embodiment of the present disclosure is shown.
[0019] Figure 2 A scenario diagram is shown illustrating the collection of web page content from a public network via a private network manager according to at least one embodiment of the present disclosure.
[0020] Figure 3 A flowchart illustrating a method for collecting web page content from a public network according to at least one embodiment of the present disclosure is shown.
[0021] Figure 4 A deployment application scenario diagram of an actual computer cluster according to at least one embodiment of the present disclosure is shown.
[0022] Figure 5 A functional partitioning diagram of a private network manager and a web crawler according to at least one embodiment of the present disclosure is shown.
[0023] Figure 6 A block diagram of an apparatus for collecting web page content from a public network according to at least one embodiment of the present disclosure is shown.
[0024] Figure 7 Another block diagram of an apparatus for collecting web page content from a public network according to at least one embodiment of the present disclosure is shown. Detailed Implementation
[0025] Referring now to specific embodiments of this disclosure, examples of which are illustrated in the accompanying drawings. Although this application will be described in conjunction with specific embodiments, it will be understood that it is not intended to limit this application to the described embodiments. Rather, it is intended to cover variations, modifications, and equivalents included within the spirit and scope of this disclosure. It should be noted that the method steps described herein can be implemented by any functional block or functional arrangement, and any functional block or functional arrangement can be implemented as a physical entity or a logical entity, or a combination of both.
[0026] In this article, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0027] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0028] Figure 1 A scenario diagram illustrating the collection of web page content from a public network according to at least one embodiment of the present disclosure is shown.
[0029] like Figure 1As shown, a user 111 in private network 110 may need to access public network 120 to retrieve web pages 121 via computer 112, for example, to regularly check the latest content in certain sections of certain news websites. However, due to the isolation between private network 110 and public network 120, it is difficult to access and query websites on public network 120 from private network 110. Moreover, manually querying these websites regularly is labor-intensive and increases uncertainty. In addition, differences in network speed, website load, and website security levels of the websites being queried may lead to failures when users manually query the website, or it may be unclear when to perform another query. Furthermore, if the website being queried is in a foreign language, it may also need to be translated into Chinese. Sometimes, different users may query the same content on the same website, so it is time-consuming for each user to collect website content themselves. This series of complex operations leads to insecurity, uncertainty, cumbersomeness, and time consumption in the process of cross-network content collection from public websites.
[0030] The defects and problems existing in the above-mentioned prior art solutions are also the result of the inventor's careful research after practical and creative labor. The discovery process of the above problems and the solutions proposed by at least one embodiment disclosed below for the above problems are all creative contributions of the inventor during the invention process.
[0031] According to at least one embodiment of this disclosure, the technique of collecting web page content from a public network involves a private network manager obtaining instructions to collect web page content from the private network, wherein the instructions include at least the URL of the web page to be collected from the public network; the private network manager invoking a web page collector in the public network according to the instructions, wherein the private network and the public network do not communicate directly; the private network manager or the web page collector setting a collection mode according to the instructions, so that the web page collector obtains web page content from the URL according to the collection mode; and the private network manager storing the web page content in the private network.
[0032] In this way, by using the private network manager as an intermediary to coordinate the interaction between the requests for collecting network content issued in the private network and the public network, the isolation between the private network and the public network and the security of the private network are guaranteed. At the same time, since users only need to issue a command to collect web page content, the private network manager can automatically collect the content of websites in the public network according to the command, so that people in the private network can safely and conveniently obtain the content of websites in the public network.
[0033] Since users may not know the network download speed, website load, security level, or importance of the webpage sections they want to collect from the URL, they cannot adjust their webpage content collection mode based on this information. This could result in collecting webpage content too frequently or too sparsely, or failing to collect the desired content. However, according to the embodiments of this disclosure, the webpage content collection mode can be set based on instructions and the URL's network download speed, website load, security level, and importance of the webpage sections, thereby collecting webpage content in an appropriate manner.
[0034] Figure 2 A scenario diagram is shown illustrating the collection of web page content from a public network via a private network manager according to at least one embodiment of the present disclosure.
[0035] Figure 2 The diagram illustrates a private network 110, a user 111 within the private network, and a computer 112 used by the user 111. Computer 112 may include a central processing unit (CPU), memory, input devices, and output devices. These hardware components work together to enable the computer to perform various tasks.
[0036] A private network manager 113 deployed in the private network can act as an intermediary to coordinate the interaction between the requests for collecting network content issued from the private network 110 and the public network 120. The public network 120 may contain webpage content 121 at a URL, and user 111 wishes to collect webpage content 121. Upon receiving the instruction from user 111, the private network manager 113 can invoke the webpage collector 122 in the public network 120 to complete the collection of webpage content 121.
[0037] Figure 3 A flowchart illustrating a method for collecting web page content from a public network according to at least one embodiment of the present disclosure is shown.
[0038] The following is combined Figure 2 and 3 This document describes a process for collecting web page content from a public network according to at least one embodiment of the present disclosure.
[0039] First, user 111 in private network 110 issues a command to collect web page content via computer 112. This command may include at least the URL of the web page on the public network to be collected, such as a URL: http: / / www.*****.com In some embodiments, the instructions may further include at least one of the following: webpage category to be collected, collection method, collection frequency, whether translation is required, format of the collected content, and filename prefix.
[0040] For example, a user can click "Add Order" on the download order list of a private network website, fill in the download request list, and the list can include the URL of the website to be collected, the page categories to be collected, the collection method, the collection frequency, whether translation is required, the document format, the file name prefix (dictionary maintenance), etc. After completing the form, the user can save it and choose to submit or delete the order. The submitted order contains the instructions for collecting web page content.
[0041] In the download order list, users can select an order to view its detailed information. Users can see all the information entered when the order was created, as well as the order's approval and execution status. Data to be downloaded (or collected) in an order may require multiple downloads or be performed periodically; users can view a list of all tasks that have started downloading in the order view.
[0042] An approval process can be added, allowing department heads to view and approve download orders submitted by users within their department from the download order list. Before approval, the order status is "Pending Approval," and the instruction can be blocked. After approval, the order status becomes "Approved," and the instruction becomes available.
[0043] Orders can also be deleted by the user before submission or when they are returned due to failure to pass manager approval.
[0044] exist Figure 3 In step 310, the private network manager 113 deployed in the private network 110 can obtain instructions to collect web page content from the private network 110.
[0045] exist Figure 3 In step 320, the private network manager invokes the web scraper on the public network according to instructions. The private network and the public network do not communicate directly.
[0046] Alternatively, backend maintenance personnel can receive approved orders, attempt to test downloads on a public network based on the order details, and, once the download is successful, use a private network manager on the private network to invoke the web scraper on the public network according to the instructions.
[0047] exist Figure 3 In step 330, the private network manager or the web crawler sets the collection mode according to the instructions, so that the web crawler can obtain web page content from the URL according to the collection mode.
[0048] In some embodiments, step 330, in which a private network manager or a web scraper sets the scraping mode according to instructions, may include setting the scraping mode according to the instructions and the network download speed of the URL, the load of the URL, the security level of the URL, and the importance of the scraped web page sections of the URL.
[0049] Since users may not know the network download speed, website load, security level, or importance of the webpage sections they want to collect from the URL, they cannot adjust their webpage content collection mode based on this information. This could result in collecting webpage content too frequently or too sparsely, or failing to collect the desired content. Therefore, according to these embodiments of the present disclosure, the webpage content collection mode can be set based on instructions and the website's network download speed, website load, security level, and importance of the webpage sections, thereby collecting webpage content in an appropriate manner.
[0050] In some embodiments, setting the collection mode by a private network manager or a web crawler based on instructions and factors such as the network download speed of the URL, the load of the URL, the security level of the URL, and the importance of the URL's sections may include:
[0051] You can set the interval for collecting web page content. for:
[0052]
[0053]
[0054]
[0055] )
[0056] Interval between webpage content collection It can guide the time interval between two requests from a web crawler to a URL to collect content, and can reflect the frequency of collection. The longer the time, the lower the sampling frequency. If the length is short, the sampling frequency is long. The unit can be seconds.
[0057] Among them, T base This is the base data acquisition interval. This is the default acquisition interval without dynamic adjustment, providing a baseline value for the data acquisition time interval. The base acquisition interval can be 86400 seconds, or 1 day.
[0058] R current This is the current network download speed (in Mbps) for the URL. This is a real-time measurement of the network speed, used to reflect the current network conditions.
[0059] R maxThis represents the ideal maximum network download speed (in Mbps) for a given URL. It's a threshold for network speed used to compare against the current network download speed to assess potential changes in network speed.
[0060] k R This is a network download speed adjustment factor. It's a non-linear adjustment parameter used to adjust the sensitivity and curve shape of the effect of network speed on the data acquisition time interval. This is achieved by adjusting k... R It can control how quickly the data collection time interval changes with the network speed.
[0061] ω R The weighting of network download speed is determined by this parameter. This is a weighting parameter used to adjust the relative importance of network speed in the overall calculation. It is adjusted by ω. R This can balance the impact of network speed and other factors (such as load, security level, etc.) on the data collection time interval.
[0062] L current This represents the current load of the website. This is a real-time measured value of the website load, which can be the number of requests, concurrent users, etc., used to reflect the current load status of the website.
[0063] L max This represents the ideal maximum or threshold for the load on a URL. It's a threshold for website load, used to compare it to the URL's current load to assess load variations.
[0064] k L This is the load adjustment factor for the website. It's a non-linear adjustment parameter used to adjust the sensitivity and curve shape of the impact of website load on the data collection time interval. This is achieved by adjusting k... L It can control how quickly the data collection time interval changes with the website load.
[0065] ω L The load factor of a website affects its weight. This is a weighting parameter used to adjust the relative importance of website load in the overall calculation. By adjusting ω... L This can balance the impact of website load and other factors on the data collection time interval.
[0066] S current This indicates the security level of the website. It's a real-time measurement of the website's security level, which can include security scores, vulnerability remediation rates, etc., reflecting the current security status of the website.
[0067] S max This represents the ideal maximum level of security for a website. It's a threshold used to compare a website's security level to its actual security level in order to assess how much security has changed.
[0068] kS This is a security adjustment factor for the website address. It's a non-linear adjustment parameter used to adjust the sensitivity and curve shape of the effect of website security level on the data collection time interval. This is achieved by adjusting k... S It can control how quickly the data collection time interval changes with the website's security level.
[0069] ω S The security level of a website affects its weight. This is a weighting parameter used to adjust the relative importance of website security in the overall calculation. By adjusting ω... S This can balance the impact of website security and other factors on the data collection time interval.
[0070] I importance This score represents the importance of a website's website section. It's based on a comprehensive evaluation of factors such as user traffic and content update frequency, reflecting the current importance of each section.
[0071] δ represents the influence coefficient of the website's category importance. This is a coefficient parameter used to adjust the direct impact of the website category importance on the data collection time interval. By adjusting δ, the rate at which the data collection time interval changes with the website category importance can be controlled.
[0072] ω I This is a non-linear adjustment factor that influences the weight of website category importance. It's a non-linear adjustment parameter used to adjust the shape of the curve showing how the importance of website categories affects their weight. By adjusting ω... I This allows for further fine-tuning of the impact of website section importance on the data collection time interval.
[0073] In order to take into account the impact of website network speed and load, a similar non-linear adjustment method was used, namely... Where X represents network speed R or load L. This form allows for a more gradual reduction in the sampling interval when the network speed or load is close to its maximum, avoiding overly frequent sampling.
[0074] To consider the impact on website security, use This format allows for a more significant reduction in the data collection interval at lower security levels, enabling more frequent monitoring of potential security issues.
[0075] To account for the importance of website sections, a linear weighting method was used to directly add the values to the above formula, but a non-linear adjustment factor ω was employed. L To control the shape of the curve that affects it.
[0076] By adjusting these parameters, the influence of various factors on the data acquisition time interval can be flexibly controlled to adapt to different application scenarios and needs.
[0077] Web scrapers can obtain web page content from URLs according to the scraping patterns defined above. For example, web scrapers can obtain web page content from URLs through web crawlers, by calling the URL's Application Programming Interface (API), or by using browser plugins.
[0078] In some embodiments, a private network manager or a web page collector sets a collection mode according to instructions, so that the web page collector can obtain web page content from a URL according to the collection mode. This includes: sending an access request to the URL according to the collection mode; obtaining web page content result data from the URL; obtaining a list of web page content data from the web page content result data; decomposing the list of web page content data to obtain individual data address configurations; analyzing the structure of individual data based on the individual data address configurations to obtain available information; extracting content within the collected web page categories from the available information according to the collected web page categories; and if the instructions include the need for translation, calling a translation program to translate the extracted content within the collected web page categories.
[0079] Taking a web crawler as an example, it can send access requests to a URL based on the collection mode. These requests include the Uniform Resource Locator (URL) address, latency, timeout, page encoding format, and number of retries. For example:
[0080] Access request: GET
[0081] URL: ${nextPage==null?'https: / / www.****.com':nextPage}
[0082] Page encoding: UTF-8
[0083] For example, the URL returned the result data (resp) of the collected webpage.
[0084] Next, extract the list of webpage content data from the webpage content results data. For example, use XPath technology to select and extract the data, obtaining the list content `dataList`. For example:
[0085] ${resp.xpaths(' / / *[@id="block-server-theme-content"] / article / div / div / div / div[3] / div / div[2] / div / div / div[2] / div / div / div / div')}
[0086] Next, we break down the webpage content data list to obtain the individual data address configuration. For example, we break down the list and set up loop logic to process the data objects one by one. This includes the collection data {list}, the loop variable {item}, the index {index}, and the end information settings. For example, we break down the individual data address configuration as follows:
[0087] ${funNewsUrl('https: / / wmo.int',item.selector('div').selector('a').attr('href'))}
[0088] Next, based on the configuration of a single data address, analyze the structure of that single data entry to obtain usable information. For example, analyze the structure of a single data entry to obtain usable information such as title, publication time, news address, body text, and image list.
[0089] like:
[0090] Title:${resp.regx(' <title>(.*?)< / title> ')}
[0091] PublishDate:${resp.xpath(' / / *[@id="block-server-theme-content"] / article / div / div / div / div[2] / div / div[1] / div
[0092] / div[1] / div[1] / text()')}
[0093] Link: ${newsUrl}
[0094] Content: ${fun_cleanHtml(string.replaceAll(resp.html
[0095] .selector('.node__content').selectors('.container-wide'
[0096] [1].html(),'<img[^> ]*>','placeholder_for_image'))}
[0097] Images: ${resp.html.selector('.node__content')
[0098] .selectors('.container-wide')[1]
[0099] .elementImagesWithDomain('https: / / www.****.com')}
[0100] If the data to be collected is for a specific section of a webpage, the keywords for that section can be obtained from the available information above, thus identifying the content of that section.
[0101] Next, extract the content from the available information based on the webpage categories. For example, execute the data block SQL program: store the obtained available information, such as title, publication time, news address, body text, image list, etc., into a pre-defined database table or file according to the data table field structure. For example:
[0102] INSERTIGNOREINTO`spider`.`tb_wmo_int`(`id`,`title`,`content`,`url`,`publish_date`,`images`,`translate_flag`,`order_no`)VA LUES('${idx}','${title}','${content}','${link}','${publishDate}','${images.toString()}','${translateFlag}','${orderNo}');
[0103] If the instruction includes the requirement for translation, the translation program will be invoked to translate the content extracted from the webpage sections. The `translateFlag` field can be set in the instruction, with 1 indicating translation is required and 0 indicating no translation is needed. If the `translateFlag` field is set to 1, the large model translation interface can be invoked to translate the collected data, providing Chinese or bilingual (Chinese and English) content.
[0104] If the downloaded webpage content is the file itself (such as a PDF file) rather than the scraped webpage text information, the file can be automatically saved and displayed in the Internet download file list while being synchronized to the user's personal space, so that all users in the system can view and download it, and can also be searched based on information such as file name and download time.
[0105] exist Figure 3 In step 340, the private network manager stores the web page content on the private network.
[0106] In this step, the web crawler obtains webpage content from the URL and stores it in the data management area of the public network. It then extracts the webpage content from the data management area and stores it in a buffer on a private network, where the private network manager communicates with both the private and public networks. Finally, according to predetermined push rules, the webpage content is sent from the private network manager to the private network's storage (e.g., the user's computer's storage).
[0107] Specifically, in this step, the web crawler first stores the collected web page content in a data storage area on a public network. Then, the private network manager retrieves the web page content stored in the public network's data storage area. The private network manager then stores the web page content in a data storage area on the private network according to predetermined push rules (e.g., periodically, irregularly, based on the number of web page contents stored in the public network's data storage area exceeding a threshold, or based on user requests, etc.). This allows users on the private network to view and download the web page content from the private network's data storage area to their own computers.
[0108] Thus, by using a private network manager as an intermediary to coordinate the interaction between the private network's requests for content collection and the public network, the isolation between the private and public networks and the security of the private network are ensured. Simultaneously, since users only need to issue a command to collect webpage content, the private network manager can automatically collect content from websites on the public network according to the command, enabling personnel on the private network to securely and conveniently access content from websites on the public network. Furthermore, according to these embodiments of the present disclosure, the content collection mode can be set based on the command, the network download speed of the URL, the URL's load, the URL's security level, and the importance of the webpage sections to be collected, thereby collecting webpage content in an appropriate manner.
[0109] Figure 4 A deployment application scenario diagram of an actual computer cluster according to at least one embodiment of the present disclosure is shown.
[0110] like Figure 4 As shown, user 411 in private network 410 can operate proxy service cluster 412 (computers in the cluster) to submit orders for web page collection. Proxy service clusters 412 can be connected via HTTP persistent connections. Note that a cluster can consist of one or more computers.
[0111] The private network 410 also includes a gateway cluster 413, a registration and configuration center cluster 414, a public service cluster 415, a business service cluster 416, a database cluster 417, and a distributed file system cluster 418.
[0112] Gateway cluster 413 is primarily used to improve the performance and reliability of system inbound access. By having multiple gateway instances work in parallel, the following functions can be achieved: Load balancing: External access requests are evenly distributed across different gateway instances, preventing overload of a single gateway and improving system response speed and throughput. High availability and fault tolerance: When a gateway instance fails, other gateway instances can take over its work, ensuring service continuity and stability. Security protection: Gateway clusters can also serve as the first line of defense for security, protecting the internal network from external attacks through measures such as configuring firewall rules and intrusion detection systems.
[0113] The registration configuration center cluster 414 is primarily used for centralized management of configuration information, such as database connection information and application parameters. Through the configuration center, dynamic updates and version management of configurations can be easily achieved.
[0114] Public Service Cluster 415 is mainly used to provide shared services across business areas, such as authentication and authorization, logging, and message notification.
[0115] The Business Service Cluster 416 is primarily used to provide specific business functions and services. These services are typically related to specific business areas, such as web page order processing, approval processing, and payment services.
[0116] These clusters can process user 411's orders to collect web page content and ultimately issue instructions to collect the web page content.
[0117] The private network manager 430 can obtain instructions for collecting web page content from the cluster of private network 410.
[0118] User 421 on public network 420 can operate proxy service cluster 422 to access the Internet, etc. Proxy service clusters 422 can connect to each other via HTTP persistent connections.
[0119] The public network 420 may also include a data acquisition cluster 423 and a translation cluster 424. The private network manager 430 can invoke the data acquisition cluster 423 to perform a web page content acquisition process according to at least one embodiment, and can also invoke the translation cluster 424 to perform a translation process of the acquired web page content. The data acquisition cluster 423 can store the acquired web page content (or files) in a database cluster within the public network 420. The private network manager 430 can extract the acquired web page content (or files) from the database cluster within the public network 420 and store the acquired web page content in a database cluster 417 within the private network 410. The distributed file system cluster 418 within the private network 410 can store the acquired files (such as, for example, PDF files).
[0120] Figure 5 A functional partitioning diagram of a private network manager and a web crawler according to at least one embodiment of the present disclosure is shown.
[0121] To complete the entire data collection task, the private network manager can help perform the following operations: downloading the order list, adding orders, viewing orders, approving orders, processing orders, synchronizing internal and external network data (synchronizing the collected web page content from the public network to the private network), downloading the file list from the internet, and completing orders. The web page collector can help perform the following operations (e.g., depending on the collection mode): downloading task configuration and (depending on the configuration) downloading data.
[0122] Figure 6 A block diagram of an apparatus 600 for collecting web page content from a public network according to at least one embodiment of the present disclosure is shown.
[0123] The apparatus 600 for collecting web page content from public networks includes a private network manager 610 and a web page collector 620.
[0124] The private network manager 610 obtains instructions from the private network to collect web page content, wherein the instructions include at least the URL of the web page on the public network to be collected.
[0125] The private network manager 610 invokes the web page collector 620 in the public network according to the instructions. The private network and the public network do not communicate directly.
[0126] The private network manager 610 or the web page collector 620 sets the collection mode according to the instructions, so that the web page collector 620 can obtain web page content from the URL according to the collection mode.
[0127] The private network manager 610 stores web page content on a private network.
[0128] In some embodiments, the instructions may further include at least one of the following: webpage category to be collected, collection method, collection frequency, whether translation is required, collection content format, and filename prefix.
[0129] In some embodiments, the private network manager 610 or the web page collector 620 sets the collection mode according to instructions, including setting the collection mode according to instructions and the network download speed of the URL, the load of the URL, the security level of the URL, and the importance of the collection web page section of the URL.
[0130] In some embodiments, the private network manager 610 or the web page collector 620 sets the collection mode according to instructions and the network download speed of the URL, the load of the URL, the security level of the URL, and the importance of the URL's sections, including:
[0131] Set the interval for collecting web page content. for:
[0132]
[0133]
[0134]
[0135] )
[0136] Among them, T base Based on the data acquisition interval, R current R represents the current network download speed of the URL. max k represents the ideal maximum network download speed for a given URL. R ω is the network download speed adjustment factor. R The weighting of network download speed, L current L represents the current load of the URL. max k is the ideal maximum or threshold for the load on a URL. L ω is the load adjustment factor for the URL. L The load factor of a URL affects its weight, S current For the security level of a website, S max k represents the ideal maximum level of security for a website. S ω is the security adjustment factor for a website. S The security level of a URL affects its weight. importance Let δ represent the importance of a website's sections, and ω represent the influence coefficient of that importance. I This is a non-linear adjustment factor that affects the weight of the website's sections based on their importance.
[0137] In some embodiments, a private network manager 610 or a web page collector 620 sets a collection mode according to instructions, so that the web page collector can obtain web page content from a URL according to the collection mode. This includes: sending an access request to the URL according to the collection mode; obtaining web page content result data from the URL; obtaining a list of web page content data from the web page content result data; decomposing the list of web page content data to obtain a single data address configuration; analyzing the structure of a single data based on the single data address configuration to obtain available information; extracting the content within the collection web page category from the available information according to the collection web page category; and if the instructions include the need for translation, calling a translation program to translate the extracted content within the collection web page category.
[0138] In some embodiments, the web page collector 620 obtains web page content from a URL and stores the web page content in a data management area of a public network, wherein a private network manager 610 stores the web page content in a private network, including: extracting the web page content from the data management area and storing the web page content in a buffer of the private network, wherein the private network manager 610 communicates with both the private network and the public network; and sending the web page content from the private network manager 610 to the memory of the private network according to a predetermined push rule.
[0139] Figure 7 Another block diagram of an apparatus for collecting web page content from a public network according to at least one embodiment of the present disclosure is shown.
[0140] An apparatus for collecting web page content from a public network may include a processor 710 and a memory 720, the memory 720 being coupled to the processor 710 and storing computer instructions therein for performing the steps of various methods of at least one embodiment of the present disclosure when executed by the processor 710.
[0141] The processor 710 may include, but is not limited to, one or more processors or microprocessors.
[0142] The memory 720 may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (e.g., hard disk, floppy disk, solid-state drive, removable disk, CD-ROM, DVD-ROM, Blu-ray disc, etc.).
[0143] In addition, the device for collecting web page content from public networks may also include (but is not limited to) a data bus 730, an input / output (I / O) bus 740, a display 750, and input / output devices 760 (e.g., keyboard, mouse, speaker, etc.).
[0144] The processor 610 can communicate with external displays 750 and input / output devices 760 via the I / O bus 740.
[0145] In one embodiment, the at least one computer instruction may also be compiled into or comprise a computer program product or software product, wherein one or more computer instructions, when executed by a processor, perform the steps of the various functions and / or methods in the embodiments described herein.
[0146] This disclosure may also include non-transitory computer-readable storage media.
[0147] Instructions, such as computer instructions, are stored on a non-transitory computer-readable storage medium. When the computer instructions are executed by a processor, the various methods described above can be performed. Non-transitory computer-readable storage media include, but are not limited to, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (e.g., hard disks, floppy disks, solid-state drives, removable disks, CD-ROMs, DVD-ROMs, Blu-ray discs, etc.). For example, a non-transitory computer-readable storage medium can be connected to a computing device such as a computer, and then, when the computing device executes the computer instructions stored on the computer-readable storage medium, the various methods described above can be performed.
[0148] This disclosure may also include a computer program product capable of performing the methods, steps, and operations given herein. For example, such a computer program product may be a computer software package, computer code instructions, or a computer-readable tangible medium having computer instructions tangibly stored (and / or encoded) thereon, which can be executed by a processor to perform the operations described herein. The computer program product may include packaging materials.
[0149] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The term “such as / for example” as used herein refers to the phrase “such as / for example but not limited to,” and is used interchangeably with it.
[0150] The flowcharts and method descriptions in this disclosure are merely illustrative examples and are not intended to require or imply that the steps of the various embodiments must be performed in the given order. As those skilled in the art will recognize, the steps in the above embodiments can be performed in any order. Words such as "then," "next," etc., are not intended to limit the order of the steps; these words are only used to guide the reader through the description of these methods. Furthermore, any reference to a singular element, such as the use of the articles "a," "one," or "the," is not to be construed as limiting that element to the singular.
[0151] Furthermore, the steps and apparatus in the various embodiments herein are not limited to any one embodiment. In fact, new embodiments can be conceived by combining relevant steps and apparatus in the various embodiments herein based on the concepts of this disclosure, and these new embodiments are also included within the scope of this disclosure.
Claims
1. A method for collecting webpage content from a public network, comprising: The private network manager obtains instructions for collecting web page content from the private network, wherein the instructions include at least the URLs of the web pages on the public network to be collected; The private network manager invokes the web page crawler in the public network according to the instructions, wherein the private network and the public network do not communicate directly. The private network manager or the web page collector sets the collection mode according to the instructions, so that the web page collector can obtain the web page content from the URL according to the collection mode; The private network manager stores the webpage content on the private network. The step of setting the collection mode according to the instructions by the private network manager or the web page collector includes: Set the interval for collecting web page content. for: ) Among them, T base Based on the data acquisition interval, R current R represents the current network download speed of the URL. max k represents the ideal maximum network download speed for the given URL. R ω is the network download speed adjustment factor. R The weighting of network download speed, L current L represents the current load of the URL. max k represents the ideal maximum or threshold load for the URL. L ω is the load adjustment factor for the URL. L S is the load impact weight of the URL. current S represents the security level of the stated URL. max k represents the ideal maximum level of security for the given URL. S ω is the security adjustment factor for the URL. S The security level of the URL affects the weight, I importance Let δ represent the importance of the website's sections, and ω represent the influence coefficient of the importance of the website's sections. I This is a non-linear adjustment factor that influences the weight of the website's sections based on their importance.
2. The method according to claim 1, wherein, The instructions also include at least one of the following: webpage category to be collected, collection method, collection frequency, whether translation is required, collection content format, and filename prefix.
3. The method according to claim 1, wherein, The step of setting a collection mode according to the instructions by the private network manager or the web crawler, so that the web crawler can obtain the web page content from the URL according to the collection mode, includes: According to the collection mode, an access request is sent to the URL; Obtain webpage content result data from the URL; Obtain a list of webpage content data from the webpage content result data; Decompose the webpage content data list to obtain the individual data address configuration; Based on the configuration of the single data address, analyze the structure of the single data and obtain available information; Extract the content within the collected webpage sections from the available information according to the collected webpage sections; If the instruction includes the requirement for translation, then a translation program is invoked to translate the content extracted from the collected webpage sections.
4. The method according to claim 1, wherein, The web page crawler obtains the web page content from the URL and stores the web page content in the data management area of the public network. The storage of the webpage content on the private network by the private network manager includes: The webpage content is extracted from the data management area and stored in a buffer of the private network, wherein the private network manager communicates with the private network and the public network. The web page content is sent from the private network manager to the storage of the private network according to the predetermined push rules.
5. An apparatus for collecting web page content from a public network, comprising a private network manager and a web page collector, wherein, The private network manager obtains instructions for collecting web page content from the private network, wherein the instructions include at least the URLs of the web pages on the public network to be collected; The private network manager invokes the web page crawler in the public network according to the instructions, wherein the private network and the public network do not communicate directly. The private network manager or the web page collector sets the collection mode according to the instructions, so that the web page collector can obtain the web page content from the URL according to the collection mode; The private network manager stores the webpage content on the private network. The interval for collecting web page content is set by the private network manager or by the web page collector. for: ) Among them, T base Based on the data acquisition interval, R current R represents the current network download speed of the URL. max k represents the ideal maximum network download speed for the given URL. R ω is the network download speed adjustment factor. R The weighting of network download speed, L current L represents the current load of the URL. max k represents the ideal maximum or threshold load for the URL. L ω is the load adjustment factor for the URL. L S is the load impact weight of the URL. current S represents the security level of the stated URL. max k represents the ideal maximum level of security for the given URL. S ω is the security adjustment factor for the URL. S The security level of the URL affects the weight, I importance Let δ represent the importance of the website's sections, and ω represent the influence coefficient of the importance of the website's sections. I This is a non-linear adjustment factor that influences the weight of the website's sections based on their importance.
6. An apparatus for collecting web page content from a public network, comprising: Memory, which stores computer instructions; At least one processor is configured to execute the computer instructions in the memory to perform the method according to any one of claims 1-4.
7. A non-transitory computer-readable storage medium having computer instructions stored thereon, in, When the computer instructions are executed by a processor, the processor performs the method according to any one of claims 1-4.
8. A computer program product having computer instructions stored thereon, in, When the computer instructions are executed by a processor, the processor performs the method according to any one of claims 1-4.
Citation Information
Patent Citations
HTTP network access achieving method based on serial communication
CN101651711A
Network data acquisition method and device, computer equipment and storage medium
CN112818201A