Public website information acquisition method

By setting up an error processing module in the public website information collection method to detect and handle all modules, the problem of imperfect error processing mechanism in the existing technology is solved, and the reliability and user experience of the collection are improved.

CN119961511APending Publication Date: 2025-05-09XINJIANG LIANHAI INA INT INFORMATION TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411944956.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

When existing public website information collection methods face network problems or changes in the source site structure, the error handling mechanism is incomplete, resulting in a high collection failure rate and the inability to notify users in time and provide solutions.

Method used

During the information collection process, an error processing module is set up to detect all modules, including URL format verification, collection condition verification, network request detection and database data format compatibility detection, timely output the error cause and notify the user. The user can correct the error based on the error cause.

Benefits of technology

By detecting and error processing of all modules, the reliability and stability of information collection of public websites is improved, the collection failure rate is reduced, and the user experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961511A_ABST
    Figure CN119961511A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of information acquisition, and particularly discloses a public website information acquisition method, which comprises the following steps that: a user firstly inputs a URL (Uniform Resource Locator) address of a source website, and an error processing module performs format verification on the URL; after entering the webpage, the user sets a collection condition, and the error processing module verifies the collection condition; the information acquisition module sends an HTTP request according to a URL address and an acquisition condition, and the error processing module regularly detects a network request, extracts page content and sorts and packs information; the information is stored in a database, and an error processing module detects the database regularly; the information in the database is analyzed, restored and subjected to format unification, and a collection result is output to the user interaction module; a user can receive an acquisition result or an error prompt at a user side, and the user feeds back the error prompt at the user side, so that the problem that the reliability of public website information acquisition is reduced due to the lack of a detection function for each module during public website information acquisition is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information collection, and in particular to a method for collecting information from a public website. Background Art

[0002] Public website information collection refers to the process of collecting and organizing information on public websites through web crawlers or related tools. This process usually includes steps such as determining collection tasks, configuring collection parameters, scheduling collection tasks, processing collection results, and publishing data to application platforms. Problems in the existing public information collection field include: limited collection source stations, rigid collection strategies, incomplete information acquisition, poor stability and adaptability; when facing network problems (such as network fluctuations, temporary maintenance of source stations) or changes in source station structure (such as page layout adjustments, URL rule changes), due to its imperfect error handling mechanism, it is impossible to automatically retry and cannot reasonably adjust the retry strategy according to the error type. For example, for the problem of changes in source station structure, it is impossible to record error information in real time and relocate the information position through intelligent analysis, resulting in a high collection failure rate and failure to notify users and provide solutions in time. Therefore, it is necessary to set up detection functions in each module to detect all modules, thereby increasing the reliability of public website information collection. Summary of the invention

[0003] The purpose of the present invention is to provide a method for collecting information from public websites, so as to solve the problem that the reliability of collecting information from public websites is reduced due to the lack of setting detection functions for each module during the collection of information from public websites.

[0004] To achieve the above purpose, the present invention provides a basic solution: a method for collecting information from a public website, comprising the following steps: S1. The administrator sets the time interval for scheduled collection in the system in advance. The user enters the URL address of the source station for which information needs to be collected in the management interface of the source station management module. The error handling module performs format verification on the input URL. S2. After entering the web page, the user sets the collection conditions according to actual needs in the user end of the collection strategy formulation module, and the error handling module verifies the collection conditions; S3. The information collection module sends HTTP requests from the Python network request library to the source site according to the URL address of the source site and the collection conditions set by the user. The error handling module regularly detects the network requests, obtains, parses and extracts the page content, and then organizes and packages the collected information. S4. The information collection module then stores the packaged information in the database, and the error handling module regularly checks the stored database and cleans and deduplicates the restored information in the database; S5. The information processing module parses and restores the information of the source station, then unifies the format of different information and integrates the relevant information, and outputs the collection results to the user interaction module; S6. In the user interaction module, the user can receive the collection results or error prompts on the user side, and the user can provide feedback on the error prompts on the user side; S7. After receiving the collection results, the user triggers the information collection module in S3. The information collection module collects information regularly according to the preset time interval in S1, the source site URL address and the collection conditions in S2, and stores the latest collected information in the database in S4.

[0005] The principle and beneficial effect of the present invention are as follows: in the process of information collection, the error handling module increases the reliability of public website information collection by detecting all modules; in the source station management module, by format checking the input URL, the error handling module outputs the detected error cause to the user, and the user can correct the URL format according to the error cause; in the collection strategy formulation module, by verifying the collection conditions, the detected error cause is output to the user, and the user enters the correct collection conditions according to the error cause; in the information collection module, by detecting the network request, the user can be notified in time, and by detecting the stored database, the user can be notified in time, which can increase the reliability of public website information collection; when the user collects the same information again, the timed execution of information collection can output the latest information to the user in a short time, thereby increasing the user experience.

[0006] Solution 2 is the preferred basic solution. In step S1, the error handling module verifies the source site URL address format. If the verification fails, the error cause is stored and the verification failure cause is output to the user interaction module, prompting the user to enter the correct format. If the connection is successful, step S2 is entered. By verifying the URL format, URLs containing illegal characters or not meeting security standards can be discovered and avoided in a timely manner, thereby reducing potential security risks. It also helps to reduce errors and abnormal situations, improve the reliability and security of the website, and help search engines better crawl and index website content, thereby increasing the reliability of public website information collection.

[0007] Solution three is the preferred basic solution. In step S2, the error handling module performs syntax and logic verification on the collection conditions to check whether the keyword format is correct, whether the time range is reasonable, and whether the dynamic address splicing rules are executable. If the verification fails, the error handling module will output the error information for storage, and output the reason for the verification failure to the user interaction module, prompting the user to modify the collection conditions; if the collection conditions are verified, they will be saved in the policy database, and a unique identifier will be assigned to the policy, and then enter step S3; verifying the collection conditions can ensure that the collected data is accurate and reliable, and it can also help avoid logical contradictions or errors, further ensure the accuracy of the data, ensure that users obtain the data they really need, avoid confusion or dissatisfaction caused by data errors or omissions, and increase the user experience.

[0008] Solution 4 is the preferred basic solution. In step S3, Python's parsing library (such as BeautifulSoup or lxml library) is used to parse the page, and BeautifulSoup is used to extract the page content. The error handling module detects whether the web page is connectable every 3 hours. If the connection fails, the error reasons such as network interruption, timeout, and server return of error status code are recorded, and the error reasons are output to the user interaction module. If the connection is successful, the page content is obtained. Through regular detection, these risky links can be discovered and isolated in time, thereby protecting the data security of the website and users, and ensuring that these web page functions can work normally, such as form submission, link jump, shopping cart processing, etc., to avoid functional failures caused by link problems, thereby maintaining the overall stability and reliability of the website.

[0009] Solution 5 is the preferred basic solution. In step S4, combined with database unique constraints and index optimization, as well as temporary storage and comparison of data structure cleaning and deduplication restoration information, the error handling module checks whether the data format is compatible and whether data insertion or update fails. If incompatible or failed, the error cause is recorded and output to the user interaction module; if the data is compatible or successful, information cleaning and deduplication are performed; compatible data formats can ensure that data remains consistent during transmission and storage, avoiding data distortion or loss due to format conversion errors.

[0010] Solution 6 is the preferred basic solution. In step S5, Python's text processing and format conversion library is used to convert information in different formats into information in a unified format. The unified format helps to speed up information transmission and reduce delays caused by format conversion, while having better compatibility and scalability.

[0011] Option seven is the preferred option among options three to five. When the information collection module receives the source site URL address and the relevant collection conditions in the database, the information collection module will automatically call out the related information according to the identifier in S2, and deduplicate the information in the identifier with the updated information in S4, and then enter step S5. This can reduce the time for information collection, speed up the output of information, and thus increase the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a flow chart of a method for collecting information from a public website of the present invention. DETAILED DESCRIPTION

[0013] The present invention is further described in detail below through specific embodiments: Example like Figure 1 As shown: A method for collecting information from a public website, comprising the following steps: S1. The administrator sets the time interval for scheduled collection in the system in advance. The user enters the URL address of the source station for which information needs to be collected in the management interface of the source station management module. The error handling module performs format verification on the input URL. If the verification fails, the error cause is stored and the verification failure reason is output to the user interaction module, prompting the user to enter the correct format; if the connection is successful, proceed to step S2; S2. After entering the web page, the user sets the collection conditions according to actual needs on the user side of the collection strategy formulation module. The error handling module verifies the collection conditions, performs syntax and logic verification on the collection conditions, checks whether the keyword format is correct, whether the time range is reasonable, and whether the dynamic address splicing rules are executable. If the verification fails, the error handling module will output the error information for storage, and output the verification failure reason to the user interaction module, prompting the user to modify the collection conditions; if the collection condition verification passes, it will be saved in the policy database, and a unique identifier will be assigned to the policy, and then enter step S3; S3. The information collection module sends HTTP requests from the Python network request library to the source station according to the URL address of the source station and the collection conditions set by the user. The error handling module regularly detects the network request and checks whether the web page can be connected every 3 hours. If the connection fails, the error reasons such as network connection interruption, timeout, and server return error status code are recorded, and the error reasons are output to the user interaction module. If the connection is successful, the page content is obtained. After obtaining the page content, the page is parsed using the Python parsing library (such as BeautifulSoup or lxml library), and the page content is extracted using BeautifulSoup, and then the collected information is sorted and packaged. S4. The information collection module then stores the packaged information in the database. The error handling module regularly checks the stored database to see if the data format is compatible and if data insertion or update fails. If not, the error cause is recorded and output to the user. If the data is compatible or successful, the information is cleaned and deduplicated, combined with database unique constraints and index optimization, as well as temporary storage and comparison of data structures, cleaned and deduplicated information. When the information collection module receives the source site URL address and relevant collection conditions in the database, the information collection module automatically calls out the associated information according to the identifier in S2, and deduplicates the information in the identifier with the information updated in S4, and then enters step S5. S5. The information processing module parses and restores the information of the source station, and then unifies the format of different information. It uses Python's text processing and format conversion library to convert information of different formats into information of a unified format, and outputs the collection results to the user interaction module; S6. In the user interaction module, the user can receive the collection results or error prompts on the user side, and the user can provide feedback on the error prompts on the user side; S7. After receiving the collection results, the user triggers the information collection module in S3. The information collection module collects information regularly according to the preset time interval in S1, the source site URL address and the collection conditions in S2, and stores the latest collected information in the database in S4.

[0014] The implementation method of this embodiment is as follows: the administrator sets the time interval for scheduled collection in the system in advance, the user enters the URL address of the source station whose information needs to be collected in the management interface of the source station management module, the error handling module performs format verification on the input URL, if the verification fails, the error cause is stored, and the verification failure reason is output to the user interaction module, prompting the user to enter the correct format; if the connection is successful, the user sets the collection conditions according to actual needs on the user end of the collection strategy formulation module, the error handling module verifies the collection conditions, if the verification fails, the error handling module outputs the error information for storage, and outputs the verification failure reason to the user interaction module, prompting the user to modify the collection conditions, if the collection conditions are verified, they are saved in the policy database, and a unique identifier is assigned to the policy, the user corrects the error according to the prompted error on the client, and then obtains the collection results on the client.

[0015] The information collection module in the system sends HTTP requests from Python's network request library to the source station according to the URL address of the source station and the collection conditions set by the user. The error handling module regularly detects network requests and checks whether the web page can be connected every 3 hours. If the connection fails, the error reasons such as network connection interruption, timeout, and server return of error status code are recorded, and the error reasons are output to the user interaction module. If the connection is successful, the page content is obtained. After obtaining the page content, the page is parsed using Python's parsing library (such as BeautifulSoup or lxml library), and the page content is extracted using BeautifulSoup. Then the collected information is sorted and packaged, and the packaged information is stored in the database. The error handling module regularly checks the stored database to see if the data format is compatible and whether the data insertion or update fails. If it is incompatible or fails, the error reason is recorded and output to the user. If the data is compatible or successful, the information is cleaned and deduplicated, combined with database unique constraints and index optimization, as well as temporary storage of data structure and comparison of cleaning and deduplication restoration information. When the information collection module receives the source site URL address and related collection conditions in the database, the information collection module will automatically call out the related information according to the identifier in the policy database, and deduplicate the information in the identifier and the re-collected and updated information. The information processing module will parse and restore the information of the source site, and then unify the formats of different information. It will use Python's text processing and format conversion library to convert information in different formats into information in a unified format, and output the collection results to the user interaction module. After receiving the collection results, the user will trigger the information collection module in S3. The information collection module will collect information regularly according to the preset time interval in S1, the source site URL address and the collection conditions in S2, and store the latest collected information in the database in S4.

[0016] The above is only an embodiment of the present invention, and the common knowledge such as the known specific structure and characteristics in the scheme is not described in detail here. It should be pointed out that for those skilled in the art, several deformations and improvements can be made without departing from the structure of the present invention, which should also be regarded as the protection scope of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A method for collecting information from a public website, characterized in that: The following steps are involved: S1. The administrator sets the time interval for scheduled collection in the system in advance. The user enters the URL address of the source station for which information needs to be collected in the management interface of the source station management module. The error handling module performs format verification on the input URL. S2. After entering the web page, the user sets the collection conditions according to actual needs in the user end of the collection strategy formulation module, and the error handling module verifies the collection conditions; S3. The information collection module sends HTTP requests from the Python network request library to the source site according to the URL address of the source site and the collection conditions set by the user. The error handling module regularly detects the network requests, obtains, parses and extracts the page content, and then organizes and packages the collected information. S4. The information collection module then stores the packaged information in the database, and the error handling module regularly checks the stored database and cleans and deduplicates the restored information in the database; S5. The information processing module parses and restores the information of the source station, then unifies the format of different information and integrates the relevant information, and outputs the collection results to the user interaction module; S6. In the user interaction module, the user can receive the collection results or error prompts on the user side, and the user can provide feedback on the error prompts on the user side; S7. After receiving the collection results, the user triggers the information collection module in S3. The information collection module collects information regularly according to the preset time interval in S1, the source site URL address and the collection conditions in S2, and stores the latest collected information in the database in S4.

2. A method for collecting information from public websites according to claim 1, characterized in that: In step S1, the error handling module verifies the source site URL address format. If the verification fails, the error cause is stored and the verification failure reason is output to the user interaction module, prompting the user to enter the correct format. If the connection is successful, step S2 is entered.

3. A method for collecting information from a public website according to claim 1, characterized in that: In step S2, the error handling module performs syntax and logic verification on the collection conditions, checks whether the keyword format is correct, whether the time range is reasonable, and whether the dynamic address splicing rules are executable. If the verification fails, the error handling module will output the error information for storage, and output the reason for the verification failure to the user interaction module, prompting the user to modify the collection conditions; if the collection conditions are verified, they will be saved in the policy database, and a unique identifier will be assigned to the policy, and then enter step S3.

4. A method for collecting information from a public website according to claim 1, characterized in that: In step S3, the page is parsed using Python's parsing library, and the page content is extracted using BeautifulSoup. The error handling module detects whether the web page is connectable every 3 hours. If the connection fails, the error reasons such as network connection interruption, timeout, and server return error status code are recorded, and the error reasons are output to the user interaction module. If the connection is successful, the page content is obtained.

5. A method for collecting information from public websites according to claim 1, characterized in that: In step S4, combined with database unique constraints and index optimization, as well as temporary storage and comparison of data structure cleaning and deduplication restoration information, the error handling module checks whether the data format is compatible and whether data insertion or update fails. If incompatible or failed, the cause of the error is recorded and output to the user interaction module; if the data is compatible or successful, information cleaning and deduplication are performed.

6. A method for collecting information from public websites according to claim 1, characterized in that: In step S5, the information in different formats is converted into information in a unified format using Python's text processing and format conversion library.

7. A method for collecting information from public websites according to any one of claims 3 to 5, characterized in that: When the information collection module receives the source site URL address and the relevant collection conditions in the database, the information collection module automatically calls out the associated information according to the identifier in S2, and deduplicates the information in the identifier with the updated information in S4, and then enters step S5.