License collection monitoring method and system

Through the distributed RPA system combined with rule configuration and large language model analysis engine, the problems of low efficiency and insufficient accuracy of license collection are solved, efficient and accurate automatic license collection is achieved, and the rapid review needs of merchants for online access are met.

CN120492261APending Publication Date: 2025-08-15BAOFOO COM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510534845.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The current technology has low efficiency in the collection of certificates and licenses, insufficient data consistency and accuracy, traditional RPA tools lack intelligent understanding capabilities, and cannot achieve full process automation, resulting in high manual operation costs and uneven data quality, which cannot meet the rapidly developing merchant network access needs.

Method used

It adopts a distributed RPA system, combining the rule configuration analysis engine and the Large Language Model (LLM) analysis engine, obtains license information through web parsing, and processes it in the Kubernetes cluster, realizes intelligent positioning and screenshots, supports automatic retry, and ensures successful completion of the task.

Benefits of technology

It significantly improves the efficiency of license collection, shortens the review time for merchants to access the Internet, improves data accuracy and identification capabilities, reduces labor costs, and meets the efficient and accurate merchants to access the Internet.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492261A_ABST
    Figure CN120492261A_ABST
Patent Text Reader

Abstract

The invention provides a certificate collection monitoring method and system, and solves the problems of low certificate collection efficiency and incapability of accurately identifying key information in the prior art. The method specifically comprises the following steps: acquiring a certificate collection request task, and analyzing the task; a distributed RPA is created, and an execution environment is initialized; the distributed RPA is started, the task is executed through the distributed RPA, and a collection result is formed; uploading an acquisition result to a data storage system, and updating a task execution state; and the task execution process is monitored, and abnormal conditions are processed, so that the license collection efficiency is improved, and the accuracy of key information identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of certificate collection, and in particular to a certificate collection monitoring method and system. Background Art

[0002] Currently, financial institutions and other enterprises generally rely on manual operations to collect corporate license information for customer access and risk management. The existing methods have the following shortcomings:

[0003] 1. Low efficiency and high cost

[0004] Customer account opening due diligence (KYC) is often cumbersome and time-consuming. Staff must manually access multiple web platforms to take screenshots and download information, resulting in long review cycles and high labor costs.

[0005] 2. Insufficient data consistency and accuracy

[0006] Manual operations are easily affected by factors such as fatigue and negligence, and data collection is prone to errors or omissions. Frequent personnel turnover also leads to uneven data quality.

[0007] 3. Automation tools lack intelligence

[0008] Although traditional distributed RPA (robotic process automation) can partially replace manual operations, it lacks the ability to intelligently understand and locate complex web page structures and dynamically changing content, making it difficult to achieve full process automation.

[0009] In addition, although there are general-purpose AI intelligent agent products such as Manus on the current market, which focus on the autonomous planning and execution of general tasks, in enterprise-level application scenarios (such as financial risk control and compliance monitoring), the requirements are more clear and standardized, and there is an urgent need for a dedicated system that can both efficiently collect data and ensure its accuracy.

[0010] In the payment industry, third-party payment institutions, as key players in payment services, are required by regulatory requirements to conduct strict merchant access management and continuous compliance monitoring. During merchant onboarding, payment institutions are required to conduct comprehensive due diligence on multiple documents, including business licenses, operating permits, and legal representative ID cards. They also continuously monitor the validity of existing merchant documents to mitigate fraud risks and ensure compliance.

[0011] Currently, third-party payment institutions rely primarily on manual processes when conducting merchant due diligence. Staff must manually access dozens of websites and data platforms, including those for industrial and commercial registrations, Ministry of Industry and Information Technology filings, and financial licenses, to query and archive licenses and take screenshots. The network access review cycle for a single merchant typically takes two to three business days. This manual operation model is not only inefficient, but also suffers from repetitive, mundane tasks and frequent staff turnover, resulting in uneven data quality. This makes it impossible to meet the demands of the large number of merchants joining the network during the rapid growth of payment institutions. Furthermore, while traditional distributed RPA tools can partially replace manual operations, they lack intelligent understanding capabilities and are unable to accurately identify and locate key information on web pages. Summary of the Invention

[0012] The present invention provides a method and system for monitoring the collection of certificates and licenses to solve the above problems.

[0013] In a first aspect, the present invention provides a method for monitoring the collection of certificates and licenses, which specifically includes the following steps:

[0014] Step S1: Obtain a certificate collection request task, parse the task, and form a parsed task; wherein the request includes necessary information such as the company name and the target webpage;

[0015] Step S2: Create a distributed RPA based on the parsed task type and initialize the execution environment; wherein the distributed RPA includes an RPA scheduler (for allocating RPA tasks) and a Kubernetes cluster (for processing RPA tasks);

[0016] Step S3: Start the distributed RPA and execute the task through the distributed RPA. The task execution enters the target data source platform, performs a precise search on the target data source platform based on the enterprise name, and intelligently locates and screenshots the search results to form a collection result.

[0017] Step S4: Upload the collection results to the data storage system, push the collection result notification to relevant personnel, and update the task execution status to complete the task closed loop;

[0018] Step S5: monitor the task execution process, and automatically retry the task if any abnormality occurs during the task execution process.

[0019] Preferably, in step S1, the task is parsed by a web page parsing engine; wherein, the web page parsing engine includes a rule configuration parsing engine (the rule configuration parsing engine is based on a traditional selector mechanism, and is mainly used to process web pages with stable structures) and a large language model (LLM) parsing engine (the LLM parsing engine is based on OmniParser2 technology, and is mainly used to process complex and changeable web page structures).

[0020] Preferably, parsing the task by a web page parsing engine specifically includes the following steps:

[0021] Step S101: The parsing scheduler selects a rule-configured parsing engine and parses the task using the rule-configured parsing engine.

[0022] Step S102: When the rule configuration parsing engine task fails to parse, the parsing scheduler selects the LLM parsing engine and parses the task through the LLM parsing engine.

[0023] Preferably, parsing the task by the rule configuration parsing engine specifically includes the following steps:

[0024] Step S101a: select and configure a selector according to the task;

[0025] Step S101b: Match the corresponding parsing rules through the rule matcher according to the URL pattern of the target web page;

[0026] Step S101c: accurately locate the page element according to the selector path;

[0027] Step S101d: Generate rule parsing results according to the parsing rules.

[0028] Preferably, in step S101a, the selector includes one or more of an XPath selector, a CSS selector, and a regular expression.

[0029] Preferably, the parsing rules provide a rule self-verification mechanism. When the parsing rules fail, an alarm is triggered and the parsing scheduler selects the LLM parsing engine to parse the task.

[0030] Preferably, when the rule configuration parsing engine fails to parse the task, the parsing scheduler selects the LLM parsing engine and parses the task through the LLM parsing engine, specifically including the following steps:

[0031] Step S102a: acquiring the DOM structure and visual layout information of the target web page according to the task;

[0032] Step S102b: inputting the DOM structure and visual layout information of the target webpage into the large language model;

[0033] Step S102c: describing the positioning target information in natural language;

[0034] Step S102d: accurately identify and extract target elements based on the visual layout information and the positioning target information to form an LLM analysis result.

[0035] Preferably, the tasks successfully parsed by the LLM parsing engine are recorded, new parsing rules are formed through the rule generator, and the new parsing rules are saved to the rule database; wherein, the rule database is used to save the rule configuration; the rule configuration parsing engine includes a rule updater, and the rule updater and the rule matcher are used to update the rules through the rule database.

[0036] Preferably, the parsing scheduler can dynamically adjust the usage priority of the two engines according to the history of task parsing, so as to balance the speed and success rate of task parsing.

[0037] Preferably, in step S2, the construction of the distributed RPA is implemented as follows:

[0038] a) Implemented through kasmweb and DrissionPage;

[0039] b) Implemented via Selenium Grid;

[0040] c) implemented through Playwright Grid;

[0041] d) Deploy multiple RPA nodes directly through Docker containers and implement them through a simple load balancer.

[0042] Preferably, in step S2, the Kubernetes cluster can be replaced by Docker Swarm.

[0043] Preferably, in step S2, a single-machine multi-process mode can be adopted to ensure reliable operation of the service through a process management tool (such as PM2).

[0044] In a second aspect, the present invention further provides a certificate collection and monitoring system, which specifically includes the following modules:

[0045] A task parsing module is used to obtain a certificate collection request task, parse the task, and form a parsed task; wherein the request includes necessary information such as the company name and the target webpage;

[0046] A distributed RPA construction module is used to create a distributed RPA based on the parsed task type and initialize the execution environment; wherein the distributed RPA includes an RPA scheduler (for allocating RPA tasks) and a Kubernetes cluster (for processing RPA tasks);

[0047] A collection process generation module is used to start the distributed RPA, execute the task through the distributed RPA, enter the target data source platform for execution of the task, perform a precise search on the target data source platform based on the enterprise name, and intelligently locate and screenshot the search results to form a collection result;

[0048] The task execution status update module is used to upload the collection results to the data storage system, push the collection result notification to relevant personnel, and update the task execution status to complete the task closed loop;

[0049] The exception monitoring module is used to monitor the task execution process and automatically retry the task if an exception occurs during the task execution.

[0050] Preferably, in the task parsing module, the task is parsed by a web page parsing engine; wherein, the web page parsing engine includes a rule configuration parsing engine (the rule configuration parsing engine is based on the traditional selector mechanism, mainly used to process web pages with stable structures) and a large language model (LLM) parsing engine (the LLM parsing engine is based on OmniParser2 technology, mainly used to process complex and changeable web page structures).

[0051] Preferably, in the task parsing module, the task is parsed by a web page parsing engine, which specifically includes the following submodules:

[0052] A rule configuration parsing submodule, configured to select a rule configuration parsing engine according to a parsing scheduler, and parse the task through the rule configuration parsing engine;

[0053] The LLM parsing submodule is used for selecting the LLM parsing engine by the parsing scheduler to parse the task through the LLM parsing engine when the parsing of the rule configuration parsing engine task fails.

[0054] Preferably, the web page parsing submodule specifically includes the following units:

[0055] A rule configuration parsing first unit is used to select and configure a selector according to the task;

[0056] The second rule configuration parsing unit is used to match the corresponding parsing rules through the rule matcher according to the URL pattern of the target web page;

[0057] The third unit of rule configuration parsing is used to accurately locate page elements based on the selector path;

[0058] The fourth unit of rule configuration parsing is used to generate rule parsing results according to the parsing rules.

[0059] Preferably, in the first rule configuration parsing unit, the selector includes one or more of an XPath selector, a CSS selector, and a regular expression.

[0060] Preferably, the parsing rules provide a rule self-verification mechanism. When the parsing rules fail, an alarm is triggered and the parsing scheduler selects the LLM parsing engine to parse the task.

[0061] Preferably, the LLM parsing submodule specifically includes the following units:

[0062] The first LLM parsing unit is used to obtain the DOM structure and visual layout information of the target web page according to the task;

[0063] The second LLM parsing unit is used to input the DOM structure and visual layout information of the target web page into the large language model;

[0064] The third unit of LLM parsing is used to describe the positioning target information through natural language;

[0065] The fourth LLM parsing unit is used to accurately identify and extract target elements according to the visual layout information and the positioning target information to form an LLM parsing result.

[0066] Preferably, the tasks successfully parsed by the LLM parsing engine are recorded, new parsing rules are formed through the rule generator, and the new parsing rules are saved to the rule database; wherein, the rule database is used to save the rule configuration; the rule configuration parsing engine includes a rule updater, and the rule updater and the rule matcher are used to update the rules through the rule database.

[0067] Preferably, the parsing scheduler can dynamically adjust the usage priority of the two engines according to the history of task parsing, so as to balance the speed and success rate of task parsing.

[0068] Preferably, in the distributed RPA construction module, the construction of the distributed RPA is implemented as follows:

[0069] a) Implemented through kasmweb and DrissionPage;

[0070] b) Implemented via Selenium Grid;

[0071] c) implemented through Playwright Grid;

[0072] d) Deploy multiple RPA nodes directly through Docker containers and implement them through a simple load balancer.

[0073] Preferably, in the distributed RPA building module, the Kubernetes cluster can be replaced by Docker Swarm.

[0074] Preferably, in the distributed RPA building module, a single-machine multi-process mode can be adopted, and process management tools (such as PM2) can be used to ensure reliable service operation.

[0075] In a third aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for monitoring the collection of certificates and our licenses as described in any one of the first aspects of this application.

[0076] In a fourth aspect, the present invention further provides an electronic device comprising: a memory storing a computer program; and a processor communicatively connected to the memory, for executing a method for monitoring credential collection as described in any one of the first aspects of the present application when the computer program is called.

[0077] Compared with the prior art, the present invention has the following obvious outstanding substantial features and significant advantages:

[0078] 1. Improved efficiency

[0079] Manual collection of each business license takes an average of 30 to 45 minutes, while this system reduces the collection time to 3 to 5 minutes, increasing efficiency by approximately 90%. The system supports 24 / 7 continuous operation, with a time coverage rate 200% higher than the traditional 8-hour work system. The merchant access review cycle is shortened.

[0080] 2. Improved risk control accuracy

[0081] Through the system's automatic identification and intelligent positioning, risk detection is more efficient and accurate than manual verification. This effectively improves the efficiency and accuracy of due diligence during daily merchant access reviews and prevents high-risk merchants from joining the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings:

[0083] Figure 1 This is a flow chart of a certificate collection and monitoring method according to a preferred embodiment of the present invention.

[0084] Figure 2 This is a diagram of the architecture of a web page parsing engine according to a preferred embodiment of the present invention.

[0085] Figure 3 This is a distributed RPA architecture diagram of a preferred embodiment of the present invention.

[0086] Figure 4 It is a structural diagram of a certificate collection and monitoring system according to a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0087] The present invention provides a method and system for monitoring the collection of licenses. To make the objectives, technical solutions, and effects of the present invention more clear and explicit, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the present invention.

[0088] It should be noted that the terms "first," "second," and the like in the specification and claims of the present invention and the accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatuses.

[0089] Example 1:

[0090] like Figure 1-Figure 3 As shown, the method for collecting and monitoring certificates in this embodiment specifically includes the following steps:

[0091] Step S1: Obtain a certificate collection request task, parse the task, and form a parsed task; wherein the request includes necessary information such as the company name and the target webpage.

[0092] Wherein, the task is parsed by a web page parsing engine; wherein, Figure 2 As shown, the web page parsing engine includes a rule configuration parsing engine (the rule configuration parsing engine is based on the traditional selector mechanism and is mainly used to process web pages with stable structures) and a large language model (LLM) parsing engine (the LLM parsing engine is based on OmniParser2 technology and is mainly used to process complex and changeable web page structures).

[0093] Optionally, the task is parsed by a web page parsing engine, specifically including the following steps S101 to S102.

[0094] Step S101: The parsing scheduler selects a rule-configured parsing engine, and parses the task using the rule-configured parsing engine.

[0095] Optionally, the task is parsed by the rule configuration parsing engine, specifically including the following steps S101a to S101d.

[0096] Step S101a: Select and configure a selector according to the task; wherein the selector includes one or more of an XPath selector, a CSS selector, and a regular expression.

[0097] Step S101b: Match the corresponding parsing rules through a rule matcher according to the URL pattern of the target web page.

[0098] Step S101c: accurately locate the page element according to the selector path.

[0099] Step S101d: Generate rule parsing results according to the parsing rules.

[0100] The parsing rules provide a self-verification mechanism. When the parsing rules become invalid, an alarm is triggered and the parsing scheduler selects the LLM parsing engine to parse the task.

[0101] Step S102: When the rule configuration parsing engine task fails to parse, the parsing scheduler selects the LLM parsing engine and parses the task through the LLM parsing engine.

[0102] Optionally, when the rule configuration parsing engine fails to parse the task, the parsing scheduler selects the LLM parsing engine and parses the task through the LLM parsing engine, specifically including the following steps S102a to S102d.

[0103] Step S102a: According to the task, obtain the DOM structure and visual layout information of the target web page.

[0104] Step S102b: Input the DOM structure (i.e., Document Object Model, which is a data representation of objects that constitute the structure and content of a document on the web, representing the page so that the program can change the structure, style, and content of the document) and visual layout information of the target web page into the large language model.

[0105] Step S102c: Describe the positioning target information using natural language.

[0106] Step S102d: accurately identify and extract target elements based on the visual layout information and the positioning target information to form an LLM analysis result.

[0107] Among them, the tasks successfully parsed by the LLM parsing engine are recorded, new parsing rules are formed through the rule generator, and the new parsing rules are saved to the rule database; wherein, the rule database is used to save rule configurations; the rule configuration parsing engine includes a rule updater, and the rule updater and the rule matcher are used to update the rules through the rule database.

[0108] The parsing scheduler can dynamically adjust the usage priority of the two engines according to the history of task parsing, and balance the speed and success rate of task parsing.

[0109] The dual parsing engine design not only ensures efficient processing of websites with stable structures, but also ensures adaptability to websites with complex structures or frequent structural changes, significantly improving the overall parsing success rate and adaptability.

[0110] Step S2: Create a distributed RPA based on the parsed task type and initialize the execution environment; wherein the distributed RPA includes an RPA scheduler (for allocating RPA tasks) and a Kubernetes cluster (for processing RPA tasks);

[0111] Optionally, in step S2, the distributed RPA can be constructed in four ways: a) through kasmweb and DrissionPage; b) through Selenium Grid; c) through Playwright Grid; d) through direct deployment of multiple RPA nodes through Docker containers and through a simple load balancer. In this embodiment, the distributed RPA is constructed in way a), as shown in the following example. Figure 3 shown.

[0112] Kasmweb is an open-source streaming browser platform based on Docker containers. It provides the ability to run a complete Ubuntu desktop system in a container. It has the following features: (1) Based on WebRTC technology: The browser interface running in the container is transmitted to the user end in real time through the WebRTC protocol; (2) Complete desktop environment: Each container contains a lightweight Linux desktop environment that can run browsers such as Chrome and Firefox; (3) Secure isolation: Each session runs in an independent container, ensuring complete isolation of data and execution environment; (4) Resource efficiency: Compared with traditional virtual machines, the containerization solution significantly reduces resource consumption, and a server can run dozens of instances simultaneously. In this invention, Kasmweb is used as the basic environment for RPA execution, solving the resource waste and environment consistency problems in traditional RPA deployment.

[0113] DrissionPage is a Python-based web automation tool, which serves as the RPA execution engine in this invention.

[0114] The Kubernetes cluster can be replaced by Docker Swarm as a lightweight alternative, which is particularly suitable for small and medium-sized deployment scenarios.

[0115] Optionally, in a resource-constrained environment, you can adopt a single-machine multi-process mode and use process management tools (such as PM2) to ensure reliable service operation.

[0116] Step S3: Start the distributed RPA and execute the task through the distributed RPA. The task execution enters the target data source platform, performs a precise search on the target data source platform based on the enterprise name, and intelligently locates and screenshots the search results to form a collection result.

[0117] Step S4: Upload the collection results to the data storage system, push the collection result notification to relevant personnel, and update the task execution status to complete the task closed loop;

[0118] Step S5: monitor the task execution process, and automatically retry the task if an abnormality occurs during the task execution process.

[0119] Example 2:

[0120] like Figure 4 As shown, the certificate collection and monitoring system described in this embodiment specifically includes a task parsing module, a distributed RPA construction module, a collection process generation module, a task execution status update module and an abnormality monitoring module.

[0121] The task parsing module is used to obtain the certificate collection request task, parse the task, and form a parsed task; wherein the request includes necessary information such as the company name and the target web page.

[0122] Among them, in the task parsing module, the task is parsed by the web page parsing engine; wherein, the web page parsing engine includes a rule configuration parsing engine (the rule configuration parsing engine is based on the traditional selector mechanism, mainly used to process web pages with stable structures) and a large language model (LLM) parsing engine (the LLM parsing engine is based on OmniParser2 technology, mainly used to process complex and changeable web page structures).

[0123] The task parsing module parses the task through a web page parsing engine, and specifically includes a rule configuration parsing submodule and an LLM parsing submodule.

[0124] The rule configuration parsing submodule is used to select a rule configuration parsing engine according to the parsing scheduler, and parse the task through the rule configuration parsing engine.

[0125] The rule configuration parsing submodule specifically includes a rule configuration parsing first unit, a rule configuration parsing second unit, a rule configuration parsing third unit and a rule configuration parsing fourth unit.

[0126] The rule configuration parsing first unit is used to select and configure a selector according to the task; wherein the selector includes one or more of an XPath selector, a CSS selector, and a regular expression.

[0127] The second rule configuration parsing unit is used to match the corresponding parsing rules through the rule matcher according to the URL pattern of the target web page.

[0128] The third unit of rule configuration parsing is used to accurately locate page elements based on the selector path.

[0129] The fourth unit of rule configuration parsing is used to generate rule parsing results according to the parsing rules.

[0130] The parsing rules provide a self-verification mechanism. When the parsing rules become invalid, an alarm is triggered and the parsing scheduler selects the LLM parsing engine to parse the task.

[0131] The LLM parsing submodule is used for selecting the LLM parsing engine by the parsing scheduler to parse the task through the LLM parsing engine when the parsing of the rule configuration parsing engine task fails.

[0132] The LLM parsing submodule specifically includes a first LLM parsing unit, a second LLM parsing unit, a third LLM parsing unit and a fourth LLM parsing unit.

[0133] The first LLM parsing unit is used to obtain the DOM structure and visual layout information of the target web page according to the task.

[0134] The second LLM parsing unit is used to input the DOM structure and visual layout information of the target web page into the large language model.

[0135] The third unit of LLM parsing is used to describe the positioning target information through natural language.

[0136] The fourth LLM parsing unit is used to accurately identify and extract target elements according to the visual layout information and the positioning target information to form an LLM parsing result.

[0137] Among them, the tasks successfully parsed by the LLM parsing engine are recorded, new parsing rules are formed through the rule generator, and the new parsing rules are saved to the rule database; wherein, the rule database is used to save rule configurations; the rule configuration parsing engine includes a rule updater, and the rule updater and the rule matcher are used to update the rules through the rule database.

[0138] The parsing scheduler can dynamically adjust the usage priority of the two engines according to the history of task parsing, and balance the speed and success rate of task parsing.

[0139] A distributed RPA construction module is used to create a distributed RPA according to the parsed task type and initialize the execution environment; wherein the distributed RPA includes an RPA scheduler (for allocating RPA tasks) and a Kubernetes cluster (for processing RPA tasks).

[0140] There are four specific implementation methods for building distributed RPA: a) through kasmweb and DrissionPage; b) through Selenium Grid; c) through Playwright Grid; and d) through direct deployment of multiple RPA nodes through Docker containers and a simple load balancer.

[0141] The Kubernetes cluster can be replaced by Docker Swarm.

[0142] Optionally, in the distributed RPA building module, a single-machine multi-process mode can be adopted, and process management tools (such as PM2) can be used to ensure reliable service operation.

[0143] The collection process generation module is used to start the distributed RPA, execute the task through the distributed RPA, enter the target data source platform for execution of the task, perform precise search on the target data source platform based on the enterprise name, and intelligently locate and screenshot the search results to form the collection results.

[0144] The task execution status update module is used to upload the collection results to the data storage system, push the collection result notification to relevant personnel, and update the task execution status to complete the task closed loop.

[0145] The exception monitoring module is used to monitor the task execution process and automatically retry the task when an exception occurs during the task execution.

[0146] While the specific embodiments of the present invention have been described in detail above, these are merely exemplary and the present invention is not limited thereto. It will be apparent to those skilled in the art that any equivalent modifications and substitutions to the present invention fall within the scope of the present invention. Therefore, any equivalent changes and modifications made without departing from the spirit and scope of the present invention are intended to fall within the scope of the present invention.

Claims

1. A method for monitoring the collection of certificates and licenses, characterized in that: The specific steps include: Step S1: Obtain a certificate collection request task, parse the task, and form a parsed task; wherein the request includes the company name and target webpage information; Step S2: Create a distributed RPA based on the parsed task type and initialize the execution environment; wherein the distributed RPA includes an RPA scheduler and a Kubernetes cluster; Step S3: Start the distributed RPA, execute the task through the distributed RPA, and generate a collection result; Step S4: Upload the collection results to the data storage system, push the collection result notification to relevant personnel, and update the task execution status; Step S5: monitor the task execution process, and automatically retry the task if any abnormality occurs during the task execution process.

2. A method for monitoring and collecting certificates according to claim 1, characterized in that: In step S1, the task is parsed by a web page parsing engine; wherein the web page parsing engine includes a rule configuration parsing engine and a large language model parsing engine.

3. A method for monitoring and collecting certificates according to claim 2, characterized in that: Parsing the task through a web page parsing engine specifically includes the following steps: Step S101: The parsing scheduler selects a rule-configured parsing engine and parses the task using the rule-configured parsing engine. Step S102: When the rule configuration parsing engine task fails to parse, the parsing scheduler selects the LLM parsing engine and parses the task through the LLM parsing engine.

4. A method for monitoring and collecting certificates according to claim 3, characterized in that: The task is parsed by the rule configuration parsing engine, specifically including the following steps: Step S101a: select and configure a selector according to the task; Step S101b: Match the corresponding parsing rules through the rule matcher according to the URL pattern of the target web page; Step S101c: accurately locate the page element according to the selector path; Step S101d: Generate rule parsing results according to the parsing rules.

5. A method for monitoring and collecting certificates according to claim 4, characterized in that: In step S101a, the selector includes one or more of an XPath selector, a CSS selector, and a regular expression.

6. A method for monitoring and collecting certificates according to claim 3, characterized in that: When the rule configuration parsing engine fails to parse the task, the parsing scheduler selects the LLM parsing engine and parses the task through the LLM parsing engine, specifically including the following steps: Step S102a: acquiring the DOM structure and visual layout information of the target web page according to the task; Step S102b: inputting the DOM structure and visual layout information of the target webpage into the large language model; Step S102c: describing the positioning target information in natural language; Step S102d: accurately identify and extract target elements based on the visual layout information and the positioning target information to form an LLM analysis result.

7. A method for monitoring and collecting certificates according to any one of claims 3 to 6, characterized in that: The tasks successfully parsed by the LLM parsing engine are recorded, new parsing rules are formed through a rule generator, and the new parsing rules are saved to a rule database; wherein the rule database is used to save rule configurations; the rule configuration parsing engine includes a rule updater, and the rule updater and the rule matcher are used to update the rules through the rule database.

8. The method for monitoring and collecting certificates according to claim 1, characterized in that: In step S2, the distributed RPA is constructed, and its implementation method specifically includes the following: a) Implemented through kasmweb and DrissionPage; b) Implemented via Selenium Grid; c) implemented through Playwright Grid; d) Deploy multiple RPA nodes directly through Docker containers and implement them through a simple load balancer.

9. The method for monitoring and collecting certificates according to claim 1, characterized in that: In step S2, the Kubernetes cluster is replaced by Docker Swarm.

10. A certificate collection and monitoring device, characterized in that: Specifically, it includes the following modules: A task parsing module is used to obtain a certificate collection request task, parse the task, and form a parsed task; wherein the request includes the company name and target web page information; A distributed RPA construction module is used to create a distributed RPA according to the parsed task type and initialize the execution environment; wherein the distributed RPA includes an RPA scheduler and a Kubernetes cluster; A collection process generation module is used to start the distributed RPA, execute the task through the distributed RPA, enter the target data source platform for execution of the task, perform a precise search on the target data source platform based on the enterprise name, and intelligently locate and screenshot the search results to form a collection result; The task execution status update module is used to upload the collection results to the data storage system, push the collection result notification to relevant personnel, and update the task execution status to complete the task closed loop; The exception monitoring module is used to monitor the task execution process and automatically retry the task if an exception occurs during the task execution.