Cloud resource trusted data acquisition method of novel distributed Agent
The method integrates crawlers and external tools for secure, comprehensive data acquisition from cloud resources, addressing incomplete data and security issues by validating sources and encrypting transmissions.
Patent Information
- Application Number
- CN202510401604.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-15
AI Technical Summary
When the existing distributed agents acquire trusted data of cloud resources, the data acquisition is incomplete and the security is insufficient during transmission.
Ensure the comprehensiveness and security of data through crawler acquisition, external tool set authentication, data source verification, encrypted transmission and consistency verification.
It realizes comprehensive acquisition of cloud resource data, ensures real-time synchronization and security of data, and is suitable for data processing or decision-making.
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data acquisition methods, and particularly relates to a method for acquiring trustworthy data of cloud resources of a new type of distributed Agent. Background Art
[0002] In the current mainstream computer technology field, distributed Agents are playing an increasingly important role. The system decomposes tasks into multiple subtasks, which are processed in parallel by different Agent intelligent bodies, improving the processing efficiency and the scalability of the system. Trustworthy cloud resource data refers to data originating from authoritative and trustworthy sources, which is carefully organized and managed according to the intended use. It is provided in a standardized format and updated in a timely manner to support specific users or business scenarios.
[0003] In the cloud computing environment, the trustworthiness of data is particularly important. Since data may come from multiple sources and may be processed and transformed multiple times, ensuring the trustworthiness of data becomes more complex. When existing distributed Agents acquire trustworthy cloud resource data, the data acquisition method is single, there is a problem of incomplete data, and data is not encrypted during transmission, resulting in insufficient security during data transmission.
[0004] Therefore, in view of the problems of incomplete data acquisition and insufficient security during data transmission, a method for acquiring trustworthy cloud resource data of a distributed Agent can be designed. Summary of the Invention
[0005] In order to overcome the problems of incomplete data acquisition and insufficient security during data transmission.
[0006] The technical solution of the present invention is as follows: A method for acquiring trustworthy cloud resource data of a new type of distributed Agent, the steps of which are as follows:
[0007] S1: Acquire data
[0008] S11: Acquire data by crawler
[0009] S111: Determine the website or page from which data needs to be crawled;
[0010] S112: Use a crawler program to send an HTTP request to the target website to request the content of the page;
[0011] S113: Acquire the page content returned by the target website, use a network request library to send a request, and acquire and save the HTML source code of the page;
[0012] S114: Use an HTML parsing library to parse the acquired HTML source code and extract the required data;
[0013] S115: Clean and process the extracted data;
[0014] S116: Store the processed data in a database, file, or other storage medium;
[0015] S117: According to the requirements, initiate requests, obtain and parse pages in a loop until the target data is obtained or all scraping tasks are completed;
[0016] S12: External tool set
[0017] S121: Authentication to ensure that the Agent has the permission to access the API of the cloud service;
[0018] S122: After successful authentication, send an HTTP request to the corresponding API endpoint to obtain resource data;
[0019] S123: After receiving the data, parse, format, and clean it for subsequent processing or storage;
[0020] S124: Store the processed data locally or in a database for further analysis or display;
[0021] S2: Verification of data sources to ensure that the data sources are trustworthy, verified by data signatures, certificates, etc.;
[0022] S3: Encrypted transmission of data to ensure that the data is not tampered with or stolen during network transmission;
[0023] S31: Use the SSL / TLS protocol to encrypt data transmission to ensure the security and integrity of the data during transmission;
[0024] S32: Use VPN technology to establish a secure communication pipeline;
[0025] S4: Consistency check of data. For important data, perform a consistency check to ensure that the data has not been tampered with during transmission. By acquiring and releasing locks between multiple nodes, ensure data consistency. Distributed locks can be implemented through database-based locks or distributed coordination service-based locks.
[0026] Preferably, in S11, for possible anti-crawling measures, proxies, simulated logins, and setting request headers can be used for processing to ensure normal data scraping.
[0027] Preferably, in S11, an exception handling mechanism also needs to be set up to monitor the running status of the crawler, promptly discover and handle possible errors and exceptions, and ensure the stability and reliability of the crawler.
[0028] Preferably, in S11, the crawler program is run regularly according to the update frequency of the target website data to update the captured data.
[0029] Preferably, the types of network request libraries in S113 include but are not limited to urllib, requests, httpx, aiohttp, and websocket.
[0030] Preferably, the HTML parsing library in S114 is any one of HtmlAgilityPack, Jsoup, Tidy, DOMParser, and Cheerio.
[0031] Preferably, the content of data cleaning and processing in S115 includes but is not limited to removing unnecessary tags and formatting data.
[0032] Preferably, the specific steps for storing the processed data in a storage medium in S116 are as follows:
[0033] (1) Data reception and caching: When the server receives data that needs to be written to storage, it first performs data reception and stores the received data in the buffer.
[0034] (2) Data verification and processing: Before writing the data to storage, the server verifies and processes the data, including verifying the integrity, consistency, and legality of the data. If there are errors or anomalies in the data, the server will perform corresponding processing, such as repairing or discarding the data.
[0035] (3) Storage engine and file system: The server uses a storage engine and a file system to write data. The storage engine is a software component used by the server to manage data storage and provide a data access interface, while the file system is a software module used to organize and manage files and directories on the storage device.
[0036] (4) Write policy and cache management: Before writing the data to storage, the server adopts a certain write policy and cache management mechanism. Among them, the write policy determines when to write the data in the cache to storage, such as delayed writing or synchronous writing, and the cache management mechanism determines how to manage the size of the cache and data replacement.
[0037] (5) Data write persistence: The server realizes the persistent storage of data by writing the data to the storage device. During this process, the server organizes and stores the data according to the rules of the storage engine and the file system. At the same time, the server also records relevant metadata, such as file size and file permissions.
[0038] (6) Write confirmation and error handling: After the data is written to the storage, the server will send a write confirmation or feedback to the application. If an error occurs during the storage writing process, the server may perform corresponding error handling, such as retrying, rolling back, or reporting an error.
[0039] Preferably, the authentication in S121 usually involves using API keys, OAuth tokens, or other authentication mechanisms.
[0040] Advantages of the present invention: By combining the crawler and the external tool set, the acquisition of cloud resource data is realized, making the acquired data more comprehensive, achieving real-time synchronization and acquisition of massive data, and the speed of acquiring data is faster. Through the verification of data sources, encrypted transmission of data, and consistency verification of data, the security and credibility of the data are effectively ensured for data processing or decision-making. Specific embodiments
[0041] The following embodiments further illustrate the present invention.
[0042] The present invention provides an embodiment: A method for obtaining trusted cloud resource data of a new type of distributed Agent, the steps are as follows:
[0043] S1: Obtain data
[0044] S11: The crawler obtains data
[0045] S111: Determine the website or page for which data needs to be crawled;
[0046] S112: Use the crawler program to send an HTTP request to the target website to request the content of the page;
[0047] S113: Obtain the page content returned by the target website, use the network request library to send requests, obtain and save the HTML source code of the page;
[0048] S114: Use the HTML parsing library to parse the obtained HTML source code and extract the required data;
[0049] S115: Clean and process the extracted data;
[0050] S116: Store the processed data in a database, file, or other storage medium;
[0051] S117: According to the requirements, repeatedly initiate requests, obtain and parse pages until the target data is obtained or all crawling tasks are completed;
[0052] S12: External tool set
[0053] S121: Authentication to ensure that the Agent has the permission to access the API of the cloud service;
[0054] S122: After successful authentication, send an HTTP request to the corresponding API endpoint to obtain resource data;
[0055] S123: After receiving the data, parse, format, and clean it for subsequent processing or storage;
[0056] S124: Store the processed data locally or in a database for further analysis or display;
[0057] S2: Verification of the data source to ensure the credibility of the data source, verified through data signatures and certificates;
[0058] S3: Encrypted transmission of data to ensure that the data is not tampered with or stolen during transmission over the network;
[0059] S31: Use the SSL / TLS protocol to encrypt data transmission to ensure the security and integrity of the data during transmission;
[0060] S32: Use VPN technology to establish a secure communication pipeline;
[0061] S4: Consistency check of data. For important data, perform a consistency check to ensure that the data has not been tampered with during transmission. By acquiring and releasing locks between multiple nodes, ensure data consistency. Distributed locks can be implemented through database-based locks and distributed coordination service-based locks.
[0062] Preferably, in S11, for possible anti-crawling measures, proxy, simulated login, and setting request headers can be used for processing to ensure normal data scraping.
[0063] Preferably, in S11, an exception handling mechanism also needs to be set up to monitor the running status of the crawler, promptly detect and handle possible errors and exceptions, and ensure the stability and reliability of the crawler.
[0064] Preferably, in S11, according to the update frequency of the target website data, run the crawler program regularly to update the scraped data.
[0065] Preferably, the types of network request libraries in S113 include but are not limited to urllib, requests, httpx, aiohttp, and websocket.
[0066] Preferably, the HTML parsing library in S114 is any one of HtmlAgilityPack, Jsoup, Tidy, DOMParser, and Cheerio.
[0067] Preferably, the content of data cleaning and processing in S115 includes, but is not limited to, removing unnecessary tags and formatting data.
[0068] Preferably, the specific steps of storing the processed data in a storage medium in S116 are as follows:
[0069] (1) Data reception and caching: When the server receives data that needs to be written to storage, it first performs data reception and stores the received data in the buffer.
[0070] (2) Data verification and processing: Before writing the data to storage, the server verifies and processes the data, including verifying the integrity, consistency, and legality of the data. If there are errors or anomalies in the data, the server will perform corresponding processing, such as repairing or discarding the data.
[0071] (3) Storage engine and file system: The server uses the storage engine and the file system to write data. The storage engine is a software component used by the server to manage data storage and provide data access interfaces, while the file system is a software module used to organize and manage files and directories on the storage device.
[0072] (4) Write policy and cache management: Before writing the data to storage, the server adopts certain write policies and cache management mechanisms. Among them, the write policy determines when to write the data in the cache to storage, such as delayed writing or synchronous writing, and the cache management mechanism determines how to manage the size of the cache and data replacement.
[0073] (5) Data write persistence: The server realizes the persistent storage of data by writing the data to the storage device. During this process, the server organizes and stores the data according to the rules of the storage engine and the file system. At the same time, the server also records relevant metadata, such as file size and file permissions.
[0074] (6) Write confirmation and error handling: After completing the data write to storage, the server sends a write confirmation or feedback to the application. If an error occurs during the write to storage, the server may perform corresponding error handling, such as retrying, rolling back, or reporting an error.
[0075] Preferably, authentication in S121 usually involves using API keys, OAuth tokens, or other authentication mechanisms.
[0076] When performing work, the steps are as follows:
[0077] S1: Obtain data
[0078] S11: The crawler fetches data. For possible anti-crawling measures, proxies, simulated logins, and setting request headers can be used for handling to ensure normal data fetching. An exception handling mechanism also needs to be set up to monitor the running status of the crawler, promptly detect and handle possible errors and exceptions, and ensure the stability and reliability of the crawler. According to the update frequency of the data on the target website, the crawler program is run regularly to update the fetched data;
[0079] S111: Determine the website or page from which data needs to be fetched;
[0080] S112: Use the crawler program to send an HTTP request to the target website to request the content of the page;
[0081] S113: Obtain the page content returned by the target website. Use a network request library to send requests, obtain and save the HTML source code of the page. The types of network request libraries include but are not limited to urllib, requests, httpx, aiohttp, and websocket;
[0082] S114: Use an HTML parsing library to parse the obtained HTML source code and extract the required data. The HTML parsing library can be any one of HtmlAgilityPack, Jsoup, Tidy, DOMParser, and Cheerio;
[0083] S115: Clean and process the extracted data, including but not limited to removing unnecessary tags and formatting the data;
[0084] S116: Store the processed data in a database, file, or other storage media. The specific steps are as follows:
[0085] (1) Data reception and caching: When the server receives data that needs to be written to storage, it first performs data reception and stores the received data in the buffer;
[0086] (2) Data verification and processing: Before writing the data to storage, the server verifies and processes the data, including verifying the integrity, consistency, and legality of the data. If there are errors or exceptions in the data, the server will perform corresponding processing, such as repairing or discarding the data;
[0087] (3) Storage engine and file system: The server uses the storage engine and the file system to write data. The storage engine is a software component used by the server to manage data storage and provide a data access interface, while the file system is a software module used to organize and manage files and directories on the storage device;
[0088] (4) Write Strategies and Cache Management: Before writing data to storage, the server adopts certain write strategies and cache management mechanisms. Among them, the write strategy determines when to write the data in the cache to storage, such as delayed writing or synchronous writing, and the cache management mechanism determines how to manage the size of the cache area and data replacement;
[0089] (5) Data Write Persistence: The server realizes the persistent storage of data by writing data to the storage device. During this process, the server organizes and stores the data according to the rules of the storage engine and file system. At the same time, the server also records relevant metadata, such as file size and file permissions.
[0090] (6) Write Confirmation and Error Handling: After completing the data write to storage, the server sends a write confirmation or feedback to the application. If an error occurs during the write to storage, the server may perform corresponding error handling, such as retrying, rolling back, or reporting an error;
[0091] S117: According to the requirements, repeatedly initiate requests, obtain and parse pages until the target data is obtained or all scraping tasks are completed;
[0092] S12: External Tool Set
[0093] S121: Authentication to ensure that the Agent has the permission to access the API of the cloud service. Authentication usually involves using API keys, OAuth tokens, or other authentication mechanisms;
[0094] S122: After successful authentication, send an HTTP request to the corresponding API endpoint to obtain resource data;
[0095] S123: After receiving the data, parse, format, and clean it for subsequent processing or storage;
[0096] S124: Store the processed data locally or in a database for further analysis or display;
[0097] S2: Verification of Data Sources to ensure that the data sources are trustworthy, verified through data signatures and certificates;
[0098] S3: Encrypted Transmission of Data to ensure that the data is not tampered with or stolen during the network transmission process;
[0099] S31: Use the SSL / TLS protocol to encrypt data transmission to ensure the security and integrity of the data during transmission;
[0100] S32: Use VPN technology to establish a secure communication pipeline;
[0101] S4: Consistency check of data. For important data, perform consistency check to ensure that the data has not been tampered with during transmission. By acquiring and releasing locks among multiple nodes, ensure data consistency. Distributed locks can be implemented through database-based locks and locks based on distributed coordination services.
[0102] Through the above steps, combine the crawler with the external toolset to achieve the acquisition of cloud resource data, making the acquired data more comprehensive, realizing the real-time synchronization and acquisition of massive data. Through the verification of data sources, encrypted data transmission, and data consistency check, effectively ensure the security and credibility of the data, so as to perform data processing or decision-making to solve the problems of incomplete data acquisition and insufficient data security during transmission.
[0103] The above has described the embodiments of the present invention in detail, but the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the gist of the present invention.
Claims
1. A method for obtaining trusted data of cloud resources of a new type of distributed Agent, characterized in that: The steps are as follows: S1: Obtain data S11: Obtain data through web crawler S111: Determine the website or page from which data needs to be crawled; S112: Use a web crawler program to send an HTTP request to the target website to request the content of the page; S113: Obtain the page content returned by the target website, use a network request library to send requests, obtain and save the HTML source code of the page; S114: Use an HTML parsing library to parse the obtained HTML source code and extract the required data; S115: Clean and process the extracted data; S116: Store the processed data in a database, file, or other storage medium; S117: According to the requirements, initiate requests, obtain and parse pages in a loop until the target data is obtained or all crawling tasks are completed; S12: External toolset S121: Authentication to ensure that the Agent has the permission to access the API of the cloud service; S122: After successful authentication, send an HTTP request to the corresponding API endpoint to obtain resource data; S123: After receiving the data, parse, format, and clean it for subsequent processing or storage; S124: Store the processed data locally or in a database for further analysis or display; S2: Verification of data sources to ensure that the data sources are trustworthy, verified through data signatures and certificates; S3: Encrypted transmission of data to ensure that the data is not tampered with or stolen during transmission over the network; S31: Use the SSL / TLS protocol to encrypt data transmission to ensure the security and integrity of the data during transmission; S32: Use VPN technology to establish a secure communication pipeline; S4: Consistency check of data. For important data, perform a consistency check to ensure that the data has not been tampered with during transmission. By acquiring and releasing locks between multiple nodes, ensure data consistency. Distributed locks can be implemented through database-based locks and distributed coordination service-based locks.
2. The method for obtaining trusted data of cloud resources of the new distributed Agent according to claim 1, characterized in that: In S11, for possible anti-crawling measures, proxies, simulated logins, and setting request headers can be used for processing to ensure normal data crawling.
3. The method for obtaining trusted data of cloud resources of the new distributed Agent according to claim 1, characterized in that: In S11, an exception handling mechanism also needs to be set up to monitor the running status of the web crawler, promptly discover and handle possible errors and exceptions, and ensure the stability and reliability of the web crawler.
4. The method for obtaining trusted data of cloud resources of the novel distributed Agent according to claim 1, characterized in that: In S11, according to the update frequency of the target website data, run the web crawler program regularly to update the crawled data.
5. The method for obtaining trustworthy data of cloud resources of the novel distributed Agent according to claim 1, characterized in that: The types of network request libraries in S113 include but are not limited to urllib, requests, httpx, aiohttp, and websocket.
6. The method for obtaining trusted data of cloud resources of the new distributed Agent according to claim 1, characterized in that: The HTML parsing library in S114 is any one of HtmlAgilityPack, Jsoup, Tidy, DOMParser, and Cheerio.
7. The method for obtaining trusted data of cloud resources of the new distributed Agent according to claim 1, characterized in that: The content of data cleaning and processing in S115 includes but is not limited to removing unnecessary tags and formatting data.
8. The method for obtaining trusted data of cloud resources of the new distributed Agent according to claim 1, characterized in that: The specific steps for storing the processed data in a storage medium in S116 are: (1) Data reception and caching: When the server receives data to be written to storage, it first performs data reception and stores the received data in the buffer. (2) Data verification and processing: Before writing the data to storage, the server verifies and processes the data, including verifying the integrity, consistency, and legality of the data. If the data has errors or anomalies, the server will perform corresponding processing, such as repairing or discarding the data. (3) Storage engine and file system: The server uses the storage engine and file system to write data. The storage engine is a software component used by the server to manage data storage and provide a data access interface, while the file system is a software module used to organize and manage files and directories on the storage device. (4) Writing strategy and cache management: Before writing the data to storage, the server adopts certain writing strategies and cache management mechanisms. Among them, the writing strategy determines when to write the data in the cache to storage, such as delayed writing or synchronous writing, and the cache management mechanism determines how to manage the size of the cache and data replacement. (5) Data write persistence: The server achieves persistent storage of data by writing the data to the storage device. During this process, the server organizes and stores the data according to the rules of the storage engine and file system. At the same time, the server also records relevant metadata, such as file size and file permissions. (6) Write confirmation and error handling: After completing the data write to storage, the server sends a write confirmation or feedback to the application. If an error occurs during the write to storage, the server may perform corresponding error handling, such as retrying, rolling back, or reporting an error.
9. The method for obtaining trusted data of cloud resources of the novel distributed Agent according to claim 1, wherein: The authentication in S121 usually involves using API keys, OAuth tokens, or other authentication mechanisms.