A data collection task framework based on HTTP interface hot deployment and its usage method
The data collection task framework addresses the limitations of existing HTTP data collection frameworks by enabling hot deployment and flexible task management, ensuring efficient and scalable data collection even with frequent changes in data sources and processing logic.
Patent Information
- Application Number
- CN202310180637.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-02-28
AI Technical Summary
The existing HTTP data acquisition framework is difficult to handle the additions, parameter changes, return structure changes and acquisition rules during data acquisition, and cannot meet the complex logical needs of the data middle platform.
Design a data acquisition task framework based on HTTP interface hot deployment, including resource control module, task configuration analysis module, test module, task scheduler module, data request and reception module, data processing module, data persistence module, exception processing module and algorithm library support module, and realize the life cycle management and hot deployment of data acquisition tasks through the collaborative work of these modules.
It realizes rapid configuration and update of data acquisition tasks, without repackaging or running, improves system stability and scalability, and realizes flexible data acquisition task scheduling.
Smart Images

Figure CN116233101B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer networks and communications, and particularly relates to a data collection task framework based on HTTP interface hot deployment and a usage method thereof. Background Art
[0002] Using an HTTP interface for data collection is a common technical solution. By defining API interfaces, efficient and reliable data transmission can be achieved. Currently, mainstream HTTP client request frameworks such as okHTTP3, HTTPclient, netty, etc. have been able to encapsulate relatively low-level logics such as constructing HTTP request messages, managing tcp connections, and optimizing IO models well, making it simple and efficient to initiate a single HTTP request in a program. Among them, Apache HTTP Client is an open-source HTTP client library based on Java, used to send HTTP requests and receive HTTP responses. Apache HTTPClient has the following advantages: First, the API is easy to use. It provides a set of classes and methods that are easy to understand and use, and can be used to send HTTP requests such as GET, POST, PUT, DELETE, etc.; Second, high performance. Apache HTTP Client adopts technologies such as asynchronous I / O and connection pooling, which can effectively improve performance while minimizing resource consumption. It supports HTTP / 1.1 and HTTP / 2 protocols and can automatically handle mechanisms such as redirects, compression, and caching; Third, Apache HTTP Client can be customized as needed. It provides a set of flexible options to customize parameters such as request headers, proxies, Cookies, connection timeouts, etc., as well as strategies such as connection management, thread pooling, and retry; Fourth, secure and reliable. Apache HTTP Client supports HTTPS and TLS / SSL security protocols, which can encrypt and authenticate HTTP requests, thus ensuring the security and reliability of requests and responses; Fifth, Apache HTTP Client can run in a Java environment and can be used on various operating systems and development platforms, such as Windows, Linux, MacOS, Android, etc.
[0003] However, for data collectors, data collection may require large amounts of data to be collected from various data sources on a regular or periodic basis and in batches. Due to various reasons such as business expansion, HTTP data collection interfaces frequently involve additions, parameter changes, changes in return structures, and even changes in collection rules, processing (extraction, conversion, and clarification) logic of collected data, etc. For example, for the data middle platform, during data collection, it is usually necessary to complete the classification of data request APIs, deploy data request services, receive data request API results and processing, data storage, and management of data. However, various HTTP request frameworks are not enough to complete the above complex logic. Summary of the invention
[0004] To solve the above technical problems, the present invention proposes a data collection task framework and usage method based on HTTP interface hot deployment, which regards HTTP data request API definition, HTTP data request API initiation, HTTP data request API reception, simple data processing, and data storage as the life cycle of an HTTP data collection task.
[0005] The technical solution adopted by the present invention is: a data acquisition task framework based on HTTP interface hot deployment, including a resource control module, a task configuration parsing module, a testing module, a task scheduler module, a data request and receiving module, a data processing module, a data persistence module, an exception handling module, an algorithm library support module, and a collection task monitoring module.
[0006] The task configuration parsing module is connected to the test module and the collection task monitoring module respectively; the test module is connected to the task scheduler module, the algorithm library support module and the collection task monitoring module respectively; the task scheduler module is connected to the resource control module, the data request and reception module and the collection task monitoring module respectively; the data request and reception module is connected to the data processing module and the collection task monitoring module respectively; the data processing module is connected to the data persistence module, the algorithm library support module and the collection task monitoring module respectively; the data persistence module is connected to the resource control module and the collection task monitoring module respectively; the collection task monitoring module is connected to the exception handling module.
[0007] The present invention also provides a method for using a data acquisition task framework based on HTTP interface hot deployment, and the specific steps are as follows:
[0008] S1. Submit the collection task configuration Json in the form of a string to the task configuration parsing module for parsing;
[0009] S2, the task configuration parsing module performs format verification and parsing on the task configuration Json, creates a data collection task instance and initializes its properties, hands the data collection task instance to the test module, and reports the current parsing completion status to the collection task monitoring module;
[0010] S3, the test module reads the data source IP and database connection of the data collection task instance to perform a reachability test, records the test results, and determines its test status;
[0011] S4, the task scheduler module reads the Cron expression attribute of the data collection task instance that has passed the test status in step S3, registers a scheduling trigger for it, and when the scheduling condition represented by the Cron expression is triggered, applies for executable thread resources from the resource control module. When the data collection task instance obtains the thread resources, it starts executing the data request and receiving module, and reports the current waiting scheduling execution status to the collection task monitoring module;
[0012] S5. The thread pool takes out the data request and receiving module instance carrying the data collection task instance, starts executing the data request and receiving process, starts the data processing module, and reports the current data request and receiving completion status to the collection task monitoring module;
[0013] S6. The data processing module reads the data processing template in the data collection task instance, obtains the executable algorithm from the algorithm support module, assembles the processing program, and passes the data to the processing program one by one to obtain the data processing result, and starts the data persistence module to report the current data processing completion status to the collection task monitoring module;
[0014] S7. The data persistence module reads the target database connection and data storage SQL template in the data collection task instance, injects the processed data into the data storage SQL template to construct an executable SQL script, executes SQL through the target database connection, writes the data to the database, waits for the execution result to be returned, reports the current storage status to the collection task monitoring module, returns the current thread resources to the resource control module, and the collection task is completed.
[0015] Furthermore, the step S1 is specifically as follows:
[0016] The resource control parameters of the custom resource control module, including the size of each resource pool and the data acquisition task configuration storage file directory, are initialized by file reading and the data acquisition task configuration file directory is batch loaded. All .Json data acquisition task configuration files under the data acquisition task configuration directory are traversed and all are handed over to the task configuration parsing module for data acquisition task instantiation.
[0017] Further, in the step S2, the working process of the task configuration parsing module is specifically as follows:
[0018] The task configuration parsing module parses and validates the configuration Json file of the HTTP data collection task, creates an instance and initializes the data collection task, including assigning a globally unique task identifier to it, and initializing the request data source IP, request URL parameter iterator, request body iterator, scheduling policy expression, data processing template, target database connection, and storage SQL template for the task instance; after the initialization is completed, the task configuration parsing module hands the data collection task instance to the test module for testing, and at the same time reports its parsing completion status to the collection task monitoring module.
[0019] Further, the step S3 is specifically as follows:
[0020] S31. Read the data source IP attribute in the data collection task instance, construct 10 PING packets, and record the packet loss rate as R loss , read the URL parameter iterator and obtain its size Len u , read the request body iterator and obtain its size Len b ;
[0021] S32. Estimate the actual number of HTTP requests for the collection task instance. The total number of HTTP requests that need to be successfully initiated and received is L suc :
[0022] L suc = max(Len u , Len b )
[0023] Record that a total of L suc times of failed packet loss are required for successfully initiating and receiving L loss times of HTTP requests. The probability of HTTP request failure is R loss , when R loss ∈(0, 1), L loss follows a negative binomial distribution, denoted as:
[0024] L loss ~NB(L suc , R loss )
[0025] where NB represents the negative binomial distribution; then the calculation of the average number of failures for successfully initiating and receiving L suc times:
[0026]
[0027] Calculate the initial estimated resource consumption Cost of the number of HTTP requests according to the following formulareq :
[0028]
[0029] Assign Cost req to the priority attribute of the data collection task instance.
[0030] Further, in step S3, the working process of the test module is specifically as follows:
[0031] The test module receives a request task instance, reads its data source information, performs reachability tests on the data source IP and the destination database, records the test results, and reports the parsing failure status to the collection task monitoring module and discards the data collection task instance if the target is unreachable; if the reachability requirement is met, it estimates its priority value based on the size of its URL iterator, the size of the request body parameters, the packet loss rate recorded during the source IP test, and the maximum number of HTTP clients configured for the collection task instance, assigns this value to the priority attribute of the data collection task instance, then hands it over to the task scheduler module, and at the same time reports its test passed status to the collection task monitoring module.
[0032] Further, in step S4, the working process of the task scheduler module is specifically as follows:
[0033] The task scheduler module reads the Cron expression attribute in the data collection task instance, registers a scheduling trigger corresponding to the Cron expression condition. When the Cron expression meets the system time and period conditions, the scheduling trigger will be triggered. It reads the size of the collection task HTTP request window (the number of concurrent HTTP requests during the execution of the collection task), applies to the resource control module for the corresponding number of HTTP request clients, and at the same time applies to the thread pool of the resource control module for executable threads. It places the data collection task instance with the priority attribute into the priority queue of the thread pool to wait for obtaining thread resources to start execution. When the data collection task instance obtains thread resources, it hands over the collection task instance and the HTTP request client set to the data request and reception module, and reports its queuing waiting for scheduling execution status to the collection task monitoring module.
[0034] Further, in step S5, in the data request and reception module, the request and reception process is specifically as follows:
[0035] The data request and reception module first reads the HTTP request API configuration in the data acquisition task instance, including the data source IP, port, URL parameter iterator, and request body parameter iterator. Based on the Apache HTTP Client framework, it constructs all the required HTTP request instances in one go according to the iteration order, opens the request window, and consumes the request clients in the HTTP request client pool from left to right according to the sliding window strategy. It asynchronously initiates the HTTP request instances through the clients. When the request clients are exhausted and the leftmost request instance has not obtained the result, it enters the blocked state until the leftmost request instance obtains the request result and releases the request client, and the window slides to the right; it loops until all request instances obtain the request results. First, it returns all the HTTP request clients to the resource control module, recalculates the data acquisition task priority according to the average packet loss rate and the request client window size in all requests, and assigns it to the data acquisition task instance priority field; secondly, it integrates all the request results in order, hands over the data acquisition task instance and the integrated data to the data processing module, continues to execute through the current thread resources, and reports the status of request and reception completion to the acquisition task monitoring module at the same time.
[0036] Further, in step S6, the working process of the data processing module is specifically as follows:
[0037] The data processing module reads the data processing template in the data acquisition task instance, assembles the processing process of a single piece of data. When the processing expression involves querying the algorithm library support module, it obtains the executable algorithm function body provided by the algorithm library support module, including field extraction, field operation, and renaming in a single piece of data. It integrates the processed data in order to obtain the processed data, hands it over to the data persistence module, and reports the status of the data acquisition task processing completion to the acquisition task monitoring module at the same time.
[0038] Further, the method of the present invention further includes step S8, specifically as follows:
[0039] All sub-modules in the method of the present invention report exceptions to the acquisition task monitoring module when encountering abnormal situations. The acquisition task monitoring module hands over the specific exception information to the exception handling module and terminates the current scheduling of the acquisition task instance. The exception handling module outputs the exception information to the system or the log file.
[0040] Advantages of the present invention: The framework of the present invention includes a resource control module, a task configuration parsing module, a test module, a task scheduler module, a data request and reception module, a data processing module, a data persistence module, an exception handling module, an algorithm library support module, and a collection task monitoring module. The present invention can not only simply and quickly configure a new or modify an old HTTP interface data collection task, but also quickly update it to the running environment in the form of hot deployment without repackaging or running the project, improving system stability and scalability, realizing hot deployment of the HTTP data collection interface, and thus realizing more flexible data collection task scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 FIG. is a structural diagram of a data collection task framework based on hot deployment of an HTTP interface according to the present invention.
[0042] Figure 2 FIG. is a flowchart of a method for using a data collection task framework based on hot deployment of an HTTP interface according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The method of the present invention will be further described below with reference to the drawings and embodiments.
[0044] As Figure 1 shown, a data collection task framework based on hot deployment of an HTTP interface includes a resource control module, a task configuration parsing module, a test module, a task scheduler module, a data request and reception module, a data processing module, a data persistence module, an exception handling module, an algorithm library support module, and a collection task monitoring module.
[0045] The task configuration parsing module is respectively connected to the test module and the collection task monitoring module; the test module is respectively connected to the task scheduler module, the algorithm library support module, and the collection task monitoring module; the task scheduler module is respectively connected to the resource control module, the data request and reception module, and the collection task monitoring module; the data request and reception module is respectively connected to the data processing module and the collection task monitoring module; the data processing module is respectively connected to the data persistence module, the algorithm library support module, and the collection task monitoring module; the data persistence module is respectively connected to the resource control module and the collection task monitoring module; the collection task monitoring module is connected to the exception handling module.
[0046] As Figure 2 shown, in this embodiment, the present invention also provides a method for using a data collection task framework based on hot deployment of an HTTP interface, and the specific steps are as follows:
[0047] S1. Hand over the collection task configuration Json in the form of a string to the task configuration parsing module for parsing;
[0048] S2, the task configuration parsing module performs format verification and parsing on the task configuration Json, creates a data collection task instance and initializes its properties, hands the data collection task instance to the test module, and reports the current parsing completion status to the collection task monitoring module;
[0049] S3, the test module reads the data source IP and database connection of the data collection task instance to perform a reachability test, records the test results, and determines its test status;
[0050] S4, the task scheduler module reads the Cron expression attribute of the data collection task instance that has passed the test status in step S3, registers a scheduling trigger for it, and when the scheduling condition represented by the Cron expression is triggered, applies for executable thread resources from the resource control module. When the data collection task instance obtains the thread resources, it starts executing the data request and receiving module, and reports the current waiting scheduling execution status to the collection task monitoring module;
[0051] S5. The thread pool takes out the data request and receiving module instance carrying the data collection task instance, starts executing the data request and receiving process, starts the data processing module, and reports the current data request and receiving completion status to the collection task monitoring module;
[0052] S6. The data processing module reads the data processing template in the data collection task instance, obtains the executable algorithm from the algorithm support module, assembles the processing program, and passes the data to the processing program one by one to obtain the data processing result, and starts the data persistence module to report the current data processing completion status to the collection task monitoring module;
[0053] S7. The data persistence module reads the target database connection and data storage SQL template in the data collection task instance, injects the processed data into the data storage SQL template to construct an executable SQL script, executes SQL through the target database connection, writes the data to the database, waits for the execution result to be returned, reports the current storage status to the collection task monitoring module, returns the current thread resources to the resource control module, and the collection task is completed.
[0054] In this embodiment, in step S1, the resource control parameters of the customized resource control module, including the size of each resource pool and the data acquisition task configuration storage file directory, are initialized by file reading and the data acquisition task configuration file directory is loaded in batches. All .Json data acquisition task configuration files under the data acquisition task configuration directory are traversed and all are handed over to the task configuration parsing module for data acquisition task instantiation, as follows:
[0055] S11. Read the resource control parameter configuration in the resource control configuration file, including the core thread number REQ_TASK_CORE_THREAD_POOL_SIZE of the request task thread pool (assuming the number of CPU cores of the host where the framework runs is N, then by default, the core thread number is N - 1), and the maximum thread number REQ_TASK_MAX_THREAD_POOL_SIZE of the request task thread pool (the value should not be less than the core thread number, and the default value is 2N), and create an executable thread pool for the collection task;
[0056] S12. Read the HTTP request client pool size number HTTP_MAX_CLIENT_SIZE (the default value is 10 * REQ_TASK_MAX_THREAD_POOL_SIZE), create an HTTP request client pool with a size of HTTP_MAX_CLIENT_SIZE based on the Apache open-source framework HTTP Client, the maximum number of collection tasks TASK_MAX_SIZE (by default, it takes the maximum signed 4-byte integer 2147483647), and the absolute path of the directory where the data request task description file is stored REQ_TASK_DESC_DIR; Read all.Json files in the REQ_TASK_DESC_DIR directory, convert them into an array of Json objects, and pass them to the task configuration parsing module.
[0057] In this embodiment, in step S2, the task configuration parsing module parses and validates the configuration Json file of the HTTP data collection task, creates an instance and initializes the data collection task, including assigning a globally unique task identifier to it, and initializing the request data source IP, request URL parameter iterator, request body iterator, scheduling policy expression, data processing template, target database connection, and storage SQL template for this task instance; After the initialization is completed, the task configuration parsing module hands the data collection task instance to the test module for testing, and at the same time reports its parsing completion status to the collection task monitoring module. The specific working process of the task configuration parsing module is as follows:
[0058] S21. Parse the scheduling task configuration, including the task name, scheduling policy expression, and the maximum number of HTTP request clients, and assign them to the attributes of the collection task instance; Among them, the scheduling policy Cron expression TASK_SCHE_CRON, the value is of string type, and its format is:
[0059] "seconds minutes hours dayofmonth month dayofweek[year]", following the Linux cron expression specification, used to express the periodic scheduling policy of the request task.
[0060] S22. Analyze the HTTP request API configuration, including the data source IP, port, URL parameter iteration configuration, and request body iteration configuration, create instances of the URL parameter iterator and the request body iterator, and assign them to the data collection task instance attributes;
[0061] S23. Analyze the data processing configuration, including the specific fields to be processed and the infix expression for processing, construct a data processing template, and assign it to the data collection task instance attributes;
[0062] S24. Analyze the data storage configuration, including the target database IP, port, database type, and database name, construct a database connection instance, analyze the table name and storage method, construct a data SQL DML script template, and assign it to the data collection task instance attributes.
[0063] In this embodiment, the specific steps of step S3 are as follows:
[0064] S31. Read the data source IP attribute in the data collection task instance, construct 10 PING packets, and record the packet loss rate as R loss , read the URL parameter iterator, and obtain its size Len u , read the request body iterator, and obtain its size Len b ;
[0065] S32. Estimate the actual number of HTTP requests for the collection task instance. The total number of HTTP requests that need to be successfully initiated and received is L suc :
[0066] L suc = max(Len u , Len b )
[0067] Record that a total of L suc times of failed packet loss are required for successfully initiating and receiving L loss times of HTTP requests. The probability of an HTTP request failure is R loss , when R loss ∈(0, 1), L loss follows a negative binomial distribution, denoted as:
[0068] L loss ~NB(L suc , R loss )
[0069] Among them, NB represents the negative binomial distribution; then the calculation of the average number of failures for successfully initiating and receiving L suc times is as follows:
[0070]
[0071] Calculate the resource consumption Cost of the initial estimated number of HTTP requests according to the following formula req :
[0072]
[0073] Assign Cost req to the priority attribute of the data collection task instance.
[0074] In this embodiment, in the step S3, the working process of the test module is specifically as follows:
[0075] The test module receives a request task instance, reads its data source information, performs reachability tests on the data source IP and the destination database, records the test results, and reports the parsing failure status to the collection task monitoring module and discards the data collection task instance if the target is unreachable; if the reachability requirement is met, estimate its priority value based on the size of its URL iterator, the size of the request body parameters, the packet loss rate recorded during the source IP test, and the maximum number of HTTP clients configured for the collection task instance, and assign this value to the priority attribute of the data collection task instance, and then hand it over to the task scheduler module, and at the same time report its test pass status to the collection task monitoring module.
[0076] In this embodiment, in the step S4, the working process of the task scheduler module is specifically as follows:
[0077] The task scheduler module reads the Cron expression attribute in the data collection task instance, registers a scheduling trigger corresponding to the Cron expression condition, and when the Cron expression meets the system time and cycle conditions, the scheduling trigger will be triggered, reads the size of the collection task HTTP request window (the number of concurrent HTTP requests during the execution of the collection task), applies to the resource control module for the corresponding number of HTTP request clients, and at the same time applies to the thread pool of the resource control module for executable threads, puts the data collection task instance with the priority attribute into the priority queue of the thread pool to wait for obtaining thread resources to start execution. When the data collection task instance obtains thread resources, hand over the collection task instance and the HTTP request client set to the data request and receiving module, and report its queuing waiting for scheduling execution status to the collection task monitoring module.
[0078] In this embodiment, in the step S5, in the data request and receiving module, the request and receiving process is specifically as follows:
[0079] The data request and reception module first reads the HTTP request API configuration in the data acquisition task instance, including the data source IP, port, URL parameter iterator, and request body parameter iterator. Based on the Apache HTTP Client framework, it constructs all the required HTTP request instances in one go according to the iteration order, opens the request window, and consumes the request clients in the HTTP request client pool from left to right according to the sliding window strategy. It asynchronously initiates the HTTP request instances through the clients. When the request clients are exhausted and the leftmost request instance has not obtained the result, it enters the blocked state until the leftmost request instance obtains the request result and releases the request client, and the window slides to the right; it loops until all request instances obtain the request results. First, it returns all the HTTP request clients to the resource control module, recalculates the data acquisition task priority according to the average packet loss rate and the request client window size in all requests, and assigns it to the data acquisition task instance priority field; secondly, it integrates all the request results in order, hands over the data acquisition task instance and the integrated data to the data processing module, continues to execute through the current thread resources, and at the same time reports the status of request and reception completion to the acquisition task monitoring module.
[0080] In this embodiment, step S5 is specifically as follows:
[0081] S51. Read the data source IP, port, URL parameter iterator, and request body iterator of the data acquisition task instance, and loop to put the results of the URL parameter iterator and the results iterated by the request body iterator into the HTTP request instance to construct a series of HTTP request instances numbered 0 to (L suc -1), define the initial value of the left pointer LEFT as 0, and the initial value of the right pointer RIGHT as 0;
[0082] S52. Read the maximum number of HTTP request clients of the data acquisition task instance as the window size, denoted as MAX_CLIENT. Initialize the available pool and the in-use pool of HTTP request clients, with both initial sizes being 0;
[0083] S53. Loop to determine that when the condition RIGHT - LEFT < MAX_CLIENT is satisfied, send the HTTP request instance pointed to by the right pointer RIGHT based on the HTTP request client, and increment the right pointer RIGHT by 1;
[0084] S54. When RIGHT - LEFT = MAX_CLIENT, start to determine whether the HTTP request instance pointed to by the left pointer LEFT has received the request result. If it has received the request result, increment the left pointer LEFT by 1; otherwise, enter the blocked state until the HTTP request instance pointed to by the left pointer LEFT has received the request result.
[0085] In this embodiment, in step S6, the working process of the data processing module is specifically as follows:
[0086] The data processing module reads the data processing template in the data collection task instance, assembles the processing flow of a single piece of data. When the processing expression involves querying the algorithm library support module, it obtains the executable algorithm function body provided by the algorithm library support module, including field extraction, field operation, and renaming in a single piece of data, integrates the processed data in order to obtain the processed data, hands it over to the data persistence module, and at the same time reports the status that the data collection task has been processed to the collection task monitoring module.
[0087] In this embodiment, step S6 is specifically as follows:
[0088] S61. Parse the data processing template string in the data collection task instance, strip the operands (including field names and constant values) and operators (basic operators and algorithm names) in the infix expression, and push them onto the stack respectively;
[0089] S62. Pop the operator and its corresponding operand in sequence, and the judgment process is as follows:
[0090] S621. If the operator is a basic operator, directly perform the corresponding basic operation on the operand, and push the operation result onto the stack as the operand of the next operator;
[0091] S631. If the operator is an algorithm name, apply to the algorithm library support module for the algorithm function bytecode (executable method instance) corresponding to the algorithm name, use the operand as the function input parameter, call the algorithm function, and push the call result onto the stack as the operand of the next operator;
[0092] S632. Loop the operation in step S62 until the operator is consumed. Then the final operand is the operation result, which is assigned to the corresponding field.
[0093] In this embodiment, the method of the present invention further includes step S8, which is specifically as follows:
[0094] All sub-modules in the method of the present invention report exceptions to the collection task monitoring module when an abnormal situation occurs. The collection task monitoring module hands the specific exception information to the exception handling module and terminates the current scheduling of the collection task instance. The exception handling module outputs the exception information to the system or the log file.
[0095] In this embodiment, the resource control module is used for centralized management of resources, including thread pool resources, HTTP request resources, etc.
[0096] The algorithm library support module is used to define and manage the compiled and executable algorithm programs, and open them to the data processing module to provide algorithm support during the data processing. The algorithm library support module runs independently, uses the algorithm name as the unique identifier, and provides the test module and the data processing module with a list of available algorithms and the function bodies corresponding to the algorithms that can be actually executed.
[0097] The acquisition task monitoring module is used to monitor the real-time status of the data acquisition tasks. At the same time, it receives requests for adding, modifying, and deleting data acquisition tasks from the outside, and cooperates with the task configuration parsing module and the task scheduler module to achieve hot deployment of changes to the data acquisition tasks. The acquisition task monitoring module runs independently, serves as a message center to receive message reports from each module, and provides them to the framework user or records them in the form of logs when needed. For abnormal information, the specific content will be handed over to the exception handling module for processing.
[0098] The exception handling module is used to receive the status and error information when the data acquisition task is abnormal, and parse and output the abnormal situation. The exception handling module runs independently. When the acquisition task monitoring module receives the abnormal information, it hands over the specific abnormal status, abnormal information, and abnormal acquisition task instance to the exception handling module. This module reports the exception to the framework user when needed.
[0099] In summary, the present invention provides a data acquisition task framework and usage method based on HTTP interface hot deployment. This framework better solves the problems existing in the above data acquisition scenarios from two aspects. On the one hand, it adopts a mechanism of requesting a task description file based on the HTTP interface, and parses it based on the defined file format (request URL, request parameters, data reception, simple data processing, data storage, etc. and data acquisition task strategy configuration) and the HTTP interface request task parser at runtime to dynamically and flexibly create HTTP request task instances, and updates the request task instances in a timely manner when the request task description file changes, realizing the hot deployment of the HTTP data acquisition interface, without having to repackage and recompile the project source code for running, improving the system stability and scalability.
[0100] On the other hand, it provides a usage method for the data acquisition task framework based on HTTP interface hot deployment, uses a dynamic scheduling strategy for HTTP request task instances based on priority and an HTTP request task factory of the resource pool to manage and schedule each data acquisition task instance and allocate resources, and dynamically updates the priority of the task instances according to factors such as user settings, instance request feedback, bandwidth, throughput, etc., so as to achieve more flexible data acquisition task scheduling.
[0101] Those of ordinary skill in the art will realize that the embodiments described herein are to assist the reader in understanding the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.
Claims
1. A method for using a data collection task framework based on HTTP interface hot deployment, the specific steps are as follows: S1. Submit the collection task configuration Json in the form of a string to the task configuration parsing module for parsing; S2, the task configuration parsing module performs format verification and parsing on the task configuration Json, creates a data collection task instance and initializes its properties, hands the data collection task instance to the test module, and reports the current parsing completion status to the collection task monitoring module; S3, the test module reads the data source IP and database connection of the data collection task instance to perform a reachability test, records the test results, and determines its test status; In step S3, the test module workflow is as follows: The test module receives the request task instance, reads its data source information, performs reachability tests on the data source IP and the destination database, records the test results, and reports the parsing failure status to the collection task monitoring module if the target is unreachable, and discards the data collection task instance; If the reachability requirement is met, the priority value is estimated based on the size of its URL iterator, the size of the request body parameters, the packet loss rate recorded during the source IP test, and the maximum number of HTTP clients of the collection task instance configuration instance. The value is assigned to the priority attribute of the data collection task instance, and then it is handed over to the task scheduler module, and the test pass status is reported to the collection task monitoring module at the same time; S4, the task scheduler module reads the Cron expression attribute of the data collection task instance that has passed the test status in step S3, registers a scheduling trigger for it, and when the scheduling condition represented by the Cron expression is triggered, applies for executable thread resources from the resource control module. When the data collection task instance obtains the thread resources, it starts executing the data request and receiving module, and reports the current waiting scheduling execution status to the collection task monitoring module; S5. The thread pool takes out the data request and receiving module instance carrying the data collection task instance, starts executing the data request and receiving process, starts the data processing module, and reports the current data request and receiving completion status to the collection task monitoring module; S6. The data processing module reads the data processing template in the data collection task instance, obtains the executable algorithm from the algorithm support module, assembles the processing program, and passes the data to the processing program one by one to obtain the data processing result, and starts the data persistence module to report the current data processing completion status to the collection task monitoring module; S7. The data persistence module reads the target database connection and data storage SQL template in the data collection task instance, injects the processed data into the data storage SQL template to construct an executable SQL script, executes SQL through the target database connection, writes the data to the database, waits for the execution result to be returned, reports the current storage status to the collection task monitoring module, returns the current thread resources to the resource control module, and the collection task is completed.
2. The method for using a data collection task framework based on HTTP interface hot deployment according to claim 1, characterized in that The step S1 is specifically as follows: The resource control parameters of the custom resource control module, including the size of each resource pool and the data acquisition task configuration storage file directory, are initialized by file reading and the data acquisition task configuration file directory is batch loaded. All .Json data acquisition task configuration files under the data acquisition task configuration directory are traversed and all are handed over to the task configuration parsing module for data acquisition task instantiation.
3. A method for using a data collection task framework based on HTTP interface hot deployment according to claim 1, characterized in that In step S2, the working process of the task configuration parsing module is as follows: The task configuration parsing module parses and validates the configuration Json file of the HTTP data collection task, creates an instance and initializes the data collection task, including assigning a globally unique task identifier to it, and initializing the request data source IP, request URL parameter iterator, request body iterator, scheduling policy expression, data processing template, target database connection, and storage SQL template for the task instance; After initialization, the task configuration parsing module hands the data collection task instance to the test module for testing, and at the same time reports its parsing completion status to the collection task monitoring module.
4. The method for using a data acquisition task framework based on HTTP interface hot deployment according to claim 1, wherein Step S3 is specifically as follows: S31. Read the data source IP attribute in the data collection task instance, construct 10 PING packets, and record the packet loss rate as R loss , read the URL parameter iterator and obtain its size Len u , read the request body iterator and obtain its size Len b ; S32. Estimate the actual number of HTTP requests for the data collection task instance, which is the total number of HTTP requests L that need to be successfully initiated and received suc : L suc = max(Len u , Len b ) Record successfully initiating and receiving L suc times of HTTP requests in total require L loss times of failed packet loss. The probability of an HTTP request failing is R loss , when R loss ∈(0,1), L loss follows a negative binomial distribution, denoted as: L loss ~NB(L suc ,R loss ) Among them, NB represents the negative binomial distribution; then the successful initiation and reception of L suc Calculation of the average number of failures in L times: Calculate the resource consumption Cost of the initial estimated number of HTTP requests according to the following formula req : Assign Cost req to the priority attribute of the data collection task instance.
5. The method for using a data acquisition task framework based on HTTP interface hot deployment according to claim 1, characterized in that In step S4, the working process of the task scheduler module is as follows: The task scheduler module reads the Cron expression attribute in the data collection task instance, registers the scheduling trigger corresponding to the Cron expression condition. When the Cron expression meets the system time and cycle conditions, the scheduling trigger will be triggered. It reads the HTTP request window size of the collection task, applies to the resource control module for the corresponding number of HTTP request clients, and at the same time applies to the thread pool of the resource control module for executable threads. It puts the data collection task instance with the priority attribute into the priority queue of the thread pool to wait for obtaining thread resources to start execution. When the data collection task instance obtains thread resources, it hands the collection task instance and the set of HTTP request clients to the data request and reception module, and reports its queuing waiting for scheduling execution status to the collection task monitoring module.
6. The method for using a data collection task framework based on HTTP interface hot deployment according to claim 1, characterized in that In step S5, in the data request and reception module, the request and reception process is as follows: The data request and reception module first reads the HTTP request API configuration in the data collection task instance, including the data source IP, port, URL parameter iterator, and request body parameter iterator. Based on the Apache HTTP Client framework, it constructs all the required HTTP request instances in one go according to the iteration order, opens the request window, and consumes the request clients in the HTTP request client pool from left to right according to the sliding window strategy. It asynchronously initiates the HTTP request instances through the client. When the request clients are exhausted and the leftmost request instance has not obtained the result, it enters the blocked state until the leftmost request instance obtains the request result and releases the request client, and the window slides to the right; it loops until all request instances obtain the request results. First, it returns all the HTTP request clients to the resource control module, recalculates the data collection task priority according to the average packet loss rate and the request client window size in all requests, and assigns it to the priority field of the data collection task instance; Secondly, it integrates all the request results in order, hands the data collection task instance and the integrated data to the data processing module, continues to execute through the current thread resources, and at the same time reports the request and reception completion status to the collection task monitoring module.
7. A method for using a data collection task framework based on HTTP interface hot deployment according to claim 1, characterized in that, In step S6, the working process of the data processing module is as follows: The data processing module reads the data processing template in the data acquisition task instance, assembles the processing flow for a single piece of data. When the processing expression involves querying the algorithm library support module, it obtains the executable algorithm function body provided by the algorithm library support module, including field extraction, field operation, and renaming in a single piece of data. The processed data is integrated in sequence to obtain the processed data, which is then handed over to the data persistence module. At the same time, it reports the completion status of the data acquisition task to the acquisition task monitoring module.
8. A method for using a data collection task framework based on HTTP interface hot deployment according to claim 1, characterized in that The method further includes step S8, which is specifically as follows: When encountering an abnormal situation, all sub-modules report the abnormality to the acquisition task monitoring module. The acquisition task monitoring module hands the specific abnormality information to the abnormality handling module and terminates the current scheduling of the acquisition task instance. The abnormality handling module outputs the abnormality information to the system or the log file.
9. A data acquisition task framework based on HTTP interface hot deployment, using the method of a data acquisition task framework based on HTTP interface hot deployment described in any one of claims 1-8, including a resource control module, a task configuration parsing module, a test module, a task scheduler module, a data request and reception module, a data processing module, a data persistence module, an abnormality handling module, an algorithm library support module, and an acquisition task monitoring module; The task configuration parsing module is respectively connected to the test module and the acquisition task monitoring module; the test module is respectively connected to the task scheduler module, the algorithm library support module, and the acquisition task monitoring module; the task scheduler module is respectively connected to the resource control module, the data request and reception module, and the acquisition task monitoring module; the data request and reception module is respectively connected to the data processing module and the acquisition task monitoring module; the data processing module is respectively connected to the data persistence module, the algorithm library support module, and the acquisition task monitoring module; the data persistence module is respectively connected to the resource control module and the acquisition task monitoring module; The acquisition task monitoring module is connected to the abnormality handling module.