Data acquisition and pushing system and method, electronic equipment and storage medium

Through the display layer, scheduling layer and execution layer of the data collection and push system, unified management and automated data collection of multiple databases are achieved, solving the problems of complex user operations and low efficiency in existing technologies, and improving data integration efficiency and user experience.

CN120632208APending Publication Date: 2025-09-12太保科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510725311.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies are unable to perform batch data collection for different types of databases, resulting in complex and inefficient user operations.

Method used

A data collection and push system is provided, including a display layer, a scheduling layer, a splitting layer, and an execution layer. By customizing task configuration parameters and data query scripts, it is compatible with multiple database types, automatically executes data query operations, and pushes data.

Benefits of technology

It realizes unified management and batch data collection of different types of databases, improving data integration efficiency and user operation experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632208A_ABST
    Figure CN120632208A_ABST
Patent Text Reader

Abstract

The invention provides a data acquisition and pushing system and method, electronic equipment and a storage medium. The system comprises a display layer, a scheduling layer, a splitting layer, an execution layer and a storage layer, the display layer is used for acquiring connection information and task configuration files of multiple databases; the scheduling layer is used for scheduling a current collection task related to a plurality of databases to be collected according to the execution period; the splitting layer is used for splitting the current collection task into a plurality of sub-tasks in one-to-one correspondence with a plurality of databases to be collected, and pushing the plurality of sub-tasks to the message queue; the execution layer is used for remotely connecting a plurality of to-be-collected databases according to the connection information, traversing and executing each target data query script associated with a subtask corresponding to each to-be-collected database to obtain current collected data, and pushing the current collected data to a target end in response to the situation that the current collected data meets a batch pushing condition in the execution process; the storage layer is used for recording a corresponding execution result in response to the completion of execution of the current collection task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application mainly relates to the field of data processing technology, and in particular to a data collection and push system, method, electronic device and storage medium. Background Art

[0002] In modern enterprise information systems, data often comes from multiple types of databases (for example, relational databases and non-relational databases). Currently, different scheduling tools are required for different types of databases. However, due to the wide variety of databases, users need to frequently switch scheduling tools to collect data from each type of database. Batch data collection for different types of databases is impossible, which makes user operations complicated and inefficient. Summary of the Invention

[0003] The purpose of this application is to provide a data collection and push system, method, electronic device and storage medium, which can perform batch collection, processing and push for different types of databases, thereby improving data integration efficiency and user operation experience.

[0004] In a first aspect, a data collection and push system is provided, comprising: a presentation layer, a scheduling layer, a splitting layer, an execution layer, and a storage layer; wherein,

[0005] The presentation layer is used to obtain connection information and task configuration files for multiple databases. The connection information includes the connection addresses, ports, and connection permissions of the multiple databases. The task configuration files include task configuration parameters and multiple data query scripts adapted to the database types of the multiple databases. The task configuration parameters include an execution cycle, a task type, and batch push conditions. The data query script is used to automatically execute data query operations for the corresponding database types based on pre-configured session parameters and the target end.

[0006] The scheduling layer is used to schedule current collection tasks involving several databases to be collected according to the execution cycle;

[0007] The splitting layer is used to split the current acquisition task into a plurality of subtasks corresponding one-to-one to the plurality of databases to be acquired, and push the plurality of subtasks to a message queue;

[0008] The execution layer is used to remotely connect to the plurality of databases to be collected according to the connection information, traverse and execute each target data query script associated with the subtask corresponding to each database to be collected, obtain the current collected data, and during the execution process, in response to the current collected data meeting the batch push condition, push the current collected data to the target end;

[0009] The storage layer is used to record corresponding execution results in response to the completion of the execution of the current acquisition task, and the execution results include task execution status, execution time and abnormal events.

[0010] In some embodiments, the presentation layer includes a database connection configuration module, a script repository management module, a collection task configuration module, and an approval management module; wherein,

[0011] The database connection configuration module provides a configuration interface and a batch import and export interface for obtaining the connection information through the configuration interface and the batch import and export interface, and supports the validity verification and connection test functions of the connection information;

[0012] The script repository management module is used to create the multi-data query script and support version control, retrieval, editing, deletion, preview and history recording functions of the multi-data query script;

[0013] The acquisition task configuration module is used to obtain the task configuration parameters, generate the task configuration file according to the task configuration parameters, the multi-data query script and the preset task template, and support the task priority setting function and the task dependency configuration function;

[0014] The approval management module is used to create an approval process for the multi-data query script and the task configuration file, and record the corresponding approval history when the approval process is completed.

[0015] In some embodiments, the splitting layer includes: an instance screening module, a script screening module, a task splitting module, a task identification module and a push message queue module.

[0016] The instance screening module is used to screen out the plurality of databases to be collected corresponding to the current collection task from the multiple databases;

[0017] The script screening module is used to screen out the target data query script corresponding to the current acquisition task from the multiple data query scripts;

[0018] The task splitting module is used to split the current acquisition task into a plurality of subtasks corresponding one-to-one to the plurality of databases to be acquired;

[0019] The task identification module is used to identify the current task type corresponding to the current acquisition task;

[0020] The push message queue module is used to push the multiple subtasks to the message queue corresponding to the current task type.

[0021] In some embodiments, the execution layer includes: a database remote connection module, a script execution module, a cursor query module, a data processing module, a resource dynamic expansion module and a data push module; wherein,

[0022] The database remote connection module is used to remotely connect to the plurality of databases to be collected according to the connection information;

[0023] The script execution module is used to traverse and execute each target data query script associated with the subtask corresponding to each database to be collected, and during the execution process, capture the abnormal event according to the preset task timeout time and the execution timeout time of a single script;

[0024] The cursor query module is used to perform cursor reading on the execution result of a single target data query script during the execution process;

[0025] The data processing module is used to perform format processing on the currently collected data;

[0026] The resource dynamic expansion module is used to dynamically adjust the amount of software and hardware resources required for the current acquisition task;

[0027] The data push module is used to push the current collected data to the target end in response to the current collected data meeting the batch push condition during the execution process. The batch push condition includes: the data volume of the current collected data exceeds a preset batch processing data volume threshold.

[0028] In some embodiments, the target end includes a structured database and a stream processing platform.

[0029] In a second aspect, a data collection and push method is provided. The method is applied to any data collection and push system described in the first aspect, comprising:

[0030] Obtaining connection information and task configuration files for multiple databases, wherein the connection information includes the connection addresses, ports, and connection permissions of the multiple databases; the task configuration files include task configuration parameters and multiple data query scripts adapted to the database types of the multiple databases; the task configuration parameters include an execution cycle, a task type, and batch push conditions; and the data query script is used to automatically execute data query operations for the corresponding database types based on pre-configured session parameters and the target end;

[0031] Scheduling current collection tasks involving a plurality of databases to be collected according to the execution cycle;

[0032] Splitting the current acquisition task into a plurality of subtasks corresponding one-to-one to the plurality of databases to be acquired, and pushing the plurality of subtasks to a message queue corresponding to the current task type;

[0033] Accessing the plurality of databases to be collected through a remote database connection, traversing and executing each target data query script associated with a subtask corresponding to each database to be collected, obtaining current collected data, and in response to the current collected data satisfying the push condition, pushing the current collected data to the target end;

[0034] In response to the completion of the execution of the current acquisition task, a corresponding execution result is recorded, where the execution result includes the task execution status, execution time, and abnormal events.

[0035] In some embodiments, splitting the current acquisition task into a plurality of subtasks corresponding one-to-one to the plurality of databases to be acquired, and pushing the plurality of subtasks to a message queue includes:

[0036] Filtering the plurality of databases to be collected corresponding to the current collection task from the multiple databases;

[0037] Filtering the target data query script corresponding to the current acquisition task from the multiple data query scripts;

[0038] Splitting the current acquisition task into a plurality of subtasks corresponding one-to-one to the plurality of databases to be acquired;

[0039] Identifying a current task type corresponding to the current acquisition task; and

[0040] Push the multiple subtasks to the message queue corresponding to the current task type.

[0041] In some embodiments, traversing and executing each target data query script associated with a subtask corresponding to each database to be collected to obtain current collected data, and in response to satisfying the push condition, pushing the current collected data to the target end, includes:

[0042] Traverse and execute each target data query script associated with the subtask corresponding to each database to be collected. During the execution process, perform the following steps:

[0043] Capturing the abnormal event according to the preset task timeout and single script execution timeout;

[0044] Use a cursor to read the execution results of a single target data query script;

[0045] In response to the current collected data meeting the push condition, the current collected data is pushed to the target end, and the batch push condition includes: the data volume of the current collected data exceeds a preset batch processing data volume threshold.

[0046] In a third aspect, an electronic device is provided. The electronic device includes: one or more processors; and one or more memories coupled to the one or more processors and storing instructions thereon. When the instructions are executed individually or collectively by the one or more processors, the electronic device executes the aforementioned data collection and push method.

[0047] In a fourth aspect, a non-transitory computer-readable storage medium storing machine-executable instructions is provided. When the machine-executable instructions are executed by one or more processors of a machine, the machine is caused to perform any one of the above-mentioned data collection and push methods.

[0048] Compared with the prior art, this application has the following advantages:

[0049] The data collection and push system, method, electronic device and storage medium provided by the present application support user-defined configuration of task configuration parameters and data query script tasks to generate task configuration files, and are compatible with multiple types of databases. They can uniformly configure and manage the connection information and task configuration files of multiple databases in the presentation layer, helping users to centrally manage various collection tasks required by actual business and improve the flexibility of collection tasks. Furthermore, according to the execution cycle, the current collection task involving several databases to be collected is scheduled, and the current collection task is split into several corresponding subtasks and then the target data query script is executed. During the execution process, the current collection data is automatically pushed to the target end. Data can be automatically collected, processed and pushed in batches without the use of scheduling tools, thereby improving data integration efficiency and user operation experience.

[0050] It should be understood that the application content is not intended to identify the key or essential features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings are included to provide a further understanding of the present application. They are incorporated into and constitute a part of this application. The accompanying drawings illustrate embodiments of the present application and, together with this specification, serve to explain the principles of the present application. In the accompanying drawings:

[0052] Figure 1 This is a schematic diagram of a data collection and push system provided by this application as an example;

[0053] Figure 2 This is a flowchart of a data collection and push method provided by this application as an example;

[0054] Figure 3 This is a schematic diagram of a data push process provided by this application as an example;

[0055] Figure 4This is a timing diagram of a task execution process provided by this application as an example;

[0056] Figure 5 This is a schematic diagram of an electronic device exemplarily provided in this application. DETAILED DESCRIPTION

[0057] The principle of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is merely for illustrative purposes and helps those skilled in the art to understand and implement the present disclosure without placing any restriction on the scope of the present disclosure. The disclosure described herein can be implemented in a manner different from that described below.

[0058] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0059] References in this disclosure to "one embodiment," "an embodiment," "an exemplary embodiment," etc., indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment necessarily includes the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. In addition, when a particular feature, structure, or characteristic is described in conjunction with an exemplary embodiment, whether or not explicitly described, those skilled in the art will recognize that such feature, structure, or characteristic may be combined with other embodiments.

[0060] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed terms.

[0061] The terms used herein are intended only to describe specific embodiments and are not intended to limit exemplary embodiments. As used herein, the singular forms "a," "an," and "the" also include the plural forms, unless the context clearly indicates otherwise. As used herein, "a group of elements" or "a set of elements" is intended to include one or more elements. It should also be understood that the terms "comprise," "include," "have," "have," "include," and / or "comprising," when used herein, specify the presence of the features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0062] Figure 1 This is a schematic diagram of a data collection and push system 100 provided by this application. Figure 1 In this embodiment, the data collection and push system 100 may include: a presentation layer 10 , a scheduling layer 20 , a splitting layer 30 , an execution layer 40 and a storage layer 50 .

[0063] Among them, the presentation layer 10 is used to obtain the connection information and task configuration files of multiple databases. The connection information includes the connection addresses, ports and connection permissions of multiple databases. The task configuration files include task configuration parameters and multiple data query scripts adapted to the database types of multiple databases. The task configuration parameters include the execution cycle, task type and batch push conditions. The data query script is used to automatically execute data query operations for the corresponding database type based on pre-configured session parameters and the target end.

[0064] Specifically, the data collection and push system 100 is compatible with various database types, including but not limited to relational databases, non-relational databases, etc. Task types can include high-frequency tasks, ordinary tasks, and large tasks, and can also be defined based on actual business needs.

[0065] In some embodiments, the connection information may also include the database name, version, user name and password for logging into the database, etc. The task configuration parameters may also include the task name, task version, etc. This application does not impose any restrictions on this.

[0066] Continue to see Figure 1 In this embodiment, the presentation layer 10 may include: a database connection configuration module 11, a script warehouse management module 12, a collection task configuration module 13 and an approval management module 14; wherein,

[0067] The database connection configuration module 11 provides a configuration interface and a batch import and export interface for obtaining connection information through the configuration interface and the batch import and export interface, and supports the validity verification of the connection information and the connection test function.

[0068] Specifically, the database connection configuration module 11 provides users with a visual configuration interface, allowing them to enter new connection information within it to add newly added databases. A batch import and export interface allows for integration with third-party database management platforms, allowing these platforms to batch import or export connection information for a series of databases to the database connection configuration module 11, streamlining user operations. Furthermore, the database connection configuration module 11 supports connection information validity verification and connection testing, ensuring that subsequent data collection and push system 100 can correctly transmit data to the corresponding databases to be collected.

[0069] The script repository management module 12 is used to create multi-data query scripts and supports version control, retrieval, editing, deletion, preview and history recording functions of the multi-data query scripts.

[0070] Specifically, the script repository management module 12 allows users to create different data query language (DQL) scripts based on different database types, and supports user-defined configuration of each data query script's version, session parameters, and target end. The target end includes structured databases and stream processing platforms. For multiple created data query scripts, the script repository management module supports version control, retrieval, editing, deletion, preview, and history logging functions to facilitate script management and subsequent troubleshooting.

[0071] The acquisition task configuration module 13 is used to obtain task configuration parameters, generate a task configuration file according to the task configuration parameters, multiple data query scripts and preset task templates, and support task priority setting functions and task dependency configuration functions.

[0072] Specifically, the collection task configuration module 13 supports receiving user-defined task configuration parameters and generating task configuration files based on these task configuration parameters, associated data query scripts, and preset task templates, allowing for the rapid creation of new collection tasks. The collection task configuration module 13 supports setting task priorities, for example, prioritizing important tasks and delaying common tasks. The collection task configuration module 13 also supports configuring task dependencies, for example, configuring the logical execution order between different collection tasks so that subsequent execution follows this logical execution order.

[0073] The approval management module 14 is used to create an approval process for multiple data query scripts and task configuration files, and record the corresponding approval history when the approval process is completed.

[0074] Specifically, the approval process can be used to conduct step-by-step approval of some or all of the creation, editing, activation and deactivation operations in the established multiple data query scripts, and can also be used to conduct step-by-step approval of some or all of the start, stop, trigger, editing and other operations in the established multiple collection tasks, so as to ensure the security and standardization of data query scripts and task configuration files.

[0075] The scheduling layer 20 is used to schedule current collection tasks involving several databases to be collected according to the execution cycle.

[0076] Specifically, for example, if the business needs to collect data from database A and database B in batches every other day, the scheduling layer will automatically schedule the current collection tasks involving database A and database B at a specified time every other day.

[0077] It can be understood that the current collection task may be a separate collection task involving one database to be collected, or a batch collection task involving multiple databases to be collected.

[0078] The splitting layer 30 is used to split the current collection task into a number of subtasks corresponding one-to-one to a number of databases to be collected, and push the subtasks to the message queue.

[0079] Continue to see Figure 1 In some embodiments, the splitting layer 30 includes: an instance screening module 31 , a script screening module 32 , a task splitting module 33 , a task identification module 34 and a push message queue module 35 .

[0080] The instance screening module 31 is used to screen out several databases to be collected corresponding to the current collection task from multiple databases.

[0081] For example, assuming that the current collection task involves database A and database B, database A and database B are identified as databases to be collected.

[0082] The script screening module 32 is used to screen out a target data query script corresponding to the current acquisition task from multiple data query scripts.

[0083] For example, assuming that database A is a relational database and database B is a non-relational database, data query script a corresponding to the relational database and data query script b corresponding to the non-relational database are screened out from the established data query scripts, and data query script a and data query script b are confirmed as target data query scripts.

[0084] The task splitting module 33 is used to split the current collection task into a number of subtasks corresponding one-to-one to a number of databases to be collected.

[0085] Exemplarily, the current acquisition task is split into subtask 1 corresponding to database A and subtask 2 corresponding to database B.

[0086] The task identification module 34 is used to identify the current task type corresponding to the current acquisition task.

[0087] The push message queue module 35 is used to push several subtasks to the message queue corresponding to the current task type.

[0088] For example, assuming that the current task type is a common task, subtask 1 and subtask 2 are pushed to a message queue dedicated to executing common tasks.

[0089] The execution layer 40 is used to remotely connect to several databases to be collected based on the connection information, traverse and execute each target data query script associated with the subtask corresponding to each database to be collected, obtain the current collected data, and during the execution process, in response to the current collected data meeting the batch push conditions, push the current collected data to the target end.

[0090] Continue to see Figure 1 In some embodiments, the execution layer 40 includes: a database remote connection module 41, a script execution module 42, a cursor query module 43, a data processing module 44, a resource dynamic expansion module 45 and a data push module 46.

[0091] The database remote connection module 41 is used to remotely connect to a number of databases to be collected based on the connection information.

[0092] Exemplarily, database A and database B are remotely connected through JDBC technology.

[0093] The script execution module 42 is used to traverse and execute each target data query script associated with the subtask corresponding to each database to be collected. During the execution process, abnormal events are captured according to the preset task timeout time and the execution timeout time of a single script.

[0094] Specifically, the total task execution time is recorded starting from the time when the current collection task starts. During the execution process, ensure that the execution time of data query script a and the execution time of data query script b do not exceed the execution timeout of a single script. Abnormal events include: the total task execution time of the current collection task exceeds the task timeout, and the execution time of a data query script exceeds the execution timeout of a single script.

[0095] The cursor query module 43 is used to perform cursor reading on the execution result of a single target data query script during execution, thereby reducing the consumption of system memory by the script through cursor reading.

[0096] The data processing module 44 is used to perform format processing on the currently collected data, for example, converting the format of the currently collected data into a specified data format.

[0097] The resource dynamic expansion module 45 is used to dynamically adjust the amount of software and hardware resources required for the current acquisition task.

[0098] Exemplarily, the resource dynamic expansion module 45 can dynamically adjust the concurrency according to the task type and the amount of data to be collected to ensure efficient execution of the task.

[0099] The data push module 46 is used to push the current collected data to the target end in response to the current collected data meeting the batch push condition during execution. The batch push condition includes: the data volume of the current collected data exceeds a preset batch processing data volume threshold.

[0100] During the execution process, the current collected data that meets the preset batch processing data volume will be pushed to the target end, thereby realizing the batch data push process.

[0101] The storage layer 50 is used to record the corresponding execution result in response to the completion of the current acquisition task, and the execution result includes the task execution status, execution time and abnormal events.

[0102] Figure 2 This is a flow chart of a data collection and push method 200 provided by this application. Figure 1 The data collection and push system 100 shown includes:

[0103] S201, obtaining connection information and task configuration files of multiple databases.

[0104] Among them, the connection information includes the connection address, port and connection permissions of multiple databases. The task configuration file includes task configuration parameters and multiple data query scripts adapted to the database types of multiple databases. The task configuration parameters include execution cycle, task type and batch push conditions. The data query script is used to automatically execute data query operations for the corresponding database type based on pre-configured session parameters and target end.

[0105] S202: Scheduling current collection tasks involving several databases to be collected according to an execution cycle.

[0106] It can be understood that the current collection task may be a separate collection task involving one database to be collected, or a batch collection task involving multiple databases to be collected.

[0107] Figure 3 This is a schematic diagram of a data push process 300 provided by this application. Figure 3The data push process 300 starts at S301, obtains the task configuration file of the current collection task, and then executes S302 to determine whether the current collection task is a batch collection task. If so, executes S3031 to query multiple database types involved in the batch collection task, and then executes S3041 to encapsulate the task configuration parameters of each database type; if in step S302, the current collection task is not a batch collection task, executes S3032 to query the database type of the database to be collected, and then executes S3042 to encapsulate the task configuration parameters of the database type; after steps S3041 and S3042 are completed, execute S305 to set the session parameters and the target end, and then execute S306 to determine the message queue corresponding to the current collection task, and end the data push process 300.

[0108] S203: Split the current collection task into a number of subtasks corresponding to a number of databases to be collected, and push the subtasks to a message queue corresponding to the current task type.

[0109] In some embodiments, the current collection task is split into a number of subtasks corresponding to a number of databases to be collected, and the subtasks are pushed to the message queue, including:

[0110] Filter out several databases to be collected that correspond to the current collection task from multiple databases.

[0111] Filter out the target data query script corresponding to the current acquisition task from multiple data query scripts.

[0112] The current collection task is split into several subtasks corresponding to several databases to be collected.

[0113] Identify the current task type corresponding to the current collection task.

[0114] Push several subtasks to the message queue corresponding to the current task type.

[0115] S204, accessing several databases to be collected through a remote database connection, traversing and executing each target data query script associated with the subtask corresponding to each database to be collected, obtaining the current collected data, and during the execution process, in response to the current collected data meeting the push conditions, pushing the current collected data to the target end.

[0116] In some embodiments, traversing and executing each target data query script associated with a subtask corresponding to each database to be collected to obtain current collected data, and in response to satisfying a push condition, pushing the current collected data to the target end, includes:

[0117] Traverse and execute each target data query script associated with the subtask corresponding to each database to be collected. During the execution process, perform the following steps:

[0118] Capture exception events based on the preset task timeout and single script execution timeout.

[0119] Use a cursor to read the execution results of a single target data query script.

[0120] In response to the current collected data meeting the push condition, the current collected data is pushed to the target end, and the batch push condition includes: the data volume of the current collected data exceeds a preset batch processing data volume threshold.

[0121] S205: In response to the completion of the current acquisition task, the corresponding execution result is recorded.

[0122] The execution results include task execution status, execution time, and abnormal events.

[0123] Figure 4 This is a timing diagram of a task execution process 400 provided by this application as an example. Figure 4 The task execution process 400 involves business services, JDBC connections, SQL executors, and target terminals 404. Business services, JDBC connections, and SQL executors are processes in the data collection and push system 100. The task execution process 400 includes the following steps:

[0124] S401: The business service requests JDBC to obtain a connection.

[0125] S402: The business service sends the execution timeout of a single script and the SQL cursor size to the SQL executor.

[0126] S403, the SQL executor executes the SQL paging query and returns the data to the business service.

[0127] S404, the business service counts the current data volume.

[0128] S405: When the amount of currently collected data reaches a threshold for batch processing data, the business service pushes the currently collected data to the target end.

[0129] S406: The business service summarizes and records the execution results.

[0130] S407: When the execution time of the current collection task exceeds the task timeout, the business service requests JDBC to cancel the current collection task.

[0131] Based on the above method, this application schedules the current collection tasks involving several databases to be collected according to the execution cycle, splits the current collection tasks into several corresponding sub-tasks, and then executes each target data query script, and automatically pushes the current collection data to the target end during the execution process. It can automatically collect, process and push data in batches without using scheduling tools, thereby improving data integration efficiency and user operation experience.

[0132] Further, if Figure 5 An exemplary embodiment of the present application also provides an electronic device 500, comprising one or more memories 501 and one or more processors 502, wherein the one or more memories 501 are coupled to the one or more processors 502 and store instructions thereon, and the instructions can be executed individually or collectively by the one or more processors 502, so that the electronic device 500 performs any method as in the first aspect.

[0133] It should be understood that the processor mentioned in the embodiments of the present application may be a CPU, or may be other general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0134] It should also be understood that the memory mentioned in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory, dynamic random access memory, synchronous dynamic random access memory, double data rate synchronous dynamic random access memory, enhanced synchronous dynamic random access memory, synchronously linked dynamic random access memory, and direct memory bus random access memory.

[0135] The present application also provides a non-transitory computer-readable storage medium storing machine-executable instructions, wherein the computer-executable instructions can be executed by one or more processors of a machine. The machine may include the electronic device mentioned above, etc. When the computer-executable instructions are executed by the one or more processors, the machine performs any of the methods mentioned above.

[0136] A computer-readable storage medium may include a propagated data signal embodying computer program code, for example, in baseband or as part of a carrier wave. The propagated signal may be in a variety of forms, including electromagnetic, optical, etc., or a suitable combination thereof. The computer-readable storage medium may be connected to an instruction execution system, device, or apparatus to communicate, propagate, or transmit the program for use. The program code on the computer-readable storage medium may be transmitted via any suitable medium, including radio, cable, fiber optic cable, radio frequency signal, or similar medium, or any combination of the above.

[0137] The basic concepts have been described above. It will be apparent to those skilled in the art that the above disclosures are merely examples and do not limit the present application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and revisions to the present application. Such modifications, improvements, and revisions are suggested in the present application and remain within the spirit and scope of the exemplary embodiments of the present application.

[0138] At the same time, this application uses specific terms to describe the embodiments of this application. For example, "one embodiment," "an embodiment," and / or "some embodiments" refer to a certain feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that "one embodiment," "an embodiment," or "an alternative embodiment" mentioned twice or multiple times in different locations in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application may be appropriately combined.

[0139] Some aspects of the present application can be performed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software can be referred to as "data blocks", "modules", "engines", "units", "components" or "systems". The processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DAPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors or combinations thereof. In addition, various aspects of the present application may be expressed as computer products located in one or more computer-readable media, which include computer-readable program code. For example, computer-readable media may include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, tapes...), optical disks (e.g., compact disks CDs, digital versatile disks DVDs...), smart cards, and flash memory devices (e.g., cards, sticks, key drives...).

[0140] A computer-readable medium may include a propagated data signal embodying computer program code, for example, in baseband or as part of a carrier wave. The propagated signal may be in a variety of forms, including electromagnetic, optical, etc., or a suitable combination thereof. A computer-readable medium may be any computer-readable medium other than a computer-readable storage medium that can be connected to an instruction execution system, apparatus, or device to communicate, propagate, or transmit the program for use. The program code on the computer-readable medium may be transmitted via any suitable medium, including radio, cable, fiber optic cable, radio frequency signal, or similar medium, or any combination of the above.

[0141] Similarly, it should be noted that, in order to simplify the description of this application and thus facilitate understanding of one or more embodiments of the application, the foregoing description of the embodiments of this application sometimes combines multiple features into a single embodiment, figure, or description thereof. However, this disclosure method does not mean that the subject matter of this application requires more features than those recited in the claims. In fact, the features of an embodiment may be fewer than all the features of the individual embodiments disclosed above.

[0142] In some embodiments, numbers are used to describe the quantity of components and attributes. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise stated, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the description and claims are approximate values, which may change according to the required features of individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of the present application are approximate values, in specific embodiments, the settings of such numerical values ​​are as accurate as possible within the feasible range.

[0143] Although the present application has been described with reference to the current specific embodiments, ordinary technicians in this technical field should recognize that the above embodiments are only used to illustrate the present application, and various equivalent changes or substitutions can be made without departing from the spirit of the present application. Therefore, as long as the changes and modifications to the above embodiments are within the scope of the essential spirit of the present application, they will fall within the scope of the claims of the present application.

Claims

1. A data collection and push system, characterized in that: include: Presentation layer, scheduling layer, splitting layer, execution layer and storage layer; among them, The presentation layer is used to obtain connection information and task configuration files for multiple databases. The connection information includes the connection addresses, ports, and connection permissions of the multiple databases. The task configuration files include task configuration parameters and multiple data query scripts adapted to the database types of the multiple databases. The task configuration parameters include an execution cycle, a task type, and batch push conditions. The data query script is used to automatically execute data query operations for the corresponding database types based on pre-configured session parameters and the target end. The scheduling layer is used to schedule current collection tasks involving several databases to be collected according to the execution cycle; The splitting layer is used to split the current acquisition task into a plurality of subtasks corresponding one-to-one to the plurality of databases to be acquired, and push the plurality of subtasks to a message queue; The execution layer is used to remotely connect to the plurality of databases to be collected according to the connection information, traverse and execute each target data query script associated with the subtask corresponding to each database to be collected, obtain the current collected data, and during the execution process, in response to the current collected data meeting the batch push condition, push the current collected data to the target end; The storage layer is used to record corresponding execution results in response to the completion of the execution of the current acquisition task, and the execution results include task execution status, execution time and abnormal events.

2. The data collection and push system according to claim 1, wherein: The presentation layer includes a database connection configuration module, a script warehouse management module, a collection task configuration module and an approval management module; wherein, The database connection configuration module provides a configuration interface and a batch import and export interface for obtaining the connection information through the configuration interface and the batch import and export interface, and supports the validity verification and connection test functions of the connection information; The script repository management module is used to create the multi-data query script and support version control, retrieval, editing, deletion, preview and history recording functions of the multi-data query script; The acquisition task configuration module is used to obtain the task configuration parameters, generate the task configuration file according to the task configuration parameters, the multi-data query script and the preset task template, and support the task priority setting function and the task dependency configuration function; The approval management module is used to create an approval process for the multi-data query script and the task configuration file, and record the corresponding approval history when the approval process is completed.

3. The data collection and push system according to claim 1 or 2, characterized in that: The splitting layer includes: instance screening module, script screening module, task splitting module, task identification module and push message queue module. The instance screening module is used to screen out the plurality of databases to be collected corresponding to the current collection task from the multiple databases; The script screening module is used to screen out the target data query script corresponding to the current acquisition task from the multiple data query scripts; The task splitting module is used to split the current acquisition task into a plurality of subtasks corresponding one-to-one to the plurality of databases to be acquired; The task identification module is used to identify the current task type corresponding to the current acquisition task; The push message queue module is used to push the multiple subtasks to the message queue corresponding to the current task type.

4. The data collection and push system according to claim 1 or 2, characterized in that: The execution layer includes: a database remote connection module, a script execution module, a cursor query module, a data processing module, a resource dynamic expansion module and a data push module; wherein, The database remote connection module is used to remotely connect to the plurality of databases to be collected according to the connection information; The script execution module is used to traverse and execute each target data query script associated with the subtask corresponding to each database to be collected, and during the execution process, capture the abnormal event according to the preset task timeout time and the execution timeout time of a single script; The cursor query module is used to perform cursor reading on the execution result of a single target data query script during the execution process; The data processing module is used to perform format processing on the currently collected data; The resource dynamic expansion module is used to dynamically adjust the amount of software and hardware resources required for the current acquisition task; The data push module is used to push the current collected data to the target end in response to the current collected data meeting the batch push condition during the execution process. The batch push condition includes: the data volume of the current collected data exceeds a preset batch processing data volume threshold.

5. The data collection and push system according to claim 1 or 2, characterized in that: The target end includes a structured database and a stream processing platform.

6. A data collection and push method, characterized in that: The data collection and push system according to any one of claims 1 to 5 comprises: Obtaining connection information and task configuration files for multiple databases, wherein the connection information includes the connection addresses, ports, and connection permissions of the multiple databases; the task configuration files include task configuration parameters and multiple data query scripts adapted to the database types of the multiple databases; the task configuration parameters include an execution cycle, a task type, and batch push conditions; and the data query script is used to automatically execute data query operations for the corresponding database types based on pre-configured session parameters and the target end; Scheduling current collection tasks involving a plurality of databases to be collected according to the execution cycle; Splitting the current acquisition task into a plurality of subtasks corresponding one-to-one to the plurality of databases to be acquired, and pushing the plurality of subtasks to a message queue corresponding to the current task type; Accessing the plurality of databases to be collected through a remote database connection, traversing and executing each target data query script associated with a subtask corresponding to each database to be collected, obtaining current collected data, and in response to the current collected data satisfying the push condition, pushing the current collected data to the target end; In response to the completion of the execution of the current acquisition task, a corresponding execution result is recorded, where the execution result includes the task execution status, execution time, and abnormal events.

7. The method according to claim 6, wherein Splitting the current acquisition task into a plurality of subtasks corresponding to the plurality of databases to be acquired, and pushing the plurality of subtasks to a message queue, including: Filtering the plurality of databases to be collected corresponding to the current collection task from the multiple databases; Filtering the target data query script corresponding to the current acquisition task from the multiple data query scripts; Splitting the current acquisition task into a plurality of subtasks corresponding one-to-one to the plurality of databases to be acquired; Identifying a current task type corresponding to the current acquisition task; and Push the multiple subtasks to the message queue corresponding to the current task type.

8. The method according to claim 7, wherein Traversing and executing each target data query script associated with the subtask corresponding to each database to be collected to obtain current collected data, and in response to satisfying the push condition, pushing the current collected data to the target end, including: Traverse and execute each target data query script associated with the subtask corresponding to each database to be collected. During the execution process, perform the following steps: Capturing the abnormal event according to the preset task timeout and single script execution timeout; Use a cursor to read the execution results of a single target data query script; In response to the current collected data meeting the push condition, the current collected data is pushed to the target end, and the batch push condition includes: the data volume of the current collected data exceeds a preset batch processing data volume threshold.

9. An electronic device comprising: one or more processors; as well as One or more memories coupled to the one or more processors and storing thereon instructions, which, when executed individually or collectively by the one or more processors, cause the electronic device to perform the method according to any one of claims 6 to 8.

10. A non-transitory computer-readable storage medium storing machine-executable instructions, which, when executed by one or more processors of a machine, cause the machine to perform the method of any one of claims 6 to 8.