Lightweight data batch processing method, system and device and medium

By receiving parameters and generating configuration files through the configuration interface, building timed scheduling tasks are realized, and the automated execution of data batch processing is solved, which solves the problems of high complexity and cost of data batch processing in the existing technology, and improves processing efficiency and flexibility.

CN120030037APending Publication Date: 2025-05-23SHANGHAI E&P INT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510106742.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing technology is difficult to meet the data batch processing needs under lightweight conditions. The mainstream service framework is heavier, the learning cost is high, and the operation of pure SQL batch processing tasks is relatively scarce.

Method used

Receive configuration parameters through the configuration interface and automatically generate configuration files, simplifying task settings and reducing operational difficulty. Build timed scheduling tasks based on configuration parameters to realize automated execution of tasks and reduce manual intervention. Executors and batch components work together to ensure accurate triggering and efficient processing of tasks.

Benefits of technology

It significantly improves data processing efficiency and flexibility, reduces the complexity and cost of data batch processing, supports multiple data sources and processing logic, and meets diverse data processing needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030037A_ABST
    Figure CN120030037A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight data batch processing method, system and device and a medium, and relates to the technical field of data processing.The method comprises the steps that configuration parameters corresponding to batch processing tasks are received; generating a configuration file corresponding to the configuration parameter through a batch processing configuration service; flexible configuration and rapid deployment of batch processing tasks are realized; constructing a timed scheduling task corresponding to the batch processing task based on the configuration parameters; based on the timed scheduling task, starting an executor and a batch processing component to process the current batch processing task; wherein task triggering is performed based on the timed scheduling task, and a current batch processing task is sent to the executor; obtaining a current configuration file corresponding to the current batch processing task through the actuator; reading the current configuration file, and performing task initialization on the current batch processing task; and executing the current batch processing task through the batch processing component, thereby remarkably reducing the complexity and cost of data batch processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a lightweight data batch processing method, system, electronic device and readable storage medium. Background Art

[0002] As batch processing tasks gradually mature, various batch processing components emerge in an endless stream. From a traditional perspective, commercial ETL tools such as Informatica and Kettle are used to run data batch processing tasks, and small batch processing tasks are also run using stored procedures. However, as big data technology gradually deepens, data access gradually increases, and the threshold requirements must be gradually lowered; moreover, the overall framework of the current mainstream services in the market is generally heavy, the learning cost is relatively high, and the operation of batch processing tasks implemented by pure SQL is relatively scarce, which cannot meet the data batch processing under relatively lightweight conditions.

[0003] Therefore, a lightweight data batch processing method, system, electronic device and readable storage medium need to be proposed. Summary of the invention

[0004] This specification provides a lightweight data batch processing method, system, electronic device and readable storage medium, which receives configuration parameters through a configuration interface and automatically generates configuration files, simplifies task settings and reduces operational difficulty. At the same time, based on the configuration parameters, a scheduled scheduling task is constructed to realize the automatic execution of tasks and reduce manual intervention. The collaborative work of the executor and batch processing components ensures the accurate triggering and efficient processing of tasks. In addition, the method also supports a variety of data sources and processing logics to meet diverse data processing needs. Through refined data extraction, processing and writing rules, efficient data flow and accurate storage are achieved. This lightweight data batch processing method significantly improves data processing efficiency and flexibility.

[0005] A lightweight data batch processing method provided in this application adopts the following technical solutions, including:

[0006] Receive configuration parameters corresponding to the batch processing task;

[0007] Generate a configuration file corresponding to the configuration parameters through a batch configuration service;

[0008] Constructing a scheduled scheduling task corresponding to the batch processing task based on the configuration parameters;

[0009] Based on the timed scheduling task, the executor and the batch processing component are started to process the current batch processing task; wherein, based on the timed scheduling task, the task is triggered and the current batch processing task is sent to the executor; the current configuration file corresponding to the current batch processing task is obtained through the executor; the current configuration file is read and the current batch processing task is initialized; and the current batch processing task is executed through the batch processing component.

[0010] Optionally, the configuration parameters include: data source information, processing logic information, and timing information;

[0011] Optionally, the generating of a configuration file corresponding to the configuration parameter by batch configuration service includes:

[0012] Building the configuration file based on the configuration parameters;

[0013] Saving the configuration file to a batch configuration database;

[0014] Optionally, the step of constructing a timed scheduling task corresponding to the batch processing task includes:

[0015] Constructing identification information for the batch processing task;

[0016] Combining the identification information and the timing information, configuring the timing task component to construct the timing scheduling task;

[0017] An association relationship between the task identifier and the configuration file is established.

[0018] Optionally, constructing the configuration file based on the configuration parameters includes:

[0019] configuring data extraction rules based on the input data source of the data source information;

[0020] An output data source configuration data writing rule based on the data source information;

[0021] configuring data processing rules based on the processing logic information;

[0022] The data extraction rules, the data writing rules and the data processing rules are summarized to construct the configuration file.

[0023] Optionally, obtaining, by the executor, a current configuration file corresponding to the current batch processing task includes:

[0024] Sending a configuration file acquisition request to the batch configuration service through the executor;

[0025] The batch configuration service searches the batch configuration database for a current configuration file corresponding to the current identification information;

[0026] The batch configuration service returns the current configuration file to the executor.

[0027] Optionally, executing the current batch processing task by the batch processing component includes:

[0028] Based on the data extraction rule, obtaining the original data in the input data source through a reader;

[0029] Performing data processing on the original data based on the data processing rules to obtain target data;

[0030] Determine a write type of the target data based on the data write rule; and write the target data into the output data source in combination with the write type.

[0031] Optionally, the step of writing the target data to the output data source based on the write type of the target data includes:

[0032] When the data storage mode is the update mode, determining the type of the target data and generating a determination result;

[0033] Generate a corresponding write statement based on the judgment result; wherein, when the target data is of the first write type, generate an Insert statement based on the target data, and the Insert statement is used to insert new data; when the target data is of the second write type, generate an Update statement based on the target data, and the Update statement is used to update the existing data of the output data source;

[0034] The writer is used to execute a write statement corresponding to the target data, and the target data is written into the output data source.

[0035] Optionally, also include:

[0036] Generate log records based on the execution status of the executor;

[0037] Optionally, generating a log record based on the execution status of the executor includes:

[0038] During the execution of the batch processing task, log records are obtained; when the batch processing task is completed, analysis is performed based on the log records, key information is extracted, and the key information is stored in the database; the key information includes: the execution status of the batch processing task and the number of written data.

[0039] The lightweight data batch processing system provided in this application adopts the following technical solutions, including:

[0040] The acquisition module is used to receive the configuration parameters corresponding to the batch processing task;

[0041] A batch configuration service, used to generate a configuration file corresponding to the configuration parameters;

[0042] A scheduling task construction module, used to construct a timed scheduling task corresponding to the batch processing task based on the configuration parameters;

[0043] A batch processing module, used to start the executor and batch processing component to process the current batch processing task based on the timed scheduling task;

[0044] The batch processing module comprises:

[0045] A timed task component, used to trigger a task based on the timed scheduling task and send the current batch processing task to the executor;

[0046] An executor is used to obtain a current configuration file corresponding to the current batch processing task; read the current configuration file, and perform task initialization on the current batch processing task;

[0047] The batch processing component is used to execute the current batch processing task.

[0048] Optionally, the configuration parameters include: data source information, processing logic information, and timing information;

[0049] Optionally, the batch configuration service includes:

[0050] The step of generating a configuration file corresponding to the configuration parameters through batch configuration service includes:

[0051] A configuration file construction submodule, used to construct the configuration file based on the configuration parameters;

[0052] A configuration file storage submodule, used for saving the configuration file to a batch configuration database;

[0053] Optionally, the scheduling task construction module includes:

[0054] An identification information construction submodule, used to construct identification information for the batch processing task;

[0055] A scheduling task construction submodule, used to configure the scheduled task component in combination with the identification information and the timing information, and to construct the scheduled scheduling task;

[0056] The association relationship establishing submodule is used to establish an association relationship between the task identifier and the configuration file.

[0057] Optionally, the configuration file constructs a submodule including:

[0058] An extraction rule construction unit, configured to configure data extraction rules based on an input data source of the data source information;

[0059] A write rule construction unit, configured to configure data write rules based on the output data source of the data source information;

[0060] A processing rule building unit, used to configure data processing rules based on the processing logic information;

[0061] The configuration file generating unit is used to summarize the data extraction rules, the data writing rules and the data processing rules to construct the configuration file.

[0062] Optionally, the actuator comprises:

[0063] The configuration file acquisition unit is used to send a configuration file acquisition request to the batch configuration service; based on the batch configuration service, the batch configuration service searches for a current configuration file corresponding to the current identification information from the batch configuration database, and acquires the current configuration file returned by the batch configuration service.

[0064] Optionally, the batch processing component includes:

[0065] A reader, configured to obtain the original data in the input data source based on the data extraction rule;

[0066] A processor, configured to perform data processing on the original data based on the data processing rule to obtain target data;

[0067] A writer is used to determine a write type of target data based on the data write rule; and write the target data into the output data source in combination with the write type.

[0068] Optionally, the writer includes:

[0069] A judging unit, configured to judge the type of the target data and generate a judging result when the data storage mode is the updating mode;

[0070] A generating unit, used for generating a corresponding write statement according to the judgment result;

[0071] A writing unit is used to execute a writing statement corresponding to the target data, and write the target data into the output data source.

[0072] Optionally, the generating unit includes:

[0073] A first writing subunit, configured to generate an Insert statement based on the target data when the target data is of a first writing type, wherein the Insert statement is used to insert new data;

[0074] A second writing subunit is used to generate an Update statement based on the target data when the target data is of a second writing type, wherein the Update statement is used to update the existing data of the output data source;

[0075] Optionally, it also includes: a recording module, used to generate log records based on the execution status of the executor;

[0076] Optionally, the recording module includes:

[0077] A log record acquisition submodule, used to acquire log records during the execution of the batch processing task;

[0078] An analysis submodule, used for analyzing based on the log records, extracting key information, and storing the key information in a database when the batch processing task is completed;

[0079] The key information includes: the execution status of the batch processing task and the number of written data.

[0080] This specification also provides an electronic device, wherein the electronic device includes:

[0081] processor; and,

[0082] A memory storing computer executable instructions, which when executed cause the processor to perform any of the above methods.

[0083] The present specification also provides a computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by a processor, any of the above methods is implemented.

[0084] In the present application, configuration parameters corresponding to a batch task are received; a configuration file corresponding to the configuration parameters is generated through a batch configuration service; flexible configuration and rapid deployment of batch tasks are achieved; a scheduled scheduling task corresponding to the batch task is constructed based on the configuration parameters; based on the scheduled scheduling task, an executor and a batch component are started to process the current batch task; wherein, task triggering is performed based on the scheduled scheduling task, and the current batch task is sent to the executor; the current configuration file corresponding to the current batch task is obtained through the executor; the current configuration file is read and the current batch task is initialized; the current batch task is executed through the batch component, thereby improving its lightness and efficiency, and significantly reducing the complexity and cost of data batch processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] Figure 1 A schematic diagram of the principle of a lightweight data batch processing method provided in an embodiment of this specification;

[0086] Figure 2 A schematic diagram of a local data refresh processing flow of a lightweight data batch processing method provided in an embodiment of this specification;

[0087] Figure 3 A flowchart of a lightweight data batch processing method provided in an embodiment of this specification;

[0088] Figure 4 A schematic diagram of the structure of a lightweight data batch processing system provided in an embodiment of this specification;

[0089] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of this specification;

[0090] Figure 6 A schematic diagram of a computer-readable medium provided for an embodiment of this specification. DETAILED DESCRIPTION

[0091] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art can think of other obvious variations. The basic principles of the present invention defined in the following description can be applied to other embodiments, variations, improvements, equivalents, and other technical solutions that do not deviate from the spirit and scope of the present invention.

[0092] Exemplary embodiments of the present invention will now be described more fully with reference to the accompanying drawings. However, exemplary embodiments can be implemented in a variety of forms, and should not be construed as limiting the present invention to the embodiments set forth herein. On the contrary, providing these exemplary embodiments enables the present invention to be more comprehensive and complete, and is more convenient for fully conveying the inventive concept to those skilled in the art. The same reference numerals in the figures represent the same or similar elements, components or parts, and thus their repeated description will be omitted.

[0093] Under the premise of being consistent with the technical concept of the present invention, the features, structures, characteristics or other details described in a specific embodiment do not exclude that they can be combined in one or more other embodiments in a suitable manner.

[0094] In the description of specific embodiments, the features, structures, characteristics or other details described in the present invention are intended to enable those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can practice the technical solutions of the present invention without one or more of the specific features, structures, characteristics or other details.

[0095] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0096] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0097] The term "and / or" or "and / or" includes all combinations of any one or more of the associated listed items.

[0098] Figure 1 A schematic diagram of the principle of a lightweight data batch processing method provided in an embodiment of this specification, the method comprising:

[0099] S1 receives the configuration parameters corresponding to the batch processing task;

[0100] S2 generates a configuration file corresponding to the configuration parameters through a batch configuration service;

[0101] S3 constructs a scheduled scheduling task corresponding to the batch processing task based on the configuration parameters;

[0102] S4 starts the executor and the batch processing component to process the current batch processing task based on the timed scheduling task; wherein, based on the timed scheduling task, the task is triggered and the current batch processing task is sent to the executor; the current configuration file corresponding to the current batch processing task is obtained through the executor; the current configuration file is read and the current batch processing task is initialized; and the current batch processing task is executed through the batch processing component.

[0103] As batch processing tasks gradually mature, various batch processing components emerge one after another. From a traditional perspective, commercial ETL tools such as Informatica and Kettle are used to run data batch processing tasks, and small batch processing tasks are also run using stored procedures.

[0104] However, as big data technology becomes more sophisticated and data access becomes more frequent, the threshold requirements must also be gradually lowered. However, the overall framework of the mainstream services currently on the market is generally heavy, with relatively high learning costs. There is a relative lack of pure SQL to implement batch processing tasks, and they cannot meet the needs of data batch processing under relatively lightweight conditions. Componentized operations are required for this scenario.

[0105] The mainstream batch processing products on the market now are based on the open source Hadoop MapReduce, Spark, and Flink, and are exported as SAAS products for use. Direct use can basically meet the needs, but it faces the following problems:

[0106] (1) The products are relatively heavyweight. Many of them are implemented based on the Hadoop ecosystem and provided in a SAAS manner. When the business only requires relational database output, it is often relatively cumbersome to use and there is no targeted optimization product. In addition, the private deployment versions of common batch processing products are relatively heavy and cannot be applied in relatively small-scale scenarios.

[0107] (2) Traditional batch processing components can be based on the Hadoop ecosystem and re-flush the data of file blocks according to the file segmentation method to meet the needs of data refresh when demand changes or use the concept of sliding time to increase the robustness of data batch processing tasks. However, when data is stored in a relational database as a data warehouse, data refresh usually adopts a mode of covering all table data, and it is impossible to refresh based on some data in the table.

[0108] (3) Traditional ETL tools usually rely on dragging components and writing Java, Python and other programs to implement, and cannot be implemented through pure SQL statements. The latest batch processing tools are more likely to use open source Spark and Flink frameworks, which have the ability to be implemented in pure SQL, but they also need to configure data source configuration and other content into SQL, which has a high learning cost.

[0109] In order to improve the processing efficiency of batch processing tasks, the present invention provides a lightweight data batch processing method, comprising:

[0110] S1 receives the configuration parameters corresponding to the batch processing task;

[0111] A configuration interface is pre-set; the configuration interface includes: a number of configuration information, and the configuration information includes configuration items and configuration areas.

[0112] Display the configuration interface to the user; the user can enter the configuration parameters corresponding to the configuration items in the configuration area of ​​the configuration interface. After the user completes the input, receive all the configuration parameters entered by the user.

[0113] Users refer to the configuration personnel, users, etc. of this system.

[0114] The configuration parameters include: data source information, processing logic information, and timing information;

[0115] The data source information includes: input data source and output data source.

[0116] The input data source and / or the output data source is a lightweight relational data warehouse, which can be any relational database and Hadoop. In one embodiment of the present specification, any relational database includes but is not limited to: MySQL.

[0117] The processing logic information includes, but is not limited to: the extraction logic of the batch processing task, the calculation logic of the batch processing task, and the writing logic of the batch processing task.

[0118] In one embodiment of the present specification, the writing logic of the batch processing task includes: a data storage mode.

[0119] The data storage modes include: Append mode, Update mode, and Overwrite mode.

[0120] The append mode is used to indicate: directly adding data; that is, directly adding new data to the end of the existing data table.

[0121] The update mode is used to indicate: update if there is, insert otherwise; that is, check the output data source, and if the record exists, update the record; if it does not exist, you can choose to insert a new record.

[0122] Overwrite mode is used to indicate that the table is deleted and then rebuilt; that is, the existing data table is deleted and recreated based on the new data.

[0123] Timing information includes: execution time and execution frequency.

[0124] The configuration page provided by the present invention fully standardizes the batch processing part originally implemented by code, and abstracts the reading configuration, conversion calculation configuration (SQL statement writing), and writing configuration; improves the convenience of users inputting various configuration parameters required for batch processing tasks; this configuration method reduces the research and development cost of batch processing tasks, so that users do not need to have an in-depth understanding of complex code implementation, and only need to master SQL syntax and basic concepts of batch processing tasks to complete the release and development of batch processing tasks.

[0125] S2 generates a configuration file corresponding to the configuration parameters through a batch configuration service;

[0126] Specifically, the batch configuration service uses SpringBoot as the development framework.

[0127] S21 constructs the configuration file based on the configuration parameters;

[0128] S211 configures data extraction rules based on the data source information;

[0129] configuring data extraction rules based on the input data source of the data source information;

[0130] Specifically, an SQL statement or path for data extraction is configured according to the input data source as a data extraction rule.

[0131] S212 configures data writing rules based on the data source information;

[0132] An output data source configuration data writing rule based on the data source information;

[0133] Specifically, the table name and field mapping of the written data are configured according to the output data as the data writing rule.

[0134] S213 configures data processing rules based on the processing logic information;

[0135] Use SparkSQL to write data processing logic. Specifically, perform partial code refactoring of the JDBC plug-in based on business needs.

[0136] S214 summarizes the data extraction rules, the data writing rules, and the data processing rules to construct the configuration file.

[0137] By constructing configuration files, various configuration information of batch tasks is standardized and structured, providing a solid foundation for subsequent task execution.

[0138] S22 saves the configuration file to a batch configuration database;

[0139] S3 constructs a scheduled scheduling task corresponding to the batch processing task based on the configuration parameters;

[0140] S31 constructs identification information for the batch processing task;

[0141] The identification information includes: task identification and task version number.

[0142] S32 configures the timing task component in combination with the identification information and the timing information to construct the timing scheduling task;

[0143] Encapsulate the XXL-JOB distributed scheduled tasks and build a scheduled task component, which is a distributed scheduled task component.

[0144] After building the scheduled task, save the scheduled task configuration to the calling distributed scheduled task xxl-job. The scheduled task includes: scheduled time, task identifier;

[0145] The present invention packages and integrates xxl-job into the system. Users do not need to configure in xxl-job, but perform scheduling configuration through the configuration page of the system. By calling xxl-job and storing the scheduling task in xxl-job, the configuration of xxl-job can be completed. In one embodiment of the present specification, a scheduling interface is constructed to configure the scheduling type and scheduling syntax, so as to better realize the configuration and saving of the scheduled scheduling task, so that the opening of xxl-job can be completed by clicking to save the task.

[0146] S33 establishes an association relationship between the task identifier and the configuration file.

[0147] Based on the association relationship, ensure that the corresponding configuration file can be correctly loaded when the task is executed.

[0148] The present invention implements the function of configuration management based on the batch processing scenario under lightweight scheduled tasks. The functions of configuration management include: data extraction, data calculation and data writing. Before configuring data extraction and data writing, it is necessary to configure the data source first, so that a data source can be selected in the drop-down box when configuring the batch processing task; while saving the configuration, if the task is a scheduled scheduling task, the distributed scheduled task will be called to store the task. The present invention also supports task isolation according to tenants to meet the needs of different R&D groups, multiple batch processing clusters, different R&D groups, and mixed batch processing clusters.

[0149] The batch configuration service of the present invention constructs a configuration file based on the configuration parameters input by the user and saves it to the batch configuration database. At the same time, identification information is constructed for the batch task, and a timed scheduling task is constructed in combination with the timing information. This step realizes the standardization and centralized management of configuration information, improves the maintainability of tasks, and facilitates subsequent task execution and monitoring.

[0150] S4 starts the executor and the batch processing component to process the current batch processing task based on the timed scheduling task;

[0151] The executor of the present invention is a timed task processing executor. Specifically:

[0152] S41 triggers a task based on the scheduled task and sends the current batch task to the executor;

[0153] This patent adopts the mode of configuring service + actuator architecture without connecting the actuator to the database.

[0154] S42 obtains the current configuration file corresponding to the current batch processing task through the executor;

[0155] S421 sends a configuration file acquisition request to the batch configuration service through the executor;

[0156] The configuration file acquisition request includes current identification information of the batch processing task;

[0157] S422, the batch configuration service searches the batch configuration database for a current configuration file corresponding to the current identification information;

[0158] The batch configuration service obtains a configuration file acquisition request, extracts current identification information, and searches for a corresponding current configuration file from a batch configuration database based on the current identification information.

[0159] S423 The batch configuration service returns the current configuration file to the executor.

[0160] After the batch configuration service obtains the current configuration file, the batch configuration service returns the task configuration result to the scheduled task processing executor, and the task configuration result includes: the current configuration file. The executor obtains the current configuration file returned by the batch configuration service.

[0161] S43 reads the current configuration file and initializes the current batch processing task;

[0162] Before starting the current batch processing task, you need to parse the configuration and load the Java dependencies.

[0163] Read and parse the information in the current configuration file, including data source information, processing logic information, and timing information.

[0164] Start the current batch processing task; the internal batch processing task initializes the current batch processing task;

[0165] The entire initialization process is divided into three stages: data input, batch processing logic, and writing and loading.

[0166] Data input phase: Configure and load the data reader according to the input data source in the current configuration file. This may involve finding and loading the driver package corresponding to the data source format (such as Hadoop's data source jar package).

[0167] Batch processing logic stage: According to the processing logic information in the current configuration file, configure and load the SparkSQL processing logic. This may require necessary code refactoring or configuration of the JDBC plug-in to implement specific data processing functions.

[0168] Writing and loading phase: According to the output data source in the current configuration file, configure and load the writer. Again, this may involve finding and loading the driver package corresponding to the data source format.

[0169] The present invention uses different configurations to implement batch processing tasks in different business scenarios based on the executor of batch processing task configuration. Through the above steps, it is ensured that the task is triggered at the right time point and can accurately load the required configuration files and dependencies. This provides the necessary guarantee for the smooth execution of the task.

[0170] S44 executes the current batch processing task through the batch processing component.

[0171] The executor of the present invention triggers the batch processing component to start the Spark task according to the configured Spark cluster, reads the current configuration file, and implements reader loading, SparkSQL loading and writer loading.

[0172] The batch processing component is based on the SparkSQL open source framework and is executed using Spark, so as to optimize and enhance the local data refresh scenarios of batch processing tasks based on the characteristics of lightweight relational data warehouses.

[0173] S441, based on the data extraction rule, obtaining the original data in the input data source through a reader;

[0174] S442 performs data processing on the original data based on the data processing rule to obtain target data;

[0175] Specifically, the original data is processed by SparkSQL processing logic to obtain target data;

[0176] S443 determines the write type of the target data based on the data write rule; and writes the target data into the output data source in combination with the write type.

[0177] Considering that Spark's own JDBC plug-in and overwrite mode will cause the entire table to be refreshed in a relational database, it is impossible to refresh local data based on fixed rules. In order to achieve the update of the output data source, the present invention adds an update mode (Update mode) to the data storage mode on the premise of being compatible with the append mode (Append mode) and the overwrite mode (Overwrite mode). The reader is not modified, only the writer is modified, and the code is rewritten in the configuration generation stage. In one embodiment of the present specification, the batch processing component is a Scala SDK that implements the ability to refresh local data in the data table. The capability is enhanced based on the open source framework of SparkSQL, and the Spark code is rewritten. Of course, the present invention meets the needs of batch task tuning while realizing the function, and the mainstream Spark configuration is paged.

[0178] At this time, when the current batch task requires partial update or incremental insertion of data, the user selects / adds the Update mode and configures the update strategy primary key field; the update strategy primary key field includes but is not limited to: update frequency, update data range, update data source, and how to deal with conflicts or errors that may occur during the update process.

[0179] By querying the output data source before writing, the data existing in the output data source is obtained, and the data is updated if there is data, and inserted if there is no data, to achieve partial update in jdbc mode.

[0180] S443-1: when the data storage mode is the update mode, determine the type of the target data and generate a determination result;

[0181] In one embodiment of the present specification, a batch of target data is obtained through RDD; for each batch of target data, the corresponding update field primary key list is found by batch, and a select query is performed;

[0182] Specifically, the configured updateUniqueKey node is used as a query condition to find the corresponding data in the output data source.

[0183] If the data cannot be found, it means that insertion is required, and at this time, the target data is identified as the first write type. The first write type is used to represent incremental data.

[0184] If the data is found, the data needs to be refreshed (updated), and at this time, the target data is identified as the second write type. The second write type is used to represent the overwritten data.

[0185] S443-2 generating a corresponding write statement according to the judgment result;

[0186] Generate update and insert statements according to the judgment results to facilitate data batch writing operations, thereby realizing batch processing capabilities in the batch processing scenario of lightweight relational database warehousing.

[0187] In one embodiment of the present specification, the batch task local data refresh process is as follows: Figure 2 As shown, determine the primary key of update; generate Insert and Update statements and Select statements for Update according to the written content; query by unique primary key according to batches; determine whether the query has results; if there are results, perform Update operation according to the unique primary key; if there are no results, perform Insert operation on the unique primary key that does not exist.

[0188] Wherein, when the target data is of the first write type, an Insert statement is generated based on the target data, and the Insert statement is used to insert new data; when the target data is of the second write type, an Update statement is generated based on the target data, and the Update statement is used to update the existing data of the output data source;

[0189] The executor of the present invention uses the SparkSQL framework as the underlying technology, and enhances the writing function (writer) of the relational database by implementing the plug-in mode of Spark.

[0190] Based on the open source JDBC plug-in, we refactored the local code, added the Update update mode to the original SaveMode mode, and added the Update update strategy primary key field to the Option mode in Spark to enable configuration for different update fields. After obtaining the corresponding configuration, the component generates the corresponding Insert statement and Update statement according to the final write data structure.

[0191] S443-3 uses the writer to execute the write statement corresponding to the target data, and writes the target data into the output data source.

[0192] After generating the write statement, the actual writing work is performed. Identify and generate query statements in batches to obtain existing data from the output data source; perform partial updates. Through the writer, write the corresponding target data to the output data source.

[0193] When the batch Scala SDK is configured with an updated SaveMode, the present invention generates Insert and Update statements based on the structure and configuration of batch warehousing; executes query statements in batches to obtain judgment results; generates final Insert data sets and Update data sets according to the judgment results; executes Insert statements and Update statements to implement the operation of partial data refresh.

[0194] The batch processing component of the present invention can efficiently process a large amount of data and supports multiple data writing modes to meet the needs of different scenarios. At the same time, the ability to refresh local data improves the efficiency and accuracy of data updates.

[0195] The timing scheduling task of the present invention triggers the execution of the batch processing task according to the preset execution time and execution frequency. The executor initializes the task by obtaining the current configuration file and calls the batch processing component (such as Spark) to execute the specific processing logic to realize the automatic execution and efficient processing of the task.

[0196] S5 generates a log record based on the execution status of the executor;

[0197] In one embodiment of the present specification, log storage is also provided, and based on batch processing tasks, the number of fields operated and whether the operation is successful are extracted to assist operation and maintenance personnel in facilitating operation and maintenance.

[0198] Specifically, S51 obtains log records during the execution of the batch processing task;

[0199] During the execution of the batch processing task, necessary log records are printed and returned in the executor;

[0200] S52: when the batch processing task is finished, analyzing based on the log records, extracting key information, and storing the key information in the database;

[0201] The key information includes: the execution status of the batch processing task and the number of written data.

[0202] The execution status of the batch processing task is used to indicate whether the batch processing task is successfully executed.

[0203] In one embodiment of the present specification, after key information is stored in the database, operations such as logging, batch task destruction, and clearing of corresponding execution status may be performed.

[0204] The present invention provides a convenient operation and maintenance means for operation and maintenance personnel by recording key information such as the execution status of the task and the number of written data. These log records help to discover and solve problems in a timely manner and ensure the stable operation of the task.

[0205] The present invention can be applied to multiple scenarios, and preferably, it can be applied to the data processing field of government affairs and ports. Specifically, it can extract the multi-dimensional information of the enterprise such as industry and commerce, social security, and penalties from multiple input data sources on a regular basis every day, calculate the latest annual report year, extract the latest social security number, and calculate the number of double public license penalties, and then aggregate the basic information of industry and commerce to form a basic information table of the enterprise, and put it into the ADS library of DAMO as a data warehouse; and in order to ensure that the incremental operation of information does not lose data, its writing is carried out using the Update mode.

[0206] From the perspective of the overall architecture, the batch configuration service of the present invention adopts SpringBoot as the development framework, the data batch processing component uses Spark, and the scheduled task component is encapsulated for the XXL-JOB distributed scheduled task. Therefore, the corresponding configuration and distributed scheduled tasks are connected with data.

[0207] The entire batch task processing system adopts a lightweight design concept, combining open source technologies such as SpringBoot, Spark and XXL-JOB, and realizes functions such as configuration management, task execution, and logging. The system can be applied to multiple scenarios, such as government affairs and port data processing, to improve the efficiency and accuracy of data processing; at the same time, the maintainability and scalability of the system are also well guaranteed.

[0208] like Figure 3 As shown, the processing system includes a batch configuration service, a batch configuration database, a distributed scheduled task component, a scheduled task processing executor, a Spark program, and a lightweight relational data warehouse;

[0209] Among them, the scheduled task processing executor includes a distributed task processor, a data management component, and a batch task manager; the distributed task processor includes a batch task executor;

[0210] The following briefly describes the execution logic of batch processing tasks;

[0211] The user enters the pre-configuration parameters through the configuration web page, and then publishes the configuration online.

[0212] Based on user operations, publish scheduled tasks, and / or stop scheduled tasks, and / or delete scheduled tasks;

[0213] After publishing the scheduled task, the batch task is triggered based on the node discovery of the distributed scheduled task component;

[0214] When a batch task is triggered, the scheduled task processing executor obtains the configuration files required for task configuration from the batch configuration database through the batch configuration service; that is, when the scheduled task is triggered to execute, the execution node will obtain the configuration, find the corresponding Java dependency according to the obtained configuration as the input of the loaded Jar package for task startup, and then complete the task initialization and run;

[0215] Submit batch tasks to the batch Scala SDK in the Spark program through the batch task executor; after processing, read / store them into the lightweight relational data warehouse. Use the data management component to view and manage data structures.

[0216] Figure 4 A schematic diagram of the structure of a lightweight data batch processing system provided in an embodiment of this specification, the system includes:

[0217] The acquisition module 410 is used to receive configuration parameters corresponding to the batch processing task;

[0218] Batch configuration service 420, used to generate a configuration file corresponding to the configuration parameters;

[0219] A scheduling task construction module 430 is used to construct a timed scheduling task corresponding to the batch processing task based on the configuration parameters;

[0220] The batch processing module 440 is used to start the executor and the batch processing component to process the current batch processing task based on the timed scheduling task;

[0221] The batch processing module 440 includes:

[0222] A timed task component, used to trigger a task based on the timed scheduling task and send the current batch processing task to the executor;

[0223] An executor is used to obtain a current configuration file corresponding to the current batch processing task; read the current configuration file, and perform task initialization on the current batch processing task;

[0224] The batch processing component is used to execute the current batch processing task.

[0225] Optionally, the configuration parameters include: data source information, processing logic information, and timing information;

[0226] Optionally, the batch configuration service 420 includes:

[0227] The batch configuration service 420 generates a configuration file corresponding to the configuration parameters, including:

[0228] A configuration file construction submodule, used to construct the configuration file based on the configuration parameters;

[0229] A configuration file storage submodule, used for saving the configuration file to a batch configuration database;

[0230] Optionally, the scheduling task construction module 430 includes:

[0231] An identification information construction submodule, used to construct identification information for the batch processing task;

[0232] A scheduling task construction submodule, used to configure the scheduled task component in combination with the identification information and the timing information, and to construct the scheduled scheduling task;

[0233] The association relationship establishing submodule is used to establish an association relationship between the task identifier and the configuration file.

[0234] Optionally, the configuration file constructs a submodule including:

[0235] An extraction rule construction unit, configured to configure data extraction rules based on an input data source of the data source information;

[0236] A write rule construction unit, configured to configure data write rules based on the output data source of the data source information;

[0237] A processing rule building unit, configured to configure data processing rules based on the processing logic information;

[0238] The configuration file generating unit is used to summarize the data extraction rules, the data writing rules and the data processing rules to construct the configuration file.

[0239] Optionally, the actuator includes:

[0240] The configuration file acquisition unit is used to send a configuration file acquisition request to the batch configuration service 420; based on the batch configuration service 420, the batch configuration service 420 searches for a current configuration file corresponding to the current identification information from the batch configuration database, and acquires the current configuration file returned by the batch configuration service 420.

[0241] Optionally, the batch processing component includes:

[0242] A reader, configured to obtain the original data in the input data source based on the data extraction rule;

[0243] A processor, configured to perform data processing on the original data based on the data processing rule to obtain target data;

[0244] A writer is used to determine a write type of target data based on the data write rule; and write the target data into the output data source in combination with the write type.

[0245] Optionally, the writer includes:

[0246] A judging unit, configured to judge the type of the target data and generate a judging result when the data storage mode is the updating mode;

[0247] A generating unit, used for generating a corresponding write statement according to the judgment result;

[0248] A writing unit is used to execute a writing statement corresponding to the target data, and write the target data into the output data source.

[0249] Optionally, the generating unit includes:

[0250] A first writing subunit, configured to generate an Insert statement based on the target data when the target data is of a first writing type, wherein the Insert statement is used to insert new data;

[0251] A second writing subunit is used to generate an Update statement based on the target data when the target data is of a second writing type, wherein the Update statement is used to update the existing data of the output data source;

[0252] Optionally, it also includes: a recording module, used to generate log records based on the execution status of the executor;

[0253] Optionally, the recording module includes:

[0254] A log record acquisition submodule, used to acquire log records during the execution of the batch processing task;

[0255] An analysis submodule, used for analyzing based on the log records, extracting key information, and storing the key information in a database when the batch processing task is completed;

[0256] The key information includes: the execution status of the batch processing task and the number of written data.

[0257] The functions of the system of the embodiment of the present invention have been described in the above method embodiment, so for details not fully described in this embodiment, please refer to the relevant description in the above embodiment, and no further description will be given here.

[0258] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0259] The electronic device embodiment of the present invention is described below, and the electronic device can be regarded as a physical implementation of the method and device embodiments of the present invention described above. The details described in the electronic device embodiment of the present invention should be regarded as a supplement to the above method or device embodiments; details not disclosed in the electronic device embodiment of the present invention can be implemented with reference to the above method or device embodiments.

[0260] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of this specification. Figure 5 The computer device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0261] like Figure 5 As shown, the computer device 500 of this exemplary embodiment is in the form of a general data processing device. The components of the computer device 500 may include but are not limited to: at least one processor 510, at least one memory 520, a network interface 530, a display unit 540, an input component 550, etc.

[0262] The memory 520 stores a computer-readable program, which may be a source program or a code of a read-only program. The program may be executed by the processor 510, so that the processor 510 performs the steps of various embodiments of the present invention. For example, the processor 510 may perform the following steps: Figure 1 Steps shown.

[0263] The memory 520 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) and / or a cache memory unit, and may further include a read-only memory unit (ROM). The memory 520 may also include a program / utility having a set (at least one) of program modules, such program modules including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include the implementation of a network environment.

[0264] Also included is a bus (not shown) which may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0265] The computer device 500 may also communicate with one or more external devices (e.g., keyboard, display, network device, Bluetooth device, etc.) so that a user can interact with the computer device 500 via these external devices, and / or so that the computer device 500 can communicate with one or more other data processing devices (e.g., routers, modems, etc.). Such communication may be performed through the network interface 530, and may also be performed through a network adapter with one or more networks (e.g., local area network (LAN), wide area network (WAN) and / or public network, such as the Internet). The network adapter may communicate with other modules of the computer device 500 via a bus. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in the computer device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0266] Figure 6 Schematic diagram of a computer readable medium embodiment of the present invention. Figure 6 As shown, the computer program can be stored on one or more computer-readable media. The computer-readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. When the computer program is executed by one or more data processing devices, the computer-readable medium is enabled to implement the above method of the present invention.

[0267] Through the description of the above implementation modes, it is easy for those skilled in the art to understand that the exemplary embodiments described in the present invention can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation mode of the present invention can be embodied in the form of a software product, which can be stored in a computer-readable storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a data processing device (which can be a personal computer, a server, or a network device, etc.) to perform the above method according to the present invention.

[0268] The computer readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by an instruction execution system, an apparatus, or a device or used in combination with it. The program code contained on the readable storage medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0269] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0270] In summary, the present invention can be implemented by a method, apparatus, electronic device or computer-readable medium that executes a computer program. In practice, a general data processing device such as a microprocessor or a digital signal processor (DSP) can be used to implement some or all functions of the present invention.

[0271] The specific embodiments described above further describe the purpose, technical solutions and beneficial effects of the present invention in detail. It should be understood that the present invention is not inherently related to any specific computer, virtual device or electronic device, and various general devices can also implement the present invention. The above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A lightweight data batch processing method, characterized in that: include: Receive configuration parameters corresponding to the batch processing task; Generate a configuration file corresponding to the configuration parameters through a batch configuration service; Constructing a scheduled scheduling task corresponding to the batch processing task based on the configuration parameters; Based on the timed scheduling task, the executor and the batch processing component are started to process the current batch processing task; wherein, based on the timed scheduling task, the task is triggered and the current batch processing task is sent to the executor; the current configuration file corresponding to the current batch processing task is obtained through the executor; the current configuration file is read and the current batch processing task is initialized; and the current batch processing task is executed through the batch processing component.

2. A lightweight data batch processing method as claimed in claim 1, characterized in that: The configuration parameters include: data source information, processing logic information, and timing information; The step of generating a configuration file corresponding to the configuration parameters through batch configuration service includes: Building the configuration file based on the configuration parameters; Saving the configuration file to a batch configuration database; The step of constructing a timed scheduling task corresponding to the batch processing task includes: Constructing identification information for the batch processing task; Combining the identification information and the timing information, configuring the timing task component to construct the timing scheduling task; An association relationship between the task identifier and the configuration file is established.

3. A lightweight data batch processing method as claimed in claim 2, characterized in that: The step of constructing the configuration file based on the configuration parameters includes: configuring data extraction rules based on the input data source of the data source information; An output data source configuration data writing rule based on the data source information; configuring data processing rules based on the processing logic information; The data extraction rules, the data writing rules and the data processing rules are summarized to construct the configuration file.

4. A lightweight data batch processing method as claimed in claim 2, characterized in that: The obtaining, through the executor, a current configuration file corresponding to the current batch processing task includes: Sending a configuration file acquisition request to the batch configuration service through the executor; The batch configuration service searches the batch configuration database for a current configuration file corresponding to the current identification information; The batch configuration service returns the current configuration file to the executor.

5. A lightweight data batch processing method as claimed in claim 3, characterized in that: The executing the current batch processing task by the batch processing component includes: Based on the data extraction rule, obtaining the original data in the input data source through a reader; Performing data processing on the original data based on the data processing rules to obtain target data; Determine a write type of the target data based on the data write rule; and write the target data into the output data source in combination with the write type.

6. A lightweight data batch processing method as claimed in claim 5, characterized in that: The step of writing the target data to the output data source based on the write type of the target data comprises: When the data storage mode is the update mode, determining the type of the target data and generating a determination result; Generate a corresponding write statement based on the judgment result; wherein, when the target data is of the first write type, generate an Insert statement based on the target data, and the Insert statement is used to insert new data; when the target data is of the second write type, generate an Update statement based on the target data, and the Update statement is used to update the existing data of the output data source; The writer is used to execute a write statement corresponding to the target data, and the target data is written into the output data source.

7. A lightweight data batch processing method as claimed in claim 1, characterized in that: Also includes: Generate log records based on the execution status of the executor; wherein, during the execution of the batch task, obtain log records; when the batch task is completed, analyze based on the log records, extract key information, and store the key information in the database; the key information includes: the execution status of the batch task and the number of written data.

8. A lightweight data batch processing system, characterized in that: include: The acquisition module is used to receive the configuration parameters corresponding to the batch processing task; A batch configuration service, used to generate a configuration file corresponding to the configuration parameters; A scheduling task construction module, used to construct a timed scheduling task corresponding to the batch processing task based on the configuration parameters; A batch processing module, used to start the executor and batch processing component to process the current batch processing task based on the timed scheduling task; The batch processing module comprises: A timed task component, used to trigger a task based on the timed scheduling task and send the current batch processing task to the executor; An executor is used to obtain a current configuration file corresponding to the current batch processing task; read the current configuration file, and perform task initialization on the current batch processing task; The batch processing component is used to execute the current batch processing task.

9. An electronic device, wherein: The electronic device includes: processor; and, A memory storing computer executable instructions which, when executed, cause the processor to perform a method according to any one of claims 1-7.

10. A computer-readable storage medium, wherein: The computer-readable storage medium stores one or more programs, and when the one or more programs are executed by a processor, the method of any one of claims 1 to 7 is implemented.