Data processing scheduling system and method based on unified SQL (Structured Query Language) architecture
By adopting a data processing scheduling system based on a unified SQL architecture in the big data processing system, the problems of waste of computing resources and low development, operation and maintenance efficiency are solved, and efficient computing resource utilization and flexible data processing processes are realized.
Patent Information
- Application Number
- CN202510071548.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-13
AI Technical Summary
The existing big data processing architecture has problems such as waste of computing resources and low development and operation and maintenance efficiency, especially when handling complex data tasks, the dependencies between tasks are complex and the scheduling process requires a lot of manual intervention, resulting in limited system stability and compatibility.
The data processing and scheduling system based on a unified SQL architecture is adopted. Data processing tasks are written through interactive modules. The scheduling module schedules and executes SQL files and algorithm requests. The SQL syntax adapter module performs syntax adaptation. The SQL gateway module sends tasks to the calculation engine for execution and releases resources after execution. The algorithm request module supports asynchronous calls of external algorithms.
It effectively reduces waste of computing resources, improves development and operation and maintenance efficiency, realizes seamless integration of multiple computing engines, reduces technical thresholds, and improves the flexibility of big data processing systems and the efficiency of computing resource utilization.
Smart Images

Figure CN120144274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and particularly to a data processing and scheduling system and method based on a unified SQL architecture. Background Art
[0002] With the rapid progress of information technology, the amount of data faced by enterprises is increasing at an alarming rate, and the types of data are becoming increasingly rich and diverse. In order to fully explore and utilize the value of these massive data, enterprises urgently need to build a complex and refined data processing process, which covers multiple key links such as data cleaning, transformation, analysis, and storage, aiming to achieve the efficient integration and utilization of data.
[0003] In the traditional big data processing architecture, multiple independent Spark applications are usually used to complete a series of complex data processing tasks such as data cleaning, data integration, and algorithm invocation. Since the development and implementation methods of different tasks are not unified and may involve multiple programming languages and frameworks, the management and maintenance between programs become extremely cumbersome. Developers need to master multiple technology stacks, increasing the learning and switching costs. At the same time, the stability and compatibility of the system also face challenges; for complex data processing tasks, the dependency relationships between tasks are intricate, and the configuration of the scheduling process often requires a large amount of manual intervention, which not only increases the risk of errors but also limits the realization of automated management, making it difficult for the entire data processing process to achieve an efficient operation state; in addition, during the execution of Spark applications, when facing a large number of metric dimensions to calculate or a large volume of data to process, it is easy to trigger memory overflow problems due to an overly long task chain or unreasonable memory allocation; at the same time, some tasks (such as AI algorithm invocation) need to wait for the return results of external services, during which the computing resources of Spark cannot be released, resulting in waste of computing resources and task delays, thereby affecting the operation and maintenance efficiency. Therefore, how to reduce the waste of computing resources and improve the development and operation and maintenance efficiency during big data processing is an extremely important technical problem to be solved. Summary of the Invention
[0004] To solve the problems of waste of computing resources, low development, and operation and maintenance efficiency existing in the above-mentioned prior art, the present invention proposes a data processing and scheduling system and method based on a unified SQL architecture, which effectively reduces the waste of computing resources and improves the development and operation and maintenance efficiency.
[0005] To achieve the above technical effects, the technical solution of the present invention is as follows:
[0006] A data processing and scheduling system based on a unified SQL architecture, comprising:
[0007] An interaction module for users to write data processing tasks and deploy the written data processing tasks to the SQL gateway module or the scheduling module. The data processing tasks include multiple SQL files and algorithm requests that are executed in sequence.
[0008] A scheduling module for scheduling and recording the execution of the data processing tasks, transmitting the SQL files to the SQL syntax adapter module, and transmitting the algorithm requests to the algorithm request module.
[0009] An SQL syntax adapter module for performing syntax adaptation on the SQL files and transmitting the SQL files with adapted syntax to the SQL gateway module.
[0010] An SQL gateway module for receiving the SQL files, sending the SQL files to the computing engine for execution to obtain execution results, and outputting the execution results to the storage module.
[0011] An algorithm request module for receiving the algorithm requests, calling external algorithms to process local data to obtain algorithm request results, and outputting the algorithm request results to the storage module.
[0012] A storage module for storing the execution results and the algorithm request results.
[0013] Preferably, it further includes a computing engine module for hosting the computing engine, and the computing engine is a Spark computing engine or a Flink computing engine.
[0014] Preferably, it further includes a data source mapping module for mapping an external data source into a temporary table in the storage module through the computing engine. When querying the temporary table, the data source mapping module converts the temporary table into a query of the external data source.
[0015] Preferably, the external data source includes a standardized data source and a non-standardized data source. The standardized data source is a data source implemented based on a preset standard protocol, and the non-standardized data source is an entity-open data source.
[0016] Preferably, the performing syntax adaptation on the SQL files includes: the SQL syntax adapter module receives the SQL files transmitted by the scheduling module, and performs syntax conversion of the SQL statements in the SQL files according to the target execution engine environment where the current SQL file is located to obtain the SQL files with adapted syntax.
[0017] Preferably, the calling of an external algorithm to process local data includes: making a communication agreement between the algorithm request module and an external system, the algorithm request module transmitting the local data to the external system, the external system receiving the local data, calling an external algorithm to perform synchronous request decoupling on the local data to obtain a synchronous request decoupling result, and storing the synchronous request decoupling result in a data table to obtain the algorithm request result.
[0018] Preferably, both the execution result and the algorithm request result are data tables.
[0019] The present invention also proposes a data processing scheduling method based on a unified SQL architecture, which is characterized by including the following steps:
[0020] S1. A user writes a data processing task and deploys the written data processing task to a SQL gateway module or a scheduling module. The data processing task includes multiple SQL files and algorithm requests that are executed in sequence.
[0021] S2. The scheduling module schedules and records the execution of the data processing task, transmits the SQL file to a SQL syntax adapter module, and transmits the algorithm request to an algorithm request module.
[0022] S3. The SQL syntax adapter module performs syntax adaptation on the SQL file and transmits the SQL file with adapted syntax to the SQL gateway module.
[0023] S4. The SQL gateway module receives the SQL file, sends the SQL file to a computing engine for execution to obtain an execution result, and outputs the execution result to a storage module.
[0024] S5. The algorithm request module receives the algorithm request, calls an external algorithm to process local data to obtain an algorithm request result, and outputs the algorithm request result to the storage module.
[0025] The present invention also proposes a computer device, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete mutual communication through the communication bus;
[0026] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations of the data processing scheduling method based on the unified SQL architecture as described above.
[0027] The present invention also provides a computer-readable storage medium storing at least one executable instruction, which, when running on a computer device, causes the computer device to perform the operations of the data processing and scheduling method based on the unified SQL architecture as described above.
[0028] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:
[0029] The present invention provides a data processing and scheduling system and method based on a unified SQL architecture. First, in the scheduling module, by SQLizing data processing tasks, a data processing process with a unified SQL architecture is realized, effectively improving the flexibility of the big data processing system and the utilization efficiency of computing resources. Secondly, through the SQL syntax adapter module and the unified SQL gateway module, seamless integration of multiple computing engines is achieved, reducing migration and adaptation costs. Then, the task scheduling process takes the SQL file as the core, and complex data processing tasks can be completed by writing SQL and algorithm configurations through the interaction module, reducing the technical threshold and improving development efficiency. Next, after the SQL file is sent to the computing engine for execution, the computing engine resources are released, avoiding waste of computing resources. At the same time, the algorithm request module supports asynchronous calls of external algorithms, making the system run more efficiently and stably. Further, the core logic of data processing is concentrated in the SQL file and algorithm calls, and the execution results and algorithm request results are both stored in the unified storage module, facilitating subsequent analysis and management of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a structural block diagram of a data processing and scheduling system based on a unified SQL architecture proposed in an embodiment of the present invention;
[0031] Figure 2 It is a schematic diagram of the execution principle of a data processing task proposed in an embodiment of the present invention;
[0032] Figure 3 It is another structural block diagram of a data processing and scheduling system based on a unified SQL architecture proposed in an embodiment of the present invention;
[0033] Figure 4 It is a flowchart of a data processing and scheduling method based on a unified SQL architecture proposed in an embodiment of the present invention;
[0034] Figure 5 It is a schematic diagram of the structure of a computer device proposed in an embodiment of the present invention;
[0035] 51. Processor; 52. Memory; 53. Communication interface; 54. Communication bus; 55. Executable instruction. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the patent.
[0037] For those skilled in the art, it is understandable that some well-known content descriptions in the drawings may be omitted.
[0038] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0039] Embodiment 1
[0040] As Figure 1 shown, this embodiment proposes a data processing and scheduling system based on a unified SQL architecture, including:
[0041] An interaction module, which is used for users to write data processing tasks and deploy the written data processing tasks to the SQL gateway module or the scheduling module. The data processing tasks include multiple SQL files and algorithm requests that are executed in sequence. The interaction module is directly connected to the SQL gateway module and executes the specific SQL statements of the SQL files. An SQL file is a text file containing structured query language statements. SQL is a standard programming language for managing and operating relational databases, which allows users to perform various data definition, data operation, and data control tasks.
[0042] A scheduling module, which is used to schedule and record the execution of the data processing tasks, transmit the SQL files to the SQL syntax adapter module, and transmit the algorithm requests to the algorithm request module;
[0043] An SQL syntax adapter module, which is used to perform syntax adaptation on the SQL files and transmit the syntax-adapted SQL files to the SQL gateway module;
[0044] The syntax adaptation of the SQL files includes: The SQL syntax adapter module receives the SQL files transmitted by the scheduling module, performs syntax conversion of the SQL statements in the SQL files for the target execution engine environment where the current SQL file is located, and obtains the syntax-adapted SQL files; The SQL syntax adapter module is to adapt to different syntaxes of different underlying SQL file execution engines. In the SQL syntax adapter module, for the implementation of conflicting SQL files, conditional judgment is used to determine which target execution engine environment the SQL file is in, and the SQL file corresponding to the target execution engine environment is executed to eliminate the impact brought by syntax conflicts; The SQL syntax adapter module is dominated by the scheduling module. After the scheduling module calls the SQL syntax adapter, it connects to the SQL gateway to execute the specific SQL files.
[0045] The SQL gateway module is used to receive the SQL file, send the SQL file to the computing engine for execution to obtain an execution result, and output the execution result to the storage module; the SQL gateway module is the unified entry for SQL file execution. In the SQL gateway module, it connects to a specific computing engine and sends a request to the computing engine for execution.
[0046] Here, refer to Figure 2 , during the execution of the data processing task, the data processing tasks driven by the SQL file are all implemented in SQL language. For a data processing task, it can be implemented in SQL according to metrics or specific business logics, and one metric or logic is written as one SQL file. A data processing task consists of multiple SQL files plus algorithm request configurations, and the task is executed according to the specified file order and algorithm order. After an SQL file is executed, the resulting data table is saved in the storage module, and then the next SQL file is executed. The execution of external algorithms is also output to the storage module. After all SQL files and algorithm requests are executed, a data processing task is completed.
[0047] Figure 2 The data processing task consists of three SQL files and one algorithm request. After an SQL file A is executed, the corresponding data table A is generated, and then the next SQL file B, SQL file C, and SQL file D are executed in sequence. The final output execution result depends on the execution result of the last SQL file.
[0048] The algorithm request module is used to receive the algorithm request, call an external algorithm to process local data to obtain an algorithm request result, and output the algorithm request result to the storage module;
[0049] The calling of the external algorithm to process local data includes: making a communication agreement between the algorithm request module and an external system, the algorithm request module transmitting the local data to the external system, the external system receiving the local data, calling an external algorithm to perform synchronous request decoupling on the local data to obtain a synchronous request decoupling result, and writing the synchronous request decoupling result into a data table to obtain the algorithm request result; at this time, the Spark computing engine is no longer needed, and the occupied Spark computing engine resources are released. It should be specifically stated that the algorithm request module only receives requests from the scheduling module. When an external algorithm call needs to be initiated to process local data in a data processing task, a request is initiated. After being processed with the external system, the data is output back to the storage module. A communication agreement has been made with the external system, and the system input and output can be specified.
[0050] The storage module is used to store the execution result and the algorithm request result; it also stores other data during the data processing process.
[0051] It also includes a computing engine module, which is used to carry the computing engine. The computing engine is a Spark computing engine or a Flink computing engine. The Spark computing engine or the Flink computing engine is a distributed in-memory computing engine for large-scale data processing. The computing engine module is the core of the data processing task. The computing engine module receives specific requests sent by the SQL gateway module, executes relevant logic, and finally outputs the execution results to the storage module.
[0052] It also includes a data source mapping module, which is designed to process data not in the storage module and is used to map an external data source into a temporary table in the storage module through the computing engine. When querying the temporary table, the data source mapping module converts the temporary table into a query of the external data source.
[0053] The external data source includes a standardized data source and a non-standardized data source. The data source mapping module supports the mapping of standardized and non-standardized data sources, enabling the system to have high scalability and allowing various data source types to be flexibly accessed without significantly modifying the underlying architecture. The standardized data source is a data source implemented based on a preset standard protocol, such as relational databases like MySQL and Oracle, and non-relational databases like Redis and Elasticsearch. The standardized data source is mapped through the connection component of the computing engine and can generally handle most standardized data sources.
[0054] The non-standardized data source is a data source open to entities. Entities refer to organizations or companies, etc. The non-standardized data source can generally only be accessed through an interface or a specified code package. The non-standardized data source is mapped through a custom connection component. In the processing of external data sources, the data source mapping module is used to map the data into a temporary data table in the storage module through the computing engine for data source mapping. In the subsequent SQL file, this data table can be used for querying and thus processed.
[0055] In a data processing and scheduling system based on a unified SQL architecture proposed in this embodiment, the system executes two main processes as follows:
[0056] One is the user interaction process. The sequence of the user interaction process is: user, interaction module, SQL gateway module, computing engine module, storage module. The user interaction process is mainly used for developers to debug and use temporarily.
[0057] Second is the task scheduling process, which has two branch processes, including: SQL execution process and algorithm execution process; the SQL execution process is: scheduling module, SQL syntax adapter module, SQL gateway module, computing engine module, storage module. The algorithm execution process is: scheduling module, algorithm request module, storage module.
[0058] In this embodiment, the scheduling module serves as one of the entrances for data processing tasks. Each data processing task consists of multiple SQL files and algorithm requests that are executed in sequence. The SQL files are adjusted by the SQL syntax adapter module and then connected to the SQL gateway for execution; then the algorithm requests are called to execute by the algorithm request module. One SQL file generates a data table as the execution result, and one algorithm request generates a data table as the algorithm request result. Both the execution result and the algorithm request result are output to the storage module, and the data written into the data table is determined by the statements in the SQL file.
[0059] Using a data processing scheduling system based on a unified SQL architecture proposed in this embodiment, the data processing process includes that the user develops specific data processing tasks through the interaction module, deploys the data processing tasks to the scheduling module, configures the corresponding scheduling information, and executes them regularly; the specific SQL files included in the data processing tasks are executed in the specific computing engine module through the SQL gateway module under the management of the scheduling module after the data is processed, and then written into the corresponding data table. When the algorithm is called, it does not need to go through the SQL gateway and is completely managed by the scheduling module. After the algorithm is executed, it is written back to the configured data table for use in the subsequent process; in summary, a data processing process completely driven by SQL is achieved.
[0060] It should be particularly noted that in this embodiment, first, by SQLizing data processing tasks in the scheduling module, a data processing process with a unified SQL architecture is realized, effectively improving the flexibility of the big data processing system and the utilization efficiency of computing resources; second, through the SQL syntax adapter module and the unified SQL gateway module, seamless integration of multiple computing engines is achieved, reducing migration and adaptation costs; then, the task scheduling process takes the SQL file as the core, and complex data processing tasks can be completed by writing SQL and algorithm configurations through the interaction module, reducing the technical threshold and improving development efficiency; then, after the SQL file is sent to the computing engine for execution, the computing engine resources are released, avoiding waste of computing resources. At the same time, the algorithm request module supports asynchronous calls of external algorithms, making the system operation more efficient and stable; further, the core logic of data processing is concentrated in the SQL file and algorithm calls, and the execution results and algorithm request results are both stored in the unified storage module, facilitating subsequent analysis and management of data. Generally speaking, compared with the prior art, the present invention realizes the high unity of task development and scheduling, the efficient management of resource utilization, and the flexible scalability of data processing, and can significantly improve the processing capacity and operation and maintenance efficiency of the big data system.
[0061] Embodiment 2
[0062] See Figure 3 , this embodiment proposes a data processing and scheduling system based on a unified SQL architecture, including a scheduling module, an interaction module, an SQL syntax adapter module, an SQL gateway module, a data source mapping module, an algorithm request module, a computing engine module, and a storage module.
[0063] The interaction module uses tools such as Datagrip or DBeaver, both of which are excellent database management tools. They can be connected to the SQL gateway module through the tools, and then execute the SQL file edited by the user.
[0064] The SQL gateway module adopts the open-source project Kyuubi, a distributed multi-tenant gateway, as the unified access entry to the computing engine. Kyuubi can cache the background engine instances at the user level to better achieve computing resource sharing and fast response, ensure the efficient execution of tasks, and avoid waste of computing resources caused by problems such as calling synchronous interfaces. The SQL gateway module receives the SQL statements submitted by the interaction module and the scheduling module, and then calls the specific computing engine to execute. Developers can choose to access Spark SQL or Flink SQL for data processing according to the nature and requirements of the tasks.
[0065] The data source mapping module processes standardized data sources using the Connectors component of the Spark engine or the Connectors component of Flink. Spark Connectors are components used by Spark to interact with external data sources, encapsulating the logic for connection, reading, and writing, and are ready to use out of the box. The component needs to introduce the corresponding Jar package into the big data cluster and can be used as needed. Flink Connectors are modules used by Flink to connect to external systems and support batch processing and stream processing. They are responsible for reading data (sources) from external data sources or writing data to external targets (sinks).
[0066] To process non-standardized data sources, custom Connector components are used. Both Spark and Filnk provide corresponding interfaces. After implementing the interfaces, a non-standardized data source can be mapped as a temporary table. In this embodiment, the data source mapping module is included in the computing engine, and the non-standardized implementation is loaded onto the call variables of the computing engine in the form of a package and can be loaded when used. It can accept requests from the interaction module or the scheduling module, and use the components of the selected engine to map the corresponding data source as a temporary data table for use in subsequent process calculations.
[0067] The scheduling module uses the open-source project Dagster, a DAG (Directed Acyclic Graph) scheduling tool developed in Python, which is overall responsible for task scheduling, dependency management, and execution order, ensuring that tasks are executed in the correct order and supporting an automatic retry mechanism. Dagster consists of three modules: The Dagit module is the management page module, mainly providing a visual interface for users. The Daemon module is responsible for specific task scheduling. The Code location stores specific Sql files and Python code. In this solution, the scheduling module defines a data processing task flow using code and also accepts task flows from other components.
[0068] The algorithm request module is integrated into the scheduling module. As mentioned before, a general algorithm scheduling capability is defined using the code of the scheduling module. The algorithm scheduling module designs two data tables: the algorithm configuration table and the algorithm scheduling record table. The algorithm configuration table writes the parameter configurations required for different algorithms. The algorithm scheduling record table writes the situation of a single call of a specific algorithm and records the status. When the scheduling module parses the algorithm request configuration, it finds the algorithm scheduling record table according to the algorithm to be called, then initiates a call and writes a record to the algorithm scheduling record table. Subsequently, it polls and waits for the algorithm to complete. After the algorithm is completed, it is marked as completed, exits this call, and allows the scheduling module to execute the next stage of SQL or algorithm call.
[0069] The SQL syntax adapter module is integrated into the scheduling module and deployed together. The SQL syntax adapter module uses the Python framework DBT. Under the DBT framework, the dependency relationships between SQL files can be specified through specific syntax, and DBT will recognize the syntax in each SQL file and give the specific dependency relationships and the order of execution. The scheduling module that inherits the SQL syntax adapter module (DBT) can provide scheduling capabilities according to the dependency order provided by the SQL syntax adapter module.
[0070] As mentioned above, combine DBT and Dagster to achieve the capabilities of SQL-driven data task scheduling, task visualization, and task management.
[0071] The storage module uses Iceberg plus HDFS as the underlying storage, which is a common storage module technology.
[0072] The computing engine module adopts the Spark engine and the Flink engine, which are common big data engines for data.
[0073] Developers use the interactive module to develop a specific data processing task A, which is divided into SQL files according to certain rules and deployed on the Code location component of the scheduling module. Then configure the specific scheduling logic. For example, start at 0:00 every day.
[0074] At 0:00 every day, the Daemon component of the scheduling module scans the configuration information and starts the invocation of the data processing task A. The scheduling module, through the DBT component, identifies the dependencies between each file and sends them to the SQL gateway for execution in sequence. The SQL gateway connects to the specific computing engine, and after the computing engine completes the SQL logic, it writes to the storage module. Then the scheduling module executes the dependencies of task A and sends the next SQL file to the SQL gateway for execution until all SQLs in process A are executed.
[0075] In the case of non-standard data sources, configure the connectors or packages of the specific engine to the engine execution path. Create the corresponding temporary data tables and use the specific table names in the SQL files to operate on the non-standard data sources.
[0076] Embodiment 3
[0077] See Figure 4 , this embodiment proposes a data processing scheduling method based on a unified SQL architecture, which is characterized by including the following steps:
[0078] S1. The user writes a data processing task and deploys the written data processing task to the SQL gateway module or the scheduling module. The data processing task includes multiple SQL files and algorithm requests that are executed in the order of execution;
[0079] S2. The scheduling module schedules and records the execution of the data processing task, transmits the SQL file to the SQL syntax adapter module, and transmits the algorithm request to the algorithm request module;
[0080] S3. The SQL syntax adapter module performs syntax adaptation on the SQL file and transmits the SQL file after syntax adaptation to the SQL gateway module;
[0081] S4. The SQL gateway module receives the SQL file, sends the SQL file to the computing engine for execution to obtain an execution result, and outputs the execution result to the storage module;
[0082] S5. The algorithm request module receives the algorithm request, invokes an external algorithm to process local data to obtain an algorithm request result, and outputs the algorithm request result to the storage module.
[0083] In this embodiment, by unifying batch processing and stream processing tasks under SQL semantics, simplifying the management of task dependencies in the scheduling framework, reusing the computing power of the big data cluster as much as possible, realizing efficient allocation and release of resources, and reducing the complexity of multi-source data integration by abstracting heterogeneous data sources into logical tables. The method of this embodiment supports asynchronous algorithm calls, solves the problem of waste of computing resources caused by waiting for algorithm results, significantly improves the task development and operation and maintenance efficiency, and meets the big data processing requirements under diverse business scenarios. More specifically, first, in the scheduling module, through the SQLization of data processing tasks, a data processing process with a unified SQL architecture is realized, effectively improving the flexibility of the big data processing system and the utilization efficiency of computing resources; secondly, through the SQL syntax adapter module and the unified SQL gateway module, seamless integration of multiple computing engines is realized, reducing the migration and adaptation costs; then the task scheduling process takes the SQL file as the core, and complex data processing tasks can be completed by writing SQL and algorithm configurations through the interaction module, reducing the technical threshold and improving the development efficiency; then, after sending the SQL file to the computing engine for execution, the computing engine resources are released, avoiding waste of computing resources, and at the same time, the algorithm request module supports asynchronous calls of external algorithms, making the system operation more efficient and stable; further, the core logic of data processing is concentrated in the SQL file and algorithm calls, and the execution results and algorithm request results are both stored in the unified storage module, facilitating subsequent analysis and management of data.
[0084] Embodiment 4
[0085] This embodiment also proposes a computer device, see Figure 5, including: a processor 51, a memory 52, a communication interface 53, and a communication bus 54. The processor 51, the memory 52, and the communication interface 53 complete communication with each other through the communication bus 54;
[0086] Among them: The processor 51, the memory 52, and the communication interface 53 complete communication with each other through the communication bus 54. The communication interface 53 is used for network communication with other devices such as clients or other servers. The processor 51 is used to execute executable instructions 55, and specifically can perform the operations of the data processing and scheduling method based on the unified SQL architecture described in the above embodiments. Specifically, the executable instructions 55 may include program code. The processor 51 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the computer device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0087] The memory 52 is used to store the executable instructions 55. The memory 52 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0088] The executable instructions 55 can specifically be called by the processor 51 to cause the computer device to perform the following operations:
[0089] S1. The user writes a data processing task and deploys the written data processing task to the SQL gateway module or the scheduling module. The data processing task includes multiple SQL files and algorithm requests that are executed in the order of execution;
[0090] S2. The scheduling module schedules and records the execution of the data processing task, transmits the SQL file to the SQL syntax adapter module, and transmits the algorithm request to the algorithm request module;
[0091] S3. The SQL syntax adapter module performs syntax adaptation on the SQL file and transmits the SQL file with syntax adaptation to the SQL gateway module;
[0092] S4. The SQL gateway module receives the SQL file, sends the SQL file to the computing engine for execution to obtain an execution result, and outputs the execution result to the storage module;
[0093] S5. The algorithm request module receives the algorithm request, calls an external algorithm to process local data to obtain an algorithm request result, and outputs the algorithm request result to the storage module.
[0094] This embodiment also provides a computer-readable storage medium, in which at least one executable instruction is stored. When the executable instruction runs on a computer device, the computer device is caused to execute the operations of the data processing and scheduling method based on the unified SQL architecture described in the above embodiment.
[0095] In this embodiment, first, in the scheduling module, by SQLizing data processing tasks, a data processing process with a unified SQL architecture is realized, effectively improving the flexibility of the big data processing system and the utilization efficiency of computing resources; second, through the SQL syntax adapter module and the unified SQL gateway module, seamless integration of multiple computing engines is achieved, reducing migration and adaptation costs; then, the task scheduling process takes the SQL file as the core, and complex data processing tasks can be completed by writing SQL and algorithm configurations through the interaction module, reducing the technical threshold and improving development efficiency; then, after the SQL file is sent to the computing engine for execution, the computing engine resources are released, avoiding waste of computing resources. At the same time, the algorithm request module supports asynchronous calls of external algorithms, making the system run more efficiently and stably; further, the core logic of data processing is concentrated in the SQL file and algorithm calls, and the execution results and algorithm request results are both stored in the unified storage module, facilitating subsequent analysis and management of data.
[0096] Obviously, the above embodiments of the present invention are only examples for clearly illustrating the present invention, rather than limiting the embodiments of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the embodiments here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A data processing scheduling system based on a unified SQL architecture, characterized in that: include: The interactive module is used for users to write data processing tasks and deploy the written data processing tasks to the SQL gateway module or the scheduling module. The data processing tasks include multiple SQL files and algorithm requests that are executed in order; A scheduling module, used for scheduling and recording the execution of the data processing task, transmitting the SQL file to the SQL syntax adapter module, and transmitting the algorithm request to the algorithm request module; An SQL grammar adapter module is used to perform grammar adaptation on the SQL file and transmit the grammar-adapted SQL file to the SQL gateway module; An SQL gateway module is used to receive the SQL file, send the SQL file to the computing engine for execution, obtain the execution result, and output the execution result to the storage module; An algorithm request module is used to receive the algorithm request, call an external algorithm to process local data, obtain the algorithm request result, and output the algorithm request result to the storage module; A storage module is used to store the execution result and the algorithm request result.
2. The data processing scheduling system based on the unified SQL architecture according to claim 1 is characterized in that: It also includes a computing engine module, which is used to carry the computing engine, and the computing engine is a Spark computing engine or a Flink computing engine.
3. The data processing scheduling system based on the unified SQL architecture according to claim 1 is characterized in that: It also includes a data source mapping module, which is used to map the external data source into a temporary table in the storage module through the computing engine. When querying the temporary table, the data source mapping module converts the temporary table into a query to the external data source.
4. The data processing scheduling system based on the unified SQL architecture according to claim 3 is characterized in that: The external data source includes a standardized data source and a non-standardized data source. The standardized data source is a data source implemented based on a preset standard protocol, and the non-standardized data source is a data source that is open to entities.
5. The data processing scheduling system based on the unified SQL architecture according to claim 1 is characterized in that: The syntax adaptation of the SQL file includes: a SQL syntax adapter module receives the SQL file transmitted by the scheduling module, and according to the target execution engine environment where the current SQL file is located, performs syntax conversion of the SQL statements in the SQL file according to the target execution engine environment to obtain a syntax-adapted SQL file.
6. The data processing scheduling system based on the unified SQL architecture according to claim 1 is characterized in that: The calling of an external algorithm to process local data includes: establishing a communication protocol between the algorithm request module and an external system, the algorithm request module transmitting the local data to the external system, the external system receiving the local data, calling an external algorithm to perform synchronization request decoupling on the local data to obtain a synchronization request decoupling result, storing the synchronization request decoupling result in a data table, and obtaining the algorithm request result.
7. The data processing scheduling system based on the unified SQL architecture according to claim 1 is characterized in that: The execution result and the algorithm request result are both data tables.
8. A data processing scheduling method based on a unified SQL architecture, characterized in that: The following steps are involved: S1. The user writes a data processing task and deploys the written data processing task to the SQL gateway module or the scheduling module. The data processing task includes multiple SQL files and algorithm requests executed in order; S2. The scheduling module schedules and records the execution of the data processing task, transfers the SQL file to the SQL syntax adapter module, and transfers the algorithm request to the algorithm request module; S3. The SQL syntax adapter module performs syntax adaptation on the SQL file and transmits the syntax-adapted SQL file to the SQL gateway module; S4. The SQL gateway module receives the SQL file, sends the SQL file to the computing engine for execution, obtains the execution result, and outputs the execution result to the storage module; S5. The algorithm request module receives the algorithm request, calls an external algorithm to process local data, obtains an algorithm request result, and outputs the algorithm request result to the storage module.
9. A computer device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform the operation of the data processing scheduling method based on the unified SQL architecture as described in claim 8.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one executable instruction. When the executable instruction is executed on a computer device, the computer device executes the operation of the data processing scheduling method based on the unified SQL architecture as claimed in claim 8.