Data processing method and server
By deploying data integration management tools in the server, obtaining target template information and target parameters, and automatically executing the data integration process, the problem of users needing to be familiar with the functional characteristics of Nifi components is solved, and the effect of simplifying the process, improving efficiency and improving user experience is achieved.
Patent Information
- Application Number
- CN202510128052.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-06-13
AI Technical Summary
In Nifi-based data integration, users need to be familiar with the functional characteristics of rich processor and controller components, resulting in complex data integration processes and reduced efficiency and user experience.
By deploying data integration management tools in the server to obtain target template information and target parameters, users can automatically execute the data integration process based on this information without manually dragging and dropping the processor and controller.
Simplifies the data integration process, improves data processing efficiency and user experience, and reduces the possibility of data errors or loss caused by human error.
Smart Images

Figure CN120146018A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of servers, and in particular, to a data processing method and a server. Background Art
[0002] Data integration is the core link of data processing, which is used to integrate data from different sources (the end that provides data) and in different formats into a unified data platform. As an open-source tool, Nifi has been increasingly widely used in the field of data integration because it can provide a visual interface that enables users to manage data streams based on this visual interface.
[0003] Currently, for data integration based on Nifi, data integration can be achieved by dragging and dropping components such as processors and controllers on the visual interface without writing complex code. However, Nifi has a rich set of built-in processors and controllers, etc. Users must be familiar with the functional characteristics of these components before dragging and dropping. Otherwise, selecting the wrong components may lead to improper data stream processing and affect the sequential progress of data integration. In addition, the implementation of a data integration often requires dragging and dropping a large number of components such as processors or controllers, resulting in a complex data integration process, reducing the efficiency of data integration, and affecting the user experience. Summary of the Invention
[0004] Embodiments of this application provide a data processing method and a server, which enable users to perform data integration without being familiar with the functional characteristics of components such as processors and controllers built in Nifi. Therefore, it helps to simplify the data integration process, improve the data integration efficiency, and enhance the user experience.
[0005] In a first aspect, embodiments of this application provide a data processing method, which is applied to a first server. The first server may be a single server or a server cluster. In a possible implementation, the first server deploys a data integration management tool, such as Nifi, and Nifi executes this data processing method.
[0006] Specifically, the first server obtains the target template information and the target parameters corresponding to the target template information. Among them, the target template information corresponds to the target data type, indicating the target template corresponding to the target data type, and the target template is used to perform data integration of the target data type. The target parameters include connection information and file location. The connection information is used to obtain data from the source end and correctly transfer the obtained data to the destination end. The file location is used to indicate the file path configuration information between the source end and the destination end. The first server obtains the target template according to the target template information, and based on the target parameters and the target template, performs data integration. Thus, the user does not need to be familiar with the functional characteristics of components such as processors and controllers. Only by understanding the target template information corresponding to the target data type and the target parameters corresponding to the target template information and notifying the first server, the first server can automatically execute the data processing process based on the target template information and the target parameters. This method eliminates the step of the user manually dragging multiple processors and controllers, greatly simplifies the process assembly, and improves the data processing efficiency and usability.
[0007] In a specific implementation, if the target template includes data integration processing steps, corresponding processor information, and controller information, configure the communication with the source end and the destination end according to the connection information; and configure to obtain the target data from the source end and write the target data after being processed by the processor to the specified location of the destination end according to the file location; activate the corresponding processor and controller in sequence according to the data integration processing steps, processor information, and controller information in the target template, and the activated processor and controller process the data according to the data integration processing steps; use the activated processor and controller to perform data integration, and the data integration includes: obtaining the target data from the source end, and performing integration processing on the target data, and writing the processed target data to the specified location of the destination end. Through the clear data integration processing steps, processor information, and controller information, it is ensured that each step of data processing is carried out according to the predetermined plan, thereby reducing the possibility of data errors or losses caused by human errors.
[0008] In another specific implementation, the target parameters further include task scheduling parameters. According to the task scheduling parameters, configure the execution time interval or trigger condition; based on the execution time interval or trigger condition, use the activated processor and controller to perform data integration. Thus, by setting the task scheduling period, the first server can perform the data integration operation at the time point considered appropriate according to the user's needs.
[0009] Among them, the target template information includes one of the process templates from the transport protocol to the file format, the process template from the file format to the message queue, the process template from the first relational database to the second relational database, the process template from the file format to the relational database, and the process template from the relational database to the message queue.
[0010] Among them, the process template from the transmission protocol to the file format indicates the entire data integration steps required to obtain data through the transmission protocol and output the data in the file format, as well as the corresponding processor information and controller information; the process template from the file format to the message queue indicates the entire data integration steps required to obtain the data in the file format and output the data through the message queue, as well as the corresponding processor information and controller information; the process template from the first relational database to the second relational database indicates the entire data integration steps required to obtain data from the first relational database and output the data to the second relational database, as well as the corresponding processor information and controller information; the process template from the file format to the relational database indicates the entire data integration steps required to obtain the data in the file format and output the data to the relational database, as well as the corresponding processor information and controller information; the process template from the relational database to the message queue indicates the entire data integration steps required to obtain data from the relational database and output the data through the message queue, as well as the corresponding processor information and controller information.
[0011] Among them, if the target data type includes a relational database, the processor information corresponding to the target template includes first processor information, and the first processor corresponding to the first processor information is used to execute a structured query statement;
[0012] If the target data type includes semi-structured data, the processor information corresponding to the target template includes second processor information, and the second processor corresponding to the second processor information is used to execute: parsing the semi-structured data, whether to set the default data delimiter and whether to include a header;
[0013] If the target data type is interface service data, the processor information corresponding to the target template includes third processor information, and the third processor corresponding to the third processor information is used to process interface requests and set the request timeout;
[0014] If the target data type includes message queue type data, the processor information corresponding to the target template includes fourth processor information, and the fourth processor corresponding to the fourth processor information is used to execute at least one of selecting a consumer client or a producer client, setting corresponding parameters according to the connection address, setting the default maximum request size and compression type parameters, according to the connection information binding information, whether to allow batch data, and setting the maximum message size per batch; among them, the consumer client is used to read data from the message queue, and the producer client is used to send data to the message queue.
[0015] Among them, if the processor information corresponding to the target template includes the first processor, the corresponding controller information includes data connection pool controller information; if the processor information corresponding to the target template includes the second processor, the corresponding controller information includes dedicated data reading controller information; if the processor information corresponding to the target template includes the third processor, the corresponding controller information includes network controller information dedicated to processing network requests and responses; if the processor information corresponding to the target template includes the fourth processor, the corresponding controller information includes consumer client / producer client controller information.
[0016] Among them, the target template includes corresponding processor information, and the corresponding processor information includes processor tuning parameters and performance parameters; the processor tuning parameters include the maximum number of records, the maximum value of each data volume file, and / or the default maximum concurrency parameter; the performance parameters include the queue data length, data size, or running time in the channel.
[0017] Among them, for the processor, when reading a large amount of data, if there is no limit on the data volume size, an extremely large file stream will be formed, eventually causing insufficient memory and leading to processor crashes. Therefore, for database query processing with a relatively high number of uses, the maximum number of records can be set to avoid the problem of insufficient memory of the processor. Configuring concurrency, i.e., load, can improve the performance of the processor. In the process template, by setting the default maximum concurrency parameter, it is possible to prevent multiple threads from causing an excessive thread switching burden on the system and not accelerating the data processing. By setting the queue data length and size in the channel, based on the backpressure mechanism, it is possible to prevent the system from crashing due to excessive data volume. By setting the maximum data stream and maximum data volume allowed to accumulate, it is possible to prevent the system from accumulating a large amount of data volume and occupying a large amount of resources. By setting the running processing time for a processor whose speed of processing a single task is greater than the preset speed threshold and the data volume of the data stream is greater than the preset data volume threshold, the processing speed can be improved.
[0018] Among them, according to the target template information, the first server receives the target template sent by the second server; the second server is used to encapsulate the target template. By encapsulating the target template through the second server, the first server executes the data integration task. Thus, it is possible to process the encapsulation of the target template and the execution of the data integration in parallel, which can improve the processing efficiency, avoid resource contention and bottlenecks, and improve resource utilization.
[0019] In a second aspect, an embodiment of the present application provides a data processing method applied to a display device. The method includes:
[0020] Displays a display page. The display page includes multiple template information, and the template information corresponds to data types, indicating a process template corresponding to the data type. The process template performs data integration of the data type; receives a trigger operation on the display page, obtains target template information, and updates the display content of the display page; the updated display page is used to input target parameters of the target template corresponding to the target template information; the target parameters include connection information and file location. The connection information is used to obtain data from the source end and correctly transfer the obtained data to the destination end. The file location is used to indicate the file path configuration information between the source end and the destination end; sends the target template information and the target parameters to perform data integration according to the target template information and the target parameters. That is, the user only needs to select the target template information and the target parameters through the display interface to execute the data integration operation. The operation is simple, and the user does not need to be familiar with the functional characteristics of the processor and the controller. The user only needs to understand the target template information corresponding to the target data type and the target parameters corresponding to the target template information, and the server can automatically execute the data processing process based on the target template information and the target parameters. This method eliminates the step of the user manually dragging multiple processors and controllers, greatly simplifies the process assembly, and improves the data processing efficiency and usability.
[0021] In a specific implementation, the display page includes: process template information corresponding to transmission protocol data, process template information corresponding to relational database data, process template information corresponding to non-relational databases, process template information corresponding to message queue data, process template information corresponding to file monitoring data, and process template information corresponding to interface service data.
[0022] In another specific implementation, the display device can communicate with the server through a load balancer. Among them, the load balancer is used to: receive the target template information and the target parameters sent by the display device, and send the target template information and the target parameters to the server.
[0023] In yet another specific implementation, the load balancer includes a primary load balancer and a standby load balancer; if the primary load balancer is working properly, the primary load balancer is used to send the target template and the target parameters to the server, and the standby load balancer is in a standby state; if the primary load balancer is abnormal, the primary load balancer stops working, and the standby load balancer is used to send the target template information and the target parameters to the server.
[0024] In a third aspect, an embodiment of the present application provides a data processing device, which is applied to a first server. The device includes an acquisition unit for: acquiring target template information and target parameters corresponding to the target template information; wherein, the target template information corresponds to a target data type, indicating a target template corresponding to the target data type, and the target template is used to perform data integration of the target data type; the target parameters include connection information and file location, the connection information is used to obtain data from a source end and correctly transfer the obtained data to a destination end, and the file location is used to indicate file path configuration information between the source end and the destination end. The device further includes a processing unit for: acquiring a target template according to the target template information, and performing data integration based on the target parameters and the target template.
[0025] In a specific implementation, if the target template includes data integration processing steps, corresponding processor information, and controller information, the processing unit is specifically configured to: configure communication with the source end and the destination end according to the connection information; and configure to obtain target data from the source end according to the file location and write the processed target data to a specified location at the destination end; activate corresponding processors and controllers in sequence according to the data integration processing steps, processor information, and controller information of the target template, and the activated processors and controllers process data according to the data integration processing steps; perform data integration by using the activated processors and controllers, and the data integration includes: obtaining target data from the source end and performing integration processing on the target data, and writing the processed target data to a specified location at the destination end.
[0026] In another specific implementation, the target parameters further include task scheduling parameters, and the processing unit is specifically configured to: perform data integration by using the activated processors and controllers based on an execution time interval or a trigger condition.
[0027] In a fourth aspect, an embodiment of the present application provides a server, including:
[0028] a memory for storing programs;
[0029] a processor for executing the programs stored in the memory, and when the programs stored in the memory are executed, the processor is used to execute the method according to any one of the first aspect.
[0030] In a fifth aspect, the present application provides a computer storage medium for storing a computer program, and when the computer program is executed, it is used to implement the method provided by any one of the implementation manners in the first aspect of the present application.
[0031] In a sixth aspect, the present application provides a computer program product containing instructions, and when it runs on at least one computing device, it enables at least one computing device to implement the method provided by any one of the implementation manners in the first aspect of the present application.
[0032] Any of the data processing methods provided above, corresponding computing devices, computer-readable storage media, computer program products, etc. are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods and will not be elaborated here. Description of the Drawings
[0033] Figure 1 It is a diagram of the application scenario of a data processing method provided by an embodiment of the present application;
[0034] Figure 2 It is a diagram of the application scenario of another data processing method provided by an embodiment of the present application;
[0035] Figure 3 It is a flowchart of a data processing method provided by an embodiment of the present application;
[0036] Figure 4 It is a schematic diagram of the display of a first page provided by an embodiment of the present application;
[0037] Figure 5A It is a schematic diagram of the display of another first page provided by an embodiment of the present application;
[0038] Figure 5B It is a schematic diagram of the display of yet another first page provided by an embodiment of the present application;
[0039] Figure 6 It is a schematic diagram of the display of a second page provided by an embodiment of the present application;
[0040] Figure 7 It is a diagram of the implementation scenario of another data processing method provided by an embodiment of the present application;
[0041] Figure 8 It is an interaction diagram of the implementation of a data processing method provided by an embodiment of the present application;
[0042] Figure 9 It is a schematic diagram of another method for encapsulating a process template provided by an embodiment of the present application;
[0043] Figure 10 It is a schematic diagram of another encapsulation of a process template provided by an embodiment of the present application;
[0044] Figure 11 It is a schematic diagram of the structure of a data processing device provided by an embodiment of the present application. Detailed Description of the Embodiment
[0045] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0046] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0047] The terms "first", "second", etc. in the description and claims of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first target object and the second target object are used to distinguish different target objects, rather than to describe a specific order of the target objects.
[0048] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0049] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of systems refers to two or more systems.
[0050] First, the professional terms related to the embodiments of the present application will be introduced.
[0051] Relational database: It is a database based on the relational model. The data in the database is stored in tables in the form of rows and columns, and the tables are associated through keys. Relational databases support operations for adding, deleting, modifying, and querying data based on the Structured Query Language. Common relational databases include MySQL, Oracle, PostgreSQL, etc.
[0052] Non-relational database: It is a database that does not rely on the traditional relational model and is suitable for processing large-scale distributed data. Non-relational databases include key-value stores, document stores, column stores, or graph databases, etc. Common non-relational databases include MongoDB.
[0053] Message Queue: It is an asynchronous communication method for sending and receiving messages between applications through a queue. Message queues can decouple services and improve the reliability and scalability of the system. Common message queues include Kafka, Pulsar, etc.
[0054] Interface Service: It is a predefined set of functions, protocols, and tools used to define the communication of software components. Interface services support the exchange of data and functions between different systems or applications. Interface services can be provided locally or over a network (usually HTTP / HTTPS) to enable developers to access and use external systems or services.
[0055] Structured Data: It refers to data with a fixed schema and format, usually stored in a tabular form, containing rows and columns. Each row represents a record, and each column represents a field or attribute. Each field has a specific data type (such as integer, string, date, etc.), and the structure of the data is predefined. Structured data is usually stored in relational databases. Common data formats for structured data include CSV, SQL databases, and Excel spreadsheets.
[0056] Unstructured Data: It refers to data without a fixed schema or structure. Common types of unstructured data include text files, images, audio, and video.
[0057] Semi-Structured Data: Semi-structured data is a data format that lies between structured data (such as relational databases) and unstructured data (such as text files). It has a certain structure but does not strictly follow a schema like structured data. Common semi-structured data formats include log files, XML files, HTML documents, JSON files, and YAML files, etc.
[0058] The data processing method provided by the embodiments of this application will be described below.
[0059] First, the application scenario of the data processing method provided by the embodiments of this application will be introduced.
[0060] The data processing method provided by the embodiments of this application is applied on a server, specifically on a data integration management tool deployed on the server, such as Nifi.
[0061] Among them, the data integration management tool is an open-source data flow processing and distribution tool that allows users to create and manage data flows from different source ends to destination ends. Among them, the source end refers to the starting point of the data flow, and the source end can be a file system, database, message queue, etc. that generate data, such as systems, services, or storage that output data. The destination end refers to the end point of the data flow, and the destination end can be a file system, data warehouse, message queue, or database, etc. that receives data.
[0062] The data integration management tool provides an intuitive visual interface. Users can quickly build data integration processes by dragging and dropping processors and controllers without writing complex code. This makes the data integration process more user-friendly. In addition, users can view the integration status of the data stream in real time based on the visual interface to ensure the accuracy and reliability of data transmission.
[0063] Appendix Figure 1 FIG. 100 is an application scenario diagram of a data processing method provided by an embodiment of the present application. In this application scenario 100, it includes a display device 101 and a server 102. Communication is carried out between the display device 101 and the server 102.
[0064] The display device 101 is used to provide a visual interface. In one example, the visual interface can be a web page. Through this visual interface, users can select the pre-packaged process template and configure the target parameters required for data integration. Among them, the pre-packaged process template refers to the entire data processing process required to integrate the source-side data to the destination-side, as well as the required processor information and controller information. The process template selected by the user becomes the target template.
[0065] It can be understood that after the user selects the process template, the entire data integration process required to integrate the source-side data to the destination-side has been determined, and the server 102 can directly execute the data integration task based on this process template without additional configuration of the data processing process and configuration of processor and controller information.
[0066] In addition, users can configure target parameters through the visual interface. For example, configure source-side or destination-side connection information, file location, task scheduling period, etc. Thus, the server 102 can correctly obtain data from the source-side and correctly transfer the obtained data to the destination-side through the target parameters.
[0067] Among them, the display device 101 sends the target template information and target parameters to the server 102. In one example, after the display device 101 obtains the target template information, it sends the target template information to the server 102, and after the display device 101 obtains the target parameters, it sends the target parameters to the server 102. In another example, the display device 101 can send the target template information and target parameters to the server 102 after obtaining the target template information and target parameters. The specific sending method is not specifically limited in the embodiments of the present application.
[0068] After receiving the target template information and target parameters, server 102 obtains the target template according to the target template information, and executes the data integration process based on the target template and the target parameters. Specifically, based on the source file path in the target parameters, data is obtained from the source end. Based on the data integration process of the target template, the data is processed such as cleaning abnormal data and converting data formats. The processed data is output to the destination end based on the target file path.
[0069] Furthermore, server 102 can be a server cluster. Figure 2 This is another schematic diagram of an application scenario provided by an embodiment of the present application. As Figure 2 shown, the server cluster 201 includes multiple computing nodes. The display device 101 sends the target template and the target parameters to the server cluster 201.
[0070] In the embodiment of the present application, the display device 101 can send the target template information and the target parameters to all available servers in the server cluster 201, or send them to some available servers, or send them to specified available servers. The embodiment of the present application does not specifically limit this. The available servers in the available server cluster execute the data integration process according to the target template and the target parameters.
[0071] Furthermore, as Figure 2 shown, a load balancer 202 can also be set between the display device 101 and the server cluster 201. Among them, the display device 101 communicates with the computing nodes in the server cluster 201 through the load balancer 202.
[0072] The load balancer 202 provided by the embodiment of the present application is a device or software used to distribute network traffic among multiple computing nodes. Its main purpose is to optimize resource utilization, maximize throughput, and reduce response time, etc. In the embodiment of the present application, the load balancer 202 is used to obtain the information sent by the display device 101 and send the information to the computing nodes. Among them, the information sent by the display device 101 includes the target template and the target parameters.
[0073] Among them, the load balancer 202 includes at least two load balancers, one is the main load balancer, and the others are standby load balancers. Under normal circumstances, the main load balancer sends the target template and the target parameters. In addition, the main load balancer will regularly perform health checks on the computing nodes to ensure that the computing nodes respond normally. If a failure occurs in a computing node, the main load balancer will automatically remove the failed computing node from the computing node cluster.
[0074] The standby load balancer is usually in a standby state and does not process any traffic. It monitors the status of the primary load balancer. When the primary load balancer fails (such as hardware failure, network interruption, software crash, etc.), the standby load balancer will automatically take over all traffic to ensure that the service will not be interrupted. This process is called "failover". Specifically, the primary load balancer and the standby load balancer usually communicate through heartbeat detection. If the standby load balancer does not receive the heartbeat signal from the primary load balancer within a certain period of time, it is assumed that the primary load balancer has failed and immediately takes over the traffic.
[0075] In the embodiment of the present application, by setting up the architecture of the primary and standby load balancers between the display device 101 and the server cluster 201, the reliability of data transmission and the fault tolerance of the system can be guaranteed.
[0076] In the embodiment of the present application, the display device displays multiple packaged process templates. These templates have integrated all data processing steps, processor, and controller information required from the source end to the destination end. Thus, the user does not need to be familiar with the functional characteristics of the processor and the controller. By selecting the process template and inputting the target parameters, the server or the server cluster can automatically execute the data processing process. Moreover, this method eliminates the steps of the user manually dragging multiple processors and controllers, greatly simplifies the process assembly, and improves the data processing efficiency and usability.
[0077] Appendix Figure 3 The present application embodiment provides a flowchart of a data processing method, and the method includes the following steps:
[0078] S310. The display device 101 displays the first page.
[0079] The first page includes process template information corresponding to multiple data types that have been packaged.
[0080] In the embodiment of the present application, the data of multiple data types can be data divided according to the data structure, including structured data, semi-structured data, and unstructured data, or data divided according to the data source, including relational database data, non-relational database data, message queue data, and interface service data, etc.
[0081] In addition, the data of multiple data types can be data divided according to technical characteristics, including transmission protocol data, file format data, database data, message queue data, file monitoring data, and interface service data.
[0082] Further, the transmission protocol data can be further subdivided into Transmission Control Protocol (TCP) data, File Transfer Protocol (FTP) data, and so on. The file format data can be further subdivided into data such as text (txt), JavaScript Object Notation (json), file, and table (excel). The database data includes relational database data and non-relational database data. Among them, the relational database data is further subdivided into data such as MySQL, Oracle, and PostgreSQL, and the non-relational database is further subdivided into data such as ClickHouse and DM. The message queue data can be further subdivided into Pulsar and Kafka data. The file monitoring data is specifically Tailfile data. The interface service data includes HyperText Transfer Protocol (HTTP / HTTPS) data.
[0083] The data of multiple data types can also be divided in other ways, and the embodiments of the present application do not specifically limit this.
[0084] It should be noted that the transmission protocol data such as TCP data and FTP data refers to the data obtained or output through the transmission protocol. The database data refers to the data obtained or written from the database. The file format data refers to the data obtained or written from the file system. The message queue data refers to the data obtained or sent from the message queue. The interface service data refers to the data obtained or sent from the interface service.
[0085] Exemplarily, as Figure 4 shown, it is a schematic diagram of the display of a first page provided by the embodiments of the present application. The first page 400 includes various pre-packaged process template information. Figure 4 It shows the process templates from the transmission protocol to the file format, from the file format to the message queue, from the relational database to the relational database, from the file format to the relational database, from the relational database to the message queue, and so on.
[0086] It should be noted that the process template from A to B (i.e., A - B) refers to the entire data integration steps required to obtain data through A and output the data in B, as well as the corresponding processor information and controller. For example, the process template from the transmission protocol to the file format refers to the entire data integration steps required to obtain data through the transmission protocol and output the data in the file format, as well as the corresponding processor information and controller information.
[0087] The user can select a target template from the first page 400 as needed. For example, as Figure 4 shown, each process template includes two operation items, namely, a "Delete" item and a "Download" item. When the user clicks the "Download" item below the process template, it indicates that the user determines the target template to be this process template on the first page 400. For example, if the user clicks the "Download" item corresponding to the process template from a relational database to a relational database, the target template is the process template from a relational database to a relational database.
[0088] It should be noted that if the user clicks the "Delete" selection item, it is used to perform a deletion operation on this process template. After deletion, this process template will no longer be displayed on the first page 400. Thus, the embodiments of the present application can select or delete process templates as needed.
[0089] It should be noted that for the two selection items of "Delete" and "Download" described in the embodiments of the present application, the way that the user triggers the download to select the target template is only for illustrative purposes. In actual use, the triggering method can be adjusted according to needs, and the embodiments of the present application do not specifically limit it.
[0090] Furthermore, the first page can further display the refined process templates, specifically as Figure 5A shown. Figure 5A This is another display schematic diagram of the first page provided by the embodiments of the present application. The first page 400 specifically includes: process templates such as file-pulsar, tcp-pulsar, tailfile-kafka, pulsar-pulsar, mysql-clickhouse, mysql-ftp, mysql-postgresql, and excel-clickhouse, etc.
[0091] Among them, the process template of file-pulsar refers to the entire data integration steps required to obtain data in the file format named file and output the data in the message queue data format named pulsar, as well as the corresponding processor information and controller. The process template of tcp-pulsar refers to the entire data integration steps required to obtain data through the Transmission Control Protocol named tcp and output the data in the message queue data format named pulsar, as well as the corresponding processor information and controller. The process template of tailfile-kafka refers to the entire data integration steps required to obtain data by monitoring data in the file named Tailfile and output the data in the message queue data format named kafka, as well as the corresponding processor information and controller. The process template of pulsar-pulsar refers to the entire data integration steps required to obtain data in the message queue data format named pulsar and output the data in the message queue data format named pulsar, as well as the corresponding processor information and controller. The process template of mysql-clickhouse refers to the entire data integration steps required to obtain data from the database named mysql and write the data into the non-relational database named clickhouse, as well as the corresponding processor information and controller. The process template of mysql-ftp refers to the entire data integration steps required to obtain data from the database named mysql and output the data through FTP, as well as the corresponding processor information and controller.
[0092] Exemplarily, as Figure 5A shown, each process template also includes a "Delete" item and a "Download" item. Users can perform delete or confirmation operations through the above two operation items.
[0093] Furthermore, each process template on the first page can also display the creation time and the classification of the process template. Exemplary illustration: as Figure 5A shown, the process templates include two categories. One is the general template. After the user selects the general template, a well-packaged data integration process can be directly formed. The other is the custom template, where the custom template is a data integration process customized by the user according to needs.
[0094] Among them, Figures 4 to 5A shows a schematic diagram of the process templates tiled on the page. In addition, the first page can also display the process templates in other ways.
[0095] For example, as Figure 5BIt is another schematic diagram of the display of the first page provided by the embodiment of the present application. The first page 100 displays a template type drop-down menu and a template name drop-down menu. The user can select a template type through the template type drop-down menu, and display the template names of the template processes corresponding to the template type according to the template name drop-down menu. The user can select a target template through the template name drop-down menu.
[0096] Example description, such as Figure 6 As shown, the user can select a general template through the template type drop-down menu, and select a target template through the template name drop-down menu. For example, the target template is a process template such as mysql-postgresql.
[0097] In addition, the first page can also display the process template in other ways, which is not specifically limited in the embodiment of the present application.
[0098] S320. The display device 101 responds to the trigger operation on the first page and obtains the target template information.
[0099] For example, when the user clicks the "Download" selection corresponding to the process template from the transfer protocol to the file format on the first page, the target template information that the display device 101 can obtain is the process template from the transfer protocol to the file format.
[0100] S330. The display device 101 displays the second page.
[0101] In the embodiment of the present application, after the display device 101 determines the target template information, it updates the display content of the first page (the updated first page is called the second page). The second page is used to configure the target parameters.
[0102] The target parameters are used to make the data integration corresponding to the data type run correctly, that is, to correctly obtain data from the source end and correctly transfer it to the destination end.
[0103] In the embodiment of the present application, the target parameters include the following: connection information, file location, task scheduling period, etc.
[0104] Among them, the connection information refers to the necessary configuration for the data integration management tool to establish communication with the source end and the destination end, and is used to ensure that data can be stably transmitted between the source end, the data integration management tool and the destination end. In the embodiment of the present application, the connection information includes but is not limited to the following: the IP address or domain name of the source end, the IP address or domain name of the destination end, the source network port or the destination network port, the user name, the password, the database name, etc.
[0105] The file location refers to the file path configuration information between the source end and the destination end, which is used to ensure that data files can be correctly read and written. In the embodiments of the present application, the file location includes but is not limited to the following: source file path, target file path, file naming rule, file format, file encoding, etc. Among them, the source file path refers to the storage path of the source end file, which can be a local file system path, a network shared path, or a cloud storage path. The target file path refers to the storage path of the destination end file, which can also be a local file system path, a network shared path, or a cloud storage path.
[0106] The task scheduling period refers to the time interval or trigger condition for the execution of data integration tasks. Among them, the scheduling period of the task can be a fixed period or an event-based trigger, such as a database update trigger, etc. The task scheduling period is used to confirm that the data integration is executed at an appropriate time point.
[0107] For example: Attached Figure 6 FIG. is a schematic diagram of the display of a second page provided by an embodiment of the present application. Among them, the second page 500 includes target parameter items, and the target parameter items specifically include connection information items, file location items, and task scheduling period items. Among them, the connection information items include the source end IP address item, the destination end IP address item, the source end network port item, and the destination end network port item. The file location items include the source file path item and the target file path item.
[0108] It should be noted that the above second page is only a schematic representation. In actual use, the display form and display content can be adjusted according to needs.
[0109] S340. The display device 101 obtains target parameters in response to a configuration operation on the second page.
[0110] The configuration operation is used to configure target parameters.
[0111] For example, the user can Figure 6 input corresponding connection information such as the source end IP address, the destination end IP address, the source end network port, and the destination end network port in the connection information item, input the source file path and the target file path in the file location item, and input the task scheduling period in the task scheduling period item.
[0112] S350. The display device 101 sends the target parameters and the target template information to the server 102.
[0113] Among them, the display device 101 can send the target parameters and the target template to the server 102 at the same time, or can send the target template to the server 102 after step S320, and send the target parameters to the server 102 after step S340.
[0114] Exemplarily, such asFigure 2 As shown, the display device 101 can send the target parameters and target parameters to the server cluster 201 through the load balancer 202.
[0115] S360. The server 102 executes a data integration task according to the target parameters and the target template.
[0116] Specifically, the server 102 (or the server cluster 201) obtains data integration processing steps, processor information, and controller information according to the target template. The server 102 initializes the corresponding processor components and controller components according to the configuration in the target template. The server 102 uses the initialized processor components and controller components to perform operations such as data acquisition, cleaning, transformation, and output.
[0117] Meanwhile, the server 102 can also configure the communication with the source end and the destination end according to the connection information in the target parameters to ensure the successful connection of the server 102 with the source end and the destination end systems. Then the server 102 correctly reads the target data from the source end according to the file location and writes the processed target data to the specified location at the destination end. Meanwhile, the server 102 sets the execution time interval or trigger condition of the data integration task based on the task scheduling parameters.
[0118] Then, the server 102 activates each processor and controller in sequence according to the configuration in the target template and starts to execute the data integration process. Each processor and controller processes the data in a predetermined order to ensure the smooth transfer of data from the source end to the destination end.
[0119] (Optionally) S370. The server 102 outputs the task processing process information to the display device 101.
[0120] Exemplarily, the server 102 can send the data integration task processing process information to the display device 101 through the load balancer 202.
[0121] (Optionally) S380. The display device 101 displays the task processing process through the third page.
[0122] In the embodiment of the present application, the display device 101 will display the third page, and the third page will display the current status of each processor, the status of the flow file, and the connection status. Displaying the task processing process through the third page helps the user to control the data integration task processing process and improve the task processing effect.
[0123] The current status of the processor includes but is not limited to: running status, processing progress, and error information. Among them, the processing progress includes the number of completed tasks, the number of ongoing tasks, and the number of failed tasks.
[0124] The stream file status is used to help users understand the flow of data in the process, including but not limited to: information such as the transmission path, processing time, and queue length of the stream file.
[0125] The connection status refers to the connection status between various processors, including but not limited to: information such as the number of stream files in the queue, the filling rate of the queue, and the maximum capacity of the queue.
[0126] Further, in the embodiment of the present application, the third page further includes a start task scheduling button and a stop task scheduling button. The user can trigger the start task scheduling button or the stop task scheduling button on the third page. The server 102 obtains the trigger information sent by the display device 101 and performs an operation to start the data integration task or stop the data integration task.
[0127] In summary, the data processing method provided by the embodiment of the present application displays multiple encapsulated process templates through the display device. These templates have integrated all the data processing steps, processor, and controller information required from the source end to the destination end. Thus, the user does not need to be familiar with the functional characteristics of the processor and the controller, and only needs to select a process template and input target parameters to enable the server (or server cluster) to automatically execute the data integration process. This method eliminates the step of the user manually dragging multiple processors and controllers, greatly simplifies the process assembly, and improves the data processing efficiency and usability.
[0128] The following is an introduction to a data processing method in conjunction with the accompanying drawings. This data processing method can be used to perform operations such as Figure 3 the data processing method shown, and in addition to performing data integration, it can also perform operations to package process templates in parallel.
[0129] Appendix Figure 7 This is an implementation scenario diagram of another data processing method provided by the embodiment of the present application. In this scenario, in addition to the display device 101 (this display device is called the first display device to distinguish it from the display device used for encapsulating process templates), the server cluster 201, and the load balancer 202, it also includes a second display device 702 and an encapsulation server 701.
[0130] Among them, the first display device 101, the server cluster 201, and the load balancer 202 are used to execute the data processing method shown in the appendix Figure 3 to implement data integration. The second display device 702 is for developers, that is, developers interact with the encapsulation server 701 through the second display device 702 to perform the operation of encapsulating process templates.
[0131] Specifically, the developer can send the data type of the process template to be encapsulated to the encapsulation server 701 through the first display device 702. The encapsulation server 701 encapsulates the processor information corresponding to the data type based on the data type. Among them, for the processing parameters in the processor information, the preset values of the tuned parameters can be filled based on the data processing scenario. Then, the corresponding controller information is encapsulated. After encapsulation, the encapsulation server 701 obtains the process template corresponding to the data type based on all the encapsulated processor information, controller information, and connection information. The encapsulation server 701 sends the encapsulated process template to the server cluster, and then sends it to the first display device 101 through the load balancer 202. When the user accesses the encapsulation server 701 through the first display device 101, the encapsulated process template is displayed. When the user selects the process template, the encapsulation server 701 sends the encapsulated process template to the server cluster for processing.
[0132] It should be noted that in the embodiment of the present application, the process of the encapsulation server 701 creating the process template and the process of the user triggering the data integration task by selecting the target template through the first display device are parallel processes. That is, the two can be carried out simultaneously or one can be selected for execution. The embodiment of the present application does not specifically limit this.
[0133] Appendix Figure 8 is an interaction diagram for implementing a data processing method provided by an embodiment of the present application, which is applied to the application scenario shown in Appendix Figure 7 In the application scenario shown, the method includes a data integration method and an encapsulation method. Among them, for the data integration method, please refer to Figure 3 , which will not be elaborated here. Only the specific implementation of encapsulating the process template will be elaborated here, which specifically includes the following steps:
[0134] S810. The second display device obtains the data type of the process template to be encapsulated determined by the developer.
[0135] For example, the display page of the second display device can display data types divided according to the data structure, including structured data, semi-structured data, and unstructured data. It can also be data divided according to the data source, including relational database data, non-relational database data, message queue data, and interface service data, etc. It can also be data divided according to technical characteristics, including transmission protocol data, file format data, database data, message queue data, file monitoring data, and interface service data, etc. The embodiment of the present application does not specifically limit this.
[0136] For the convenience of description, the following will be schematically described by classifying the data into: relational database data, semi-structured data, message queue data, and interface service data.
[0137] According to requirements, developers can select the data type of the template to be encapsulated on this display page. For example, it can be a process template from relational database data to semi-structured data.
[0138] S820. The second display device sends the data type of the encapsulated template to the encapsulation server 701.
[0139] In this way, the encapsulation server 701 can encapsulate the process template corresponding to this data type.
[0140] S830. The second display device obtains the processor information configured by the developer.
[0141] It should be noted that steps S830 and S820 can be carried out simultaneously, or S830 can be executed first and then S820, or S820 can be executed first and then S830. The embodiments of the present application do not specifically limit this.
[0142] S840. The second display device sends the processor information to the encapsulation server 701.
[0143] S850. The encapsulation server 701 encapsulates the processor information.
[0144] The encapsulation server 701 encapsulates the corresponding data receiving processor information for different types of data. Each processor is responsible for obtaining data from a specific source and performing corresponding operations according to the parameters configured by the user.
[0145] Exemplary illustration: If the data is relational database data, the encapsulation server 701 encapsulating the first processor information includes: using the first processor to execute a structured query statement to implement querying the data to be synchronized from the database. Among them, the first processor is used to execute the structured query statement later. For example, the first processor is the ExecuteSQL processor.
[0146] If the data is message queue type data, the encapsulation server 701 encapsulates different processor information according to different message queues. For example, if the message queue is kafka, the encapsulated processor information includes: selecting a consumer client or a producer client, setting corresponding parameters according to the connection address, setting parameters such as the default maximum request size and compression type. If the message queue is pulsar, the processor information encapsulated by the encapsulation server 701 includes: selecting a consumer client / producer client, binding information according to the connection information, whether to allow batch data, the maximum message size per batch, and the compression type (Compression Type) parameter.
[0147] Among them, the consumer client is used to read data from the message queue, and the producer client is used to send data to the message queue.
[0148] If the data is in text format, the encapsulation server 701 encapsulates the second processor information. The second processor information includes: parsing the text format data using the second processor, default data delimiter, parameters such as whether to include table headers, etc.
[0149] If the data is interface service data, the encapsulation server 701 encapsulates the third processor information. The third processor information includes: selecting the third processor, binding corresponding parameters according to the request address and parameters, setting parameters such as request timeout time, etc. In the embodiment of the present application, the third processor is used to process HTTP requests. For example, the third processor is an InvokeHTTP processor.
[0150] It should be noted that the above encapsulation processor information is only for illustrative purposes. In actual use, it can be increased or decreased according to needs, or adjusted. The embodiments of the present application do not specifically limit it.
[0151] S860. The second display device obtains the controller information configured by the developer.
[0152] S870. The second display device sends the controller information to the encapsulation server 701.
[0153] S880. The encapsulation server 701 encapsulates the controller information.
[0154] For different types of processors, corresponding controllers need to be selected.
[0155] Exemplarily, for the first processor, the encapsulation server 701 needs to select a data connection pool controller. The processor corresponding to the message queue needs to select a consumer client / producer client controller. The processor with an authentication request needs to select an authentication controller. The processor for reading data selects a data receiving and processing controller. For the second processor corresponding to semi-structured data, such as the processor corresponding to csv format data, a dedicated data reading controller needs to be selected. For the third processor, the corresponding controller is a network controller dedicated to processing network requests and responses.
[0156] In the embodiment of the present application, encapsulating the controller information specifically means encapsulating the configuration parameters corresponding to the controller. For example, for a relational database, it is necessary to encapsulate the information of the data connection pool controller (such as the DBCPConnectionPool controller), and the specific control information is: encapsulating the data connection pool controller identifier, automatically selecting the database driver name, driver location, and maximum connection according to the database type.
[0157] For example, the message queue encapsulates controller information according to different types of message queues. For example, if the message queue is Pulsar, the encapsulation server 701 will encapsulate the consumer client / producer client controller identifier, set connection information, monitor the number of processes, and parameters such as the maximum number of concurrent connections (Maximum concurrent lookup-requests);
[0158] For another example, for file format data, the encapsulation server 701 uniformly encapsulates the data read / write controller and sets parameters such as the default cache size.
[0159] S890. The second display device obtains the tuning parameters input by the developer.
[0160] S8100. The second display device sends the tuning parameters to the encapsulation server 701.
[0161] S8110. The encapsulation server 701 encapsulates the tuning parameters.
[0162] In the embodiments of the present application, the tuning parameters include processor tuning parameters and performance parameters.
[0163] The tuning parameters of the processor are used to optimize the performance of the processor.
[0164] In one example, when the processor reads a large amount of data, if there is no limit on the data stream size, an extremely large file stream will be formed, eventually causing memory shortage and leading to processor crash. Therefore, for database query processing with a relatively high usage frequency, the maximum number of records can be set. For example, the maximum number of records is set to 10,000, and the maximum value of each data stream file, for example, the maximum value of each data stream file is set to 1M.
[0165] In one example, for the processor in the data integration process, concurrency and load can be configured to improve the processor speed. Specifically, the encapsulation server 701 sets the default concurrency value and load configuration for the processor. For example, the encapsulation server 701 sets the default maximum concurrency parameter for the processor that needs to improve the concurrent processing speed. For example, the default maximum concurrency parameter is 12. In the process template, by setting the default maximum concurrency parameter, it is possible to prevent multi-threading from causing an excessive thread switching burden on the system and not accelerating the data processing.
[0166] In addition, the performance parameters are used to improve data integration performance. The performance parameters include but are not limited to parameters such as the length of queue data, data size, and running time in the channel.
[0167] Among them, by setting the length and size of the queue data in the channel, based on the pressure mechanism, it is possible to prevent the system from crashing due to excessive data volume. For example, for each data volume channel, set the maximum data flow allowed to accumulate, such as the maximum data flow being 100, and set the maximum data volume allowed to accumulate, such as the maximum data volume being 30M, to prevent the system from accumulating a large amount of data volume and occupying a large amount of resources.
[0168] Among them, for a processor whose speed of processing a single task is greater than the preset speed threshold and the data volume of the data flow is greater than the preset data volume threshold, set the running processing time. For example, set the running processing time to 1s. For example, for a processor that frequently updates data, whether setting the running processing time to 1s by default can continuously occupy the thread and reuse the task processing thread to improve the processing speed.
[0169] Among them, reusing the task processing thread means that when the data integration management tool starts, a task thread is allocated to each processor. The processor obtains the data flow file with the highest priority or a batch of data flow files from the active queue of the incoming connection for processing. If the processing of the data flow does not exceed the configured running duration, it will obtain another data volume file or a batch of data volume files from the active queue and continue all operations in the same thread until the running time limit is reached or the active queue is empty. All processed data flow files will be completed in the same thread and promoted to the appropriate relationship.
[0170] S8120. The encapsulation server 701 creates a process template corresponding to the data type based on the encapsulated processor information, controller information, connection information, and tuning parameters.
[0171] To enable those skilled in the art to better understand the process template encapsulation method provided by the embodiments of the present application, the following uses Use Case 1 and Use Case 2 for illustrative description.
[0172] Among them, Use Case 1 is the encapsulation of the data integration process template for semi-structured file types, including obtaining files, processing data formats, and finally outputting to a relational database or outputting to a message queue middleware.
[0173] Appendix Figure 9 This is a schematic diagram of another process template encapsulation method provided by the embodiments of the present application. Among them, Use Case 1 includes the processing processes of two types of template encapsulations. One is the data integration process template encapsulation from text files to message queues. The other is the data integration process template encapsulation from text files to databases.
[0174] Exemplary 1: First, the encapsulation server 701 encapsulates the file reading processor information. Exemplarily, such as Figure 9As shown, if the text file is a txt file, the file reading processor is the getFile-txt processor information for reading txt files. If the text file is a json file, the file reading processor is the getFile-json processor information for reading json files. Package the file path corresponding to the file data (such as files or folders) selected by the developer for synchronization. Based on the file path, the file reading processor reads the file data to be synchronized.
[0175] Package the processor tuning parameters. For example, set the default synchronization batch size, such as setting the synchronization batch size to 10.
[0176] Next, the encapsulation server 701 packages the ConvertRecord processor information. The record conversion processor is used to convert the read text data into a structured record format, such as Avro. Then, the encapsulation server 701 automatically packages the data reading controller and data writing controller information corresponding to the record conversion processor. For example, the data reading controller is AvroReader, and the data writing controller is AvroRecordWriter. The process template ensures that the data can be correctly converted from the text format to the number of message queues through the data reading controller and data writing controller information.
[0177] Finally, the encapsulation server 701 packages the message queue processor and the controller information for the message queue client connection. The message queue processor is used to convert the data in txt format into the data in message queue format.
[0178] As Figure 9 shown, package the message queue processor, such as the PublishPulsar processor. In addition, the encapsulation server 701 provides a connection address input interface. When the user inputs the connection address corresponding to the pulsar cluster through the input interface, the encapsulation server 701 automatically creates and configures the pulsar client connection controller (StandardPulsarClientService). The encapsulation server 701 packages the configured pulsar client connection controller and packages the performance parameters, such as the maximum number of connections and the number of monitoring processes, etc., to ensure the stability and performance of the connection.
[0179] In the embodiment of the present application, the encapsulation server 701 automatically generates the entire data integration process, including the data channels between processors, setting the adjusted performance parameters, etc. At this time, the user only needs to focus on the source and destination information, and other complex configuration and optimization work are automatically generated by the encapsulation server 701.
[0180] Thus, through the above process, users can easily read data from text files, convert the data into Pulsar messages and publish them to the message queue. The system will automatically encapsulate all necessary processors and controllers and optimize performance parameters. Users only need to focus on the information of the source and target ends. This design greatly simplifies the user's operation and improves the efficiency and reliability of data integration.
[0181] Exemplary 2: First, the service area 701 encapsulates the information of the file reading processor. As Figure 9 shown, the information of the file reading processor is the information of the excel file reading processor. And it encapsulates the file path corresponding to the file data (such as files or folders) selected by the user for synchronization. The file reading processor reads the file data to be synchronized based on the file path.
[0182] Encapsulate the processor tuning parameters. For example, set the default synchronization batch size, such as setting the synchronization batch size to 10. Then, the server 701 encapsulates the information of the parsing processor. Among them, the parsing processor is used to parse text files, such as parsing excel. The server 701 also continues to encapsulate parameters such as the default file delimiter. Then, encapsulate the record conversion processor, which is used to process text data and perform splitting (SplitJson) processing through the data splitting processor.
[0183] Furthermore, when encapsulating the splitting processor, it also encapsulates the maximum string cutting length and the processing method for empty strings. For example, the string switching length is 20M, and the processing method for empty strings is to treat them as empty strings.
[0184] The server 701 also continues to encapsulate the information of the json parsing processor and sets the entity attribute values to be parsed, the maximum string parsing length. For example, the maximum string parsing length is 20M.
[0185] The server 701 encapsulates the information of the attribute replacement (ReplaceText) processor, and the attribute replacement processor automatically binds the attributes to be modified. If modifying attributes, process them through the attribute replacement processor. If not, the data passes through the attribute replacement processor, but the attribute replacement processor does not perform any processing.
[0186] The server 701 is also used to encapsulate the writing processor and bind the destination end to the writing processor. Among them, the writing processor obtains the data sent by the attribute replacement processor and writes the data to the destination end. Furthermore, the server 701 also encapsulates the tuning parameters of the writing processor, such as the maximum value of the batch writing parameter, etc.
[0187] In the embodiment of the present application, the encapsulation server 701 automatically generates the entire data integration process, including the data channel between processors, setting the adjusted performance parameters, etc. At this time, the user only needs to focus on the information of the source end and the destination end, and other complex configuration and optimization work are automatically generated by the encapsulation server 701.
[0188] Further, the use case 2 is a process template for relational databases and message type data to relational databases, including reading data, passing through data format conversion and data formatting, and outputting to the relational database.
[0189] Appendix Figure 10 This is another schematic diagram of process template encapsulation provided by the embodiment of the present application. The process template includes the processing flows of two types of template encapsulations, namely Example 3 and Example 4.
[0190] Example 3 is a process template for relational data to relational data: Among them, the encapsulation server 701 encapsulates the data processor for reading the database. For example, when reading the data of a database table, the data processor is the data processor for reading the database table (ExecuteSQL). The encapsulation server 701 provides an input interface for the data source selected by the user. After the user inputs the source file path through the input interface, the encapsulation server 701 automatically determines the required connection controller, such as DBCPConnectionPool.
[0191] The encapsulation server 701 encapsulates the driver type and driver file corresponding to the database type, and sets the default maximum reading limit. For example, the maximum reading limit is 10M, and the maximum number of records for each data stream file is, for example, 1000, to prevent the processor from being abnormal due to excessive data volume.
[0192] The encapsulation server 701 continues to encapsulate the data stream file conversion processor, where the data stream file conversion processor is used to obtain the data read by the data processor for reading the database and convert it into the target format. For example, the data stream file conversion processor is the ConvertAvroToJSON processor, which is used to convert the data in Avro format into Json format.
[0193] The encapsulation server 701 will continue to encapsulate the data splitting processor and the parsing processor, and set the entity attribute values to be parsed, and the maximum string parsing length. For example, the maximum string parsing length is 20M.
[0194] The encapsulation server 701 encapsulates the information of the property replacement (ReplaceText) processor, and the property replacement processor automatically binds the properties to be modified. If the property is modified, it is processed through the property replacement processor. If there is no need to modify, the data passes through the property replacement processor, but the property replacement processor does not perform any processing.
[0195] The encapsulation server 701 is also used to encapsulate the write processor and the destination for binding the write processor. The write processor obtains the data sent by the attribute replacement processor and writes the data to the destination. Further, the encapsulation server 701 also encapsulates the tuning parameters of the write processor, such as the maximum value of the batch write parameter, etc.
[0196] In the embodiment of the present application, the encapsulation server 701 automatically generates the entire data integration process, including the data channel between processors, setting the adjusted performance parameters, etc. At this time, the user only needs to focus on the information of the source and the destination, and other complex configuration and optimization work are automatically generated by the encapsulation server 701.
[0197] Exemplary 4: The process template from the message queue (such as Pulsar or kafka, etc.) to the database. First, encapsulate the message queue client consumer. Then, according to the connection information selected by the user, automatically bind the corresponding parameters such as the message queue name. Encapsulate the client connection controller (StandardPulsarClientService) or the connection address (KafkaBrokers).
[0198] The encapsulation server 701 will continue to encapsulate the data splitting processor and the parsing processor, and set the entity attribute values to be parsed, the maximum string parsing length, for example, the maximum string parsing length is 20M.
[0199] The encapsulation server 701 encapsulates the information of the attribute replacement (ReplaceText) processor. The attribute replacement processor automatically binds the attributes to be modified. If the attributes are to be modified, they are processed by the attribute replacement processor. If there is no need to modify, the data passes through the attribute replacement processor, but the attribute replacement processor does not perform any processing.
[0200] The encapsulation server 701 is also used to encapsulate the write processor and the destination for binding the write processor. The write processor obtains the data sent by the attribute replacement processor and writes the data to the destination. Further, the encapsulation server 701 also encapsulates the tuning parameters of the write processor, such as the maximum value of the batch write parameter, etc.
[0201] In the embodiment of the present application, the encapsulation server 701 automatically generates the entire data integration process, including the data channel between processors, setting the adjusted performance parameters, etc. At this time, the user only needs to focus on the information of the source and the destination, and other complex configuration and optimization work are automatically generated by the encapsulation server 701.
[0202] In summary, the process template encapsulation method provided by the embodiments of the present application. Through this method, users can easily achieve the plug-and-play of data streams, and the data processing has maintainability and scalability. Moreover, by encapsulating the process template, a standardized and modular operation experience is provided for users, thereby significantly improving the deployment efficiency and usability of the data processing process. In addition, the present invention also takes into account the needs of different users, allowing users to customize and adjust the process template according to specific scenarios, ensuring the flexibility and adaptability of the data processing process. Through this encapsulation method, even non-professional users can quickly get started and efficiently manage and utilize the data integration management tool data stream to achieve automated data processing and analysis.
[0203] Furthermore, the embodiments of the present application also provide a data processing device, attached Figure 11 is a schematic structural diagram of the data processing device provided by the embodiments of the present application. The device 1100 includes an acquisition unit 1101 and a processing unit 1102.
[0204] The acquisition unit 1101 is used to: acquire target template information and target parameters corresponding to the target template information; wherein, the target template information corresponds to the target data type, indicating the target template corresponding to the target data type, and the target template is used to perform data integration of the target data type; the target parameters include connection information and file location, the connection information is used to acquire data from the source end and correctly transfer the acquired data to the destination end, and the file location is used to indicate the file path configuration information between the source end and the destination end. It further includes a processing unit 1102, which is used to: acquire the target template according to the target template information, and perform data integration based on the target parameters and the target template.
[0205] In a specific implementation, if the target template includes data integration processing steps, corresponding processor information, and controller information, the processing unit 1102 is specifically used to: configure the communication with the source end and the destination end according to the connection information; and configure to acquire the target data from the source end and write the processed target data to the specified location of the destination end according to the file location; according to the data integration processing steps, processor information, and controller information of the target template, activate the corresponding processor and controller in sequence, and the activated processor and controller process the data according to the data integration processing steps; use the activated processor and controller to perform data integration, and the data integration includes: acquiring the target data from the source end, and performing integration processing on the target data, and writing the processed target data to the specified location of the destination end.
[0206] In another specific implementation, the target parameters further include task scheduling parameters, and the processing unit 1102 is further used to: perform data integration based on the execution time interval or trigger condition by using the activated processor and controller.
[0207] An embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on a computing device, it causes the computing device to execute the above data processing method. An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center including one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the above data processing method.
[0208] The descriptions of the processes or structures corresponding to the above respective drawings each have their own focuses. For parts not detailed in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.
[0209] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data processing method, characterized in that: Applied to a first server, the method comprises: Obtain target template information and target parameters corresponding to the target template information; The target template information corresponds to the target data type, indicating the target template corresponding to the target data type, and the target template is used to perform data integration of the target data type; the target parameters include connection information and file location, the connection information is used to obtain data from the source end and correctly transmit the obtained data to the destination end, and the file location is used to indicate the file path configuration information between the source end and the destination end; The target template is acquired according to the target template information, and data integration is performed based on the target parameters and the target template.
2. The method according to claim 1, characterized in that If the target template includes data integration processing steps, corresponding processor information and controller information, performing data integration based on the target parameters and the target process template includes: According to the connection information, configure the communication with the source end and the destination end; and according to the file location, configure to obtain the target data from the source end and write the processed target data to the specified location of the destination end; According to the data integration processing steps of the target template, the processor information and the controller information, the corresponding processors and controllers are activated in sequence, and the activated processors and controllers process data according to the data integration processing steps; The activated processor and controller are used to perform data integration, and the data integration includes: acquiring target data from the source end, performing integration processing on the target data, and writing the processed target data into a designated location of the destination end.
3. The method according to claim 2, characterized in that The target parameters also include task scheduling parameters, and the method further includes: configuring execution time intervals or triggering adjustments according to the task scheduling parameters; The performing data integration by using the activated processor and controller includes: performing the data integration by using the activated processor and controller based on the execution time interval or the trigger condition.
4. The method according to any one of claims 1 to 3, characterized in that: The target template information includes: One of a process template from a transmission protocol to a file format, a process template from a file format to a message queue, a process template from a first relational database to a second relational database, a process template from a file format to a relational database, and a process template from a relational database to a message queue; Among them, the process template from the transmission protocol to the file format indicates the entire data integration steps required for obtaining data through the transmission protocol and outputting the data through the file format, the corresponding processor information and controller information; the process template from the file format to the message queue indicates the entire data integration steps required for obtaining data in the file format and outputting the data through the message queue, the corresponding processor information and controller information; the process template from the first relational database to the second relational database indicates the entire data integration steps required for obtaining data from the first relational database and outputting the data to the second relational database, the corresponding processor information and controller information; the process template from the file format to the relational database indicates the entire data integration steps required for obtaining data in the file format and outputting the data to the relational database, the corresponding processor information and controller information; the process template from the relational database to the message queue indicates the entire data integration steps required for obtaining data from the relational database and outputting the data through the message queue, the corresponding processor information and controller information.
5. The method according to claim 2, characterized in that: If the target data type includes a relational database, the processor information corresponding to the target template includes first processor information, and the first processor corresponding to the first processor information is used to execute a structured query statement; If the target data type includes semi-structured data, the processor information corresponding to the target template includes second processor information, and the second processor corresponding to the second processor information is used to execute: parsing the semi-structured data, whether to set a default data delimiter, and whether to include a header; If the target data type is interface service data, the processor information corresponding to the target template includes third processor information, and the third processor corresponding to the third processor information is used to process the interface request and set the request timeout period; If the target data type includes message queue type data, the processor information corresponding to the target template includes fourth processor information, and the fourth processor corresponding to the fourth processor information is used to execute selection of a consumer client or a producer client, set corresponding parameters according to the connection address, set the default maximum request size and compression type parameters, bind information according to the connection information, determine whether batch data is allowed, and set at least one of the maximum message size of each batch; wherein the consumer client is used to read data from a message queue, and the producer client is used to send data to a message queue.
6. The method according to claim 5, characterized in that If the processor information corresponding to the target template includes a first processor, the corresponding controller information includes data connection pool controller information; if the processor information corresponding to the target template includes a second processor, the corresponding controller information includes dedicated data reading controller information; if the processor information corresponding to the target template includes a third processor, the corresponding controller information includes network controller information dedicated to processing network requests and responses; if the processor information corresponding to the target template includes a fourth processor, the corresponding controller information includes consumer client / production client controller information.
7. The method according to claim 1, characterized in that The target template includes corresponding processor information, and the corresponding processor information includes processor tuning parameters and performance parameters; the processor tuning parameters include the maximum number of records, the maximum value of each data volume file and / or the default maximum concurrency parameters; the performance parameters include the queue data length, data size or running time in the channel.
8. The method according to claim 1, characterized in that The acquiring the target template according to the target template information includes: According to the target template information, the target template is received from the second server; the second server is used to encapsulate the target template.
9. A data processing method, characterized in that: Applied to a display device, the method comprises: Displaying a display page, the display page includes a plurality of template information, the template information corresponds to a data type, indicates a process template corresponding to the data type, and the process template executes data integration of the data type; receiving a trigger operation on the display page, acquiring target template information, and updating display content of the display page; the updated display page is used to input target parameters of the target template corresponding to the target template information; the target parameters include connection information and file location, the connection information is used to acquire data from the source end and correctly transmit the acquired data to the destination end, and the file location is used to indicate file path configuration information between the source end and the destination end; The target template information and the target parameters are sent to perform data integration according to the target template information and the target parameters.
10. A server, characterized in that: The invention comprises a memory and a processor, wherein the memory and the processor are coupled: The memory is used to store programs; The processor is configured to execute the method according to any one of claims 1 to 8 based on the program stored in the memory.