Method and system for simplifying NiFi data integration task configuration
By introducing the combination of technical implementation layer, business semantic layer and dynamic interface layer into the NiFi data stream processing tool, the isolated mapping of technical parameters and business parameters and dynamic form rendering are achieved, which solves the problem of NiFi configuration complexity, improves the operating experience of business personnel and the efficiency of enterprise-level data integration.
Patent Information
- Application Number
- CN202510782407.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Apache NiFi's data stream processing tool has problems such as high technical barriers, low configuration efficiency, prone to errors, and difficulty in large-scale reuse during the configuration process. In particular, due to the complexity of its native configuration system and the gap in business semantic understanding, it is difficult for business personnel to operate effectively.
By simplifying the configuration of NiFi data integration tasks, the combination of technical implementation layer, business semantic layer and dynamic interface layer is adopted to achieve isolated mapping of technical parameters and business parameters, and dynamic form rendering technology is used to generate a configuration interface for business domains, including developing task synchronization templates in NiFi's visual interface, filling in description fields in JSON string format, defining business semantic models and implementing a semantic two-way conversion engine, and dynamically rendering the interface based on Vue.js.
It solves the problems of high technical threshold and difficult and complex task configuration of NiFi, and makes it user-friendly for business users. It is suitable for rapid deployment and configuration management in enterprise-level data integration scenarios.
Smart Images

Figure CN120631354A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data integration, and in particular to a method and system for simplifying NiFi data integration task configuration. Background Art
[0002] In the heterogeneous data system integration scenario, Apache NiFi, as an excellent data stream processing tool, has demonstrated strong technical advantages with its automated transmission, conversion and processing capabilities.
[0003] While Apache NiFi offers comprehensive data flow pipeline building capabilities, its native configuration system presents significant technical obstacles. The underlying parameter system for developers (such as processor configuration, connection strategy, and backpressure settings) lacks understanding of business semantics. The visual orchestration interface fails to achieve multi-level abstraction, forcing business personnel to navigate hundreds of processors and complex attribute configurations. This results in low configuration efficiency, prone to errors, and difficulty in scalable reuse. Summary of the Invention
[0004] The technical task of the present invention is to address the above shortcomings and provide a method and system for simplifying the configuration of NiFi data integration tasks, which can realize the isolated mapping of technical parameters and business parameters, and combine dynamic form rendering technology to generate a configuration interface for business domains, thereby solving the problems of high NiFi technology threshold, difficult and complex task configuration, and difficulty for business users to understand and use.
[0005] The technical solution adopted by the present invention to solve its technical problem is:
[0006] A method to simplify NiFi data integration task configuration, which includes a technical implementation layer, a business semantic layer, and a dynamic interface layer.
[0007] The technical implementation layer develops a task synchronization template based on NiFi's visual interface, maps business attributes to NiFi's Parameter Context parameters, improves the business semantics of each business attribute, and fills it into the Description field in JSON string format;
[0008] The business semantic layer defines the business semantic model, implements the semantic bidirectional conversion engine, and realizes the automatic conversion between NiFi parameter context and business semantic model;
[0009] The dynamic interface layer is based on Vue.js dynamic rendering interface.
[0010] This approach uses a parameter context decoupling engine to isolate and map technical and business parameters, and combines it with dynamic form rendering technology to generate a business-specific configuration interface. This approach addresses NiFi's high technical barrier to entry, the complexity of task configuration, and the difficulty users face in understanding its use. It is suitable for rapid deployment and configuration management in enterprise-level data integration scenarios.
[0011] Furthermore, the technology implementation layer,
[0012] The technical implementation layer is based on NiFi's visual interface and uses NiFi components, including Processor, ControllerService, Process Group and other components, to develop a general task synchronization template, extract business-related attributes, and map them to parameters in NiFi's Parameter Context component for business semantic layer analysis.
[0013] Furthermore, the technical implementation layer develops a MYSQL to MYSQL synchronization template. The specific implementation process is as follows:
[0014] 1) Create a ProcessGroup named "MYSQL Synchronization MYSQL Template" and create and reference a ParameterContext named "MYSQL to MYSQL Synchronization Parameters";
[0015] 2) Create the source MYSQL acquisition component QueryDatabaseTable Processor. Since this component relies on a Controller Service component to establish a data connection, you need to first create a DBCPConnectionPoolController Service named sourceDatabase.
[0016] 3) Create the target MYSQL import component PutDatabaseRecord Processor. Since this component relies on a Controller Service component to establish a data connection, you need to first create a DBCPConnectionPoolController Service named targetDatabase.
[0017] 4) Through steps 1) to 3), the entire synchronization task is mapped to a ParameterContext containing only business attributes;
[0018] 5) The parameters in the Parameter Context only contain four attributes: Name, Value, Sensitive Value, and Description. Other attributes need to be completed according to the business semantic layer model. Use the Description attribute to fill in the required fields in JSON format in the Description, and then parse it through the business semantic layer.
[0019] 6) After the business semantics are supplemented and improved, download the task template through the Download Flow Definition of the Process Group and perform subsequent operations.
[0020] Furthermore, the DBCPConnectionPool Controller Service component named sourceDatabase contains the following key properties:
[0021] Database Connection URL: Database connection URL, which is a business attribute and needs to be mapped to ParameterContext. The mapping parameter name is sourceDatabaseURL, and the component references the mapping parameter #{sourceDatabaseURL};
[0022] Database Driver Class Name: Database driver class name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is sourceDatabaseDriver, and the component references this parameter #{sourceDatabaseDriver};
[0023] Database User: Database user name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is sourceDatabaseUser, and the component references this parameter #{sourceDatabaseUser};
[0024] Password: Database user name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is sourceDatabasePassword, and the component references this parameter #{sourceDatabasePassword};
[0025] Max Wait Time: The maximum waiting time for establishing a connection, which is a technical attribute and can be set to 500 millis;
[0026] Max Total Connections: The maximum number of connections, which is a technical attribute and can be set to 8;
[0027] The QueryDatabaseTable Processor component contains the following key properties:
[0028] Database Connection Pooling Service: Database connection pool service, a ControllerService component, used to configure data source connections. You can select the sourceDatabase established above;
[0029] Database Type: Database type, select MYSQL type, which is the default attribute and does not require mapping;
[0030] Table Name: table name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is sourceTableName, and the component references this parameter #{sourceTableName};
[0031] Columns to Return: Returns the column name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is sourceColumnName, and the component references this parameter #{sourceColumnName};
[0032] Additional WHERE clause: query conditions, which are business attributes and need to be mapped to ParameterContext. The mapping parameter name is sourceWhere, and the component references this parameter #{sourceWhere};
[0033] Custom Query: Query SQL, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter is named sourceSql, and the component references this parameter #{sourceSql};
[0034] Maximum-value Columns: Maximum value columns, used for incremental queries. This property can be ignored here.
[0035] Initial Load Strategy: Initial load strategy, used to define the first query method, starting from the beginning or starting from the maximum value. This property can be ignored here.
[0036] Max Wait Time: The maximum waiting time, which is a technical attribute. It can be configured to 0, which means it will never time out and wait forever.
[0037] Fetch Size: The amount of data to be retrieved from the result set each time. This is a technical attribute and can be configured to 0 to retrieve all data.
[0038] Max Rows Per Flow File: The number of data rows per FlowFile, which is a technical attribute and can be configured to 10,000;
[0039] Output Batch Size: The number of FlowFiles output each time. It is a technical attribute and can be configured to 0 to output all.
[0040] Furthermore, the DBCPConnectionPool Controller Service component named targetDatabase contains the following key properties:
[0041] Database Connection URL: Database connection URL, which is a business attribute and needs to be mapped to ParameterContext. The mapping parameter name is targetDatabaseURL, and the component references this parameter #{targetDatabaseURL};
[0042] Database Driver Class Name: Database driver class name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is targetDatabaseDriver, and the component references this parameter #{targetDatabaseDriver};
[0043] Database User: Database user name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is targetDatabaseUser, and the component references this parameter #{targetDatabaseUser};
[0044] Password: Database user name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter is named targetDatabasePassword, and the component references this parameter #{targetDatabasePassword};
[0045] Max Wait Time: The maximum waiting time for establishing a connection, which is a technical attribute and can be set to 500 millis;
[0046] Max Total Connections: The maximum number of connections, which is a technical attribute and can be set to 8;
[0047] The PutDatabaseRecord Processor component contains the following key properties:
[0048] Database Connection Pooling Service: A ControllerService component that is used to configure data source connections. You can choose to select the targetDatabase created above.
[0049] Record Reader: Record reader, a Controller Service component. The default is AvroReader and does not require mapping.
[0050] Database Type: Database type, select MYSQL type, default attributes, no mapping required;
[0051] Table Name: table name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is targetTableName, and the component references this parameter #{targetTableName};
[0052] Statement Type: Operation type, including INSERT, UPDATE, DELETE, and UPSERT, business attributes, which need to be mapped to the Parameter Context. The mapping parameter name is statementType, and the component references this parameter #{statementType};
[0053] Max Wait Time: The maximum waiting time, which is a technical attribute. If it is set to 0, it will never time out and will wait forever.
[0054] Rollback On Failure: error handling mechanism, a technical attribute that can be configured as false;
[0055] Table Schema Cache Size: The size of the cached table schema, a technical attribute that can be configured as 100;
[0056] Maximum Batch Size: The maximum number of SQL statements executed per database execution. This is a technical attribute and can be configured to 1000.
[0057] Furthermore, the business semantic layer,
[0058] Parsing technology is used to implement the exported NiFi task template, extract the Parameter Context attributes in the template, and convert them into business semantic layer patterns for dynamic interface layer rendering;
[0059] At the same time, the form information filled in by the user in the dynamic interface layer is converted into Parameter Context attributes, which replaces the Parameter Context attributes of the NiFi template, and the task template instance is generated and imported into NiFi for execution.
[0060] Furthermore, the business semantic layer is specifically implemented as follows:
[0061] 1) Task template parsing business semantics: Use JSON parsing tools to parse NiFi task templates, which contain the following attributes: flowContents, externalControllerServices, parameterContexts, flowEncodingVersion, parameterProviders, latest; extract the parameterContexts attribute, and then take the parameters of the first attribute of the parameterContexts attribute as the defined business attribute parameters; traverse and obtain the business semantics defined in the description field of each business attribute in parameters, and render them through the dynamic interface layer;
[0062] 2) Business semantic conversion task template: After the business personnel fill in the business attributes through the form, the parameter value paramValue of the obtained business attribute is filled in the value field of the business parameter corresponding to the template parameters.
[0063] The present invention also claims a system for simplifying the configuration of NiFi data integration tasks, comprising:
[0064] On the technical expert side, based on NiFi's visual interface, develop a task synchronization template, map business attributes to NiFi's Parameter Context parameters, improve the business semantics of each business attribute, and fill it into the Description field in JSON string format;
[0065] On the semantic engine side, define the business semantics model, implement the semantic bidirectional conversion engine, and realize the automatic conversion between NiFi parameter context and business semantic model;
[0066] On the business user side, the interface is dynamically rendered based on Vue.js;
[0067] The system simplifies the NiFi data integration task configuration through the above method.
[0068] The present invention also claims a device for simplifying NiFi data integration task configuration, comprising: at least one memory and at least one processor;
[0069] The at least one memory is configured to store a machine-readable program;
[0070] The at least one processor is configured to call the machine-readable program to implement the above method.
[0071] The present invention also claims protection for a computer-readable medium having computer instructions stored thereon, which implement the above method when executed by a processor.
[0072] Compared with the prior art, the method and system for simplifying NiFi data integration task configuration of the present invention have the following beneficial effects:
[0073] The present invention realizes the isolated mapping of technical parameters and business parameters through the parameter context decoupling engine, and combines the dynamic form rendering technology to generate a configuration interface for the business domain, thereby solving the problems of high technical threshold of NiFi, difficult and complex task configuration, and difficulty for business users to understand and use it. It is suitable for rapid deployment and configuration management in enterprise-level data integration scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 This is a diagram of a system architecture that simplifies NiFi data integration task configuration provided by an embodiment of the present invention;
[0075] Figure 2 This is a schematic diagram illustrating the principle of a method for simplifying NiFi data integration task configuration provided by an embodiment of the present invention;
[0076] Figure 3 This is an example diagram of a page for creating a ProcessGroup provided by an embodiment of the present invention;
[0077] Figure 4 This is a diagram showing properties of a DBCPConnectionPoolController Service component named sourceDatabase provided by an embodiment of the present invention;
[0078] Figure 5 This is a diagram showing properties of the QueryDatabaseTable Processor component provided by an embodiment of the present invention;
[0079] Figure 6This is a diagram showing properties of a DBCPConnectionPoolController Service component named targetDatabase provided by an embodiment of the present invention;
[0080] Figure 7 This is a diagram illustrating properties of the PutDatabaseRecord Processor component provided by an embodiment of the present invention;
[0081] Figure 8 is a diagram illustrating the properties of Parameter Context provided by an embodiment of the present invention;
[0082] Figure 9 This is an example diagram of Parameter Context parameters provided by an embodiment of the present invention;
[0083] Figure 10 This is a diagram of a page for downloading a task template through the Download Flow Definition of a Process Group provided by an embodiment of the present invention;
[0084] Figure 11 This is an example of task template parsing business semantics provided by an embodiment of the present invention Figure I ;
[0085] Figure 12 This is an example of task template parsing business semantics provided by an embodiment of the present invention Figure II ;
[0086] Figure 13 This is a page example diagram of a dynamic interface layer based on Vue.js dynamic rendering interface provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0087] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0088] The embodiment of the present invention provides a method for simplifying NiFi data integration task configuration, such as Figure 2 As shown, the implementation of this method includes a technical implementation layer, a business semantic layer and a dynamic interface layer.
[0089] The technical implementation layer develops a task synchronization template based on NiFi's visual interface, maps business attributes to the parameters of NiFi's Parameter Context, improves the business semantics of each business attribute, and fills it into the Description field in JSON string format.
[0090] The business semantic layer defines the business semantic model, implements the semantic bidirectional conversion engine, and realizes the automatic conversion between NiFi parameter context and business semantic model.
[0091] The dynamic interface layer is based on Vue.js dynamic rendering interface.
[0092] The implementation of this method is described in detail below with reference to the accompanying drawings.
[0093] 1. Technical Implementation Layer (Technical Expert Side)
[0094] The technical implementation layer is based on NiFi's visual interface. NiFi technical experts use NiFi components, including Processor, Controller Service, Process Group, etc., to develop general task synchronization templates, extract business-related attributes, and map them to parameters in NiFi's Parameter Context component for business semantic layer analysis.
[0095] The following takes the MYSQL to MYSQL synchronization template as an example. The NiFi version is 2.2.0. The specific implementation is as follows:
[0096] 1. Create a ProcessGroup named "MYSQL Synchronization MYSQL Template" and create and reference a ParameterContext named "MYSQL to MYSQL Synchronization Parameters"; Figure 3 shown.
[0097] 2. Create the QueryDatabaseTable Processor component for the source MYSQL acquisition. Since this component relies on a Controller Service component to establish a data connection, you need to first create a DBCPConnectionPoolController Service named sourceDatabase.
[0098] like Figure 4 As shown, the DBCPConnectionPool Controller Service component contains the following key properties:
[0099] Database Connection URL: Database connection URL, which is a business attribute and needs to be mapped to ParameterContext. The mapping parameter name is sourceDatabaseURL, and the component references the mapping parameter #{sourceDatabaseURL};
[0100] Database Driver Class Name: Database driver class name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is sourceDatabaseDriver, and the component references this parameter #{sourceDatabaseDriver};
[0101] Database User: Database user name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is sourceDatabaseUser, and the component references this parameter #{sourceDatabaseUser};
[0102] Password: Database user name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is sourceDatabasePassword, and the component references this parameter #{sourceDatabasePassword};
[0103] Max Wait Time: The maximum waiting time for establishing a connection, which is a technical attribute and is set to 500 millis here;
[0104] Max Total Connections: The maximum number of connections, a technical attribute, is set to 8 here.
[0105] like Figure 5 As shown, the QueryDatabaseTable Processor component contains the following key properties:
[0106] Database Connection Pooling Service: Database connection pool service, a ControllerService component, is used to configure the data source connection. Here you can select the sourceDatabase established above;
[0107] Database Type: Database type, select MYSQL type, which is the default attribute and does not require mapping;
[0108] Table Name: table name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is sourceTableName, and the component references this parameter #{sourceTableName};
[0109] Columns to Return: Returns the column name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is sourceColumnName, and the component references this parameter #{sourceColumnName};
[0110] Additional WHERE clause: query conditions, which are business attributes and need to be mapped to ParameterContext. The mapping parameter name is sourceWhere, and the component references this parameter #{sourceWhere};
[0111] Custom Query: Query SQL, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter is named sourceSql, and the component references this parameter #{sourceSql};
[0112] Maximum-value Columns: Maximum value columns, used for incremental queries. This property is ignored here.
[0113] Initial Load Strategy: Initial load strategy, used to define the first query method, starting from the beginning or starting from the maximum value. This property is ignored here.
[0114] Max Wait Time: The maximum waiting time, which is a technical attribute. Here, it is configured as 0, which means that the timeout will never expire and the system will wait forever.
[0115] Fetch Size: The number of data items to be retrieved from the result set each time. This is a technical attribute. Here, it is set to 0, which means all items are retrieved.
[0116] Max Rows Per Flow File: The number of data rows per FlowFile, which is a technical attribute and is configured as 10,000 here;
[0117] Output Batch Size: The number of FlowFiles output each time. It is a technical attribute. Here it is configured as 0, which means all are output.
[0118] 3. Create the target MYSQL import component PutDatabaseRecord Processor. Since this component relies on a Controller Service component to establish a data connection, you need to first create a DBCPConnectionPoolController Service named targetDatabase.
[0119] like Figure 6 As shown, the DBCPConnectionPool Controller Service component contains the following key properties:
[0120] Database Connection URL: Database connection URL, which is a business attribute and needs to be mapped to ParameterContext. The mapping parameter name is targetDatabaseURL, and the component references this parameter #{targetDatabaseURL};
[0121] Database Driver Class Name: Database driver class name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is targetDatabaseDriver, and the component references this parameter #{targetDatabaseDriver};
[0122] Database User: Database user name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is targetDatabaseUser, and the component references this parameter #{targetDatabaseUser};
[0123] Password: Database user name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter is named targetDatabasePassword, and the component references this parameter #{targetDatabasePassword};
[0124] Max Wait Time: The maximum waiting time for establishing a connection, which is a technical attribute and is set to 500 millis;
[0125] Max Total Connections: The maximum number of connections, which is a technical attribute, is set to 8.
[0126] like Figure 7 As shown, the PutDatabaseRecord Processor component contains the following key properties:
[0127] Database Connection Pooling Service: A ControllerService component that is used to configure data source connections. You can choose to select the targetDatabase created above.
[0128] Record Reader: Record reader, a Controller Service component. The default is AvroReader and does not require mapping.
[0129] Database Type: Database type, select MYSQL type, default attributes, no mapping required;
[0130] Table Name: table name, which is a business attribute and needs to be mapped to the Parameter Context. The mapping parameter name is targetTableName, and the component references this parameter #{targetTableName};
[0131] Statement Type: Operation type, including INSERT, UPDATE, DELETE, and UPSERT, business attributes, which need to be mapped to the Parameter Context. The mapping parameter name is statementType, and the component references this parameter #{statementType};
[0132] Max Wait Time: The maximum waiting time, which is a technical attribute. If it is set to 0, it will never time out and will wait forever.
[0133] Rollback On Failure: error handling mechanism, a technical attribute, configured as false;
[0134] Table Schema Cache Size: The size of the cached table schema, a technical attribute, configured to 100;
[0135] Maximum Batch Size: The maximum number of SQL statements executed per database execution. This is a technical attribute and is configured to be 1000.
[0136] 4. Through steps 1 to 3, the entire synchronization task is mapped to a ParameterContext with only business attributes. All attributes such as Figure 8 shown.
[0137] 5. The parameters in the Parameter Context only contain four attributes: Name, Value, Sensitive Value, and Description. Other attributes need to be completed according to the business semantic layer model. Use the Description attribute to fill in the required fields in JSON format in the Description, and parse it through the business semantic layer; for example Figure 9 shown.
[0138] 6. After the business semantics are supplemented and improved, download the task template through the Download Flow Definition of the Process Group and perform subsequent operations. Figure 10 shown.
[0139] 2. Business Semantic Layer (Semantic Engine Side)
[0140] The business semantic layer serves as a bridge between the technical implementation layer and the dynamic interface layer. It defines the business semantic model and implements a bidirectional semantic conversion engine: it can automatically convert NiFi parameter context and business semantic models. It can parse NiFi task templates exported by the technical implementation layer, extract the Parameter Context attributes in the template, and convert them into the business semantic layer model for dynamic interface layer rendering. At the same time, it can also convert the form information filled in by users in the dynamic interface layer into Parameter Context attributes, replace the Parameter Context attributes of the NiFi template, generate task template instances, and import them into NiFi for execution.
[0141] The schema for business semantics is defined as follows:
[0142] coding name Remark id Unique identifier of the form item paramCode Form item encoding paramName Form item name paramType Form item type boolean,string,number,date defaultValue Default value of form items paramsValue Form item parameter value showType Form item display type INPUT,SELECT require Is this field required? true,false description Form item description order Form item order
[0143] The specific implementation is as follows:
[0144] 1. Parsing business semantics in task templates: Use JSON parsing tools to parse NiFi task templates, which contain the following attributes: flowContents, externalControllerServices, parameterContexts, flowEncodingVersion, parameterProviders, and latest. Extract the parameterContexts attribute, and then take the first attribute of the parameterContexts attribute, parameters, as the defined business attribute parameters; traverse and obtain the business semantics defined in the description field of each business attribute in parameters, and render them through the dynamic interface layer. Figure 11 、 Figure 12 shown.
[0145] 2. Business semantic conversion task template: After the business personnel fill in the business attributes through the form, they will fill in the parameter value paramValue of the business attribute into the value field of the corresponding business parameter in the template parameters.
[0146] 3. Dynamic interface layer (business user side)
[0147] The dynamic interface layer is based on Vue.js dynamic rendering interface, such as Figure 13 shown.
[0148] This approach uses a parameter context decoupling engine to isolate and map technical and business parameters, and combines it with dynamic form rendering technology to generate a business-specific configuration interface. This approach addresses NiFi's high technical barrier to entry, the complexity of task configuration, and the difficulty users face in understanding its use. It is suitable for rapid deployment and configuration management in enterprise-level data integration scenarios.
[0149] The embodiment of the present invention also provides a system for simplifying NiFi data integration task configuration, such as Figure 1 As shown, it includes: technical expert side, semantic engine side, and business user side.
[0150] On the technical expert side, based on NiFi's visual interface, develop a task synchronization template, map business attributes to NiFi's Parameter Context parameters, improve the business semantics of each business attribute, and fill it into the Description field in JSON string format;
[0151] On the semantic engine side, define the business semantics model, implement the semantic bidirectional conversion engine, and realize the automatic conversion between NiFi parameter context and business semantic model;
[0152] On the business user side, the interface is dynamically rendered based on Vue.js;
[0153] The system simplifies NiFi data integration task configuration by using the method for simplifying NiFi data integration task configuration described in the above embodiment.
[0154] The technical experts side,
[0155] Based on NiFi's visual interface, NiFi technical experts use NiFi components, including Processor, Controller Service, Process Group, etc., to develop a general task synchronization template, extract business-related attributes, and map them to parameters in NiFi's Parameter Context component for business semantic layer analysis.
[0156] Taking the MYSQL to MYSQL synchronization template as an example, the NiFi version is 2.2.0, and the specific implementation is as follows:
[0157] 1. Create a ProcessGroup named "MYSQL Synchronization MYSQL Template" and create and reference a ParameterContext named "MYSQL to MYSQL Synchronization Parameters";
[0158] 2. Create the QueryDatabaseTable Processor component for the source MYSQL acquisition. Since this component relies on a Controller Service component to establish a data connection, you need to first create a DBCPConnectionPoolController Service named sourceDatabase.
[0159] The DBCPConnectionPool Controller Service component contains the following key properties: DatabaseConnection URL, Database Driver Class Name, Database User, Password, Max WaitTime, and Max Total Connections.
[0160] The QueryDatabaseTable Processor component contains the following key properties: Database ConnectionPooling Service, Database Type, Columns to Return, Additional WHERE clause, Custom Query, Maximum-value Columns, Initial Load Strategy, Max Wait Time, FetchSize, Max Rows Per Flow File, Output Batch Size.
[0161] 3. Create the target MYSQL import component PutDatabaseRecord Processor. Since this component relies on a Controller Service component to establish a data connection, you need to first create a DBCPConnectionPoolController Service named targetDatabase.
[0162] The DBCPConnectionPool Controller Service component contains the following key properties: DatabaseConnection URL, Database Driver Class Name, Database User, Password, Max WaitTime, and Max Total Connections.
[0163] The PutDatabaseRecord Processor component includes the following key properties: Database ConnectionPooling Service, Record Reader, Database Type, Table Name, Statement Type, MaxWait Time, Rollback On Failure, Table Schema Cache Size, and Maximum Batch Size.
[0164] 4. Through steps 1 to 3, the entire synchronization task is mapped to a ParameterContext with only business attributes. All attributes such as Figure 8 shown.
[0165] 5. The parameters in the Parameter Context contain only four attributes: Name, Value, Sensitive Value, and Description. Other attributes need to be completed according to the business semantic layer model. Use the Description attribute to fill in the required fields in JSON format in the Description, and then parse it through the business semantic layer.
[0166] 6. After the business semantics are supplemented and improved, download the task template through Download Flow Definition in the Process Group and perform subsequent operations.
[0167] On the semantic engine side,
[0168] The semantic engine serves as a bridge between the technical implementation layer and the dynamic interface layer, defining business semantic patterns and implementing a semantic bidirectional conversion engine: it can automatically convert NiFi parameter contexts and business semantic models. It can parse NiFi task templates exported by the technical implementation layer, extract the Parameter Context attributes in the templates, and convert them into business semantic layer patterns for dynamic interface layer rendering. It can also convert form information filled in by users in the dynamic interface layer into Parameter Context attributes, replacing the Parameter Context attributes of the NiFi template, generating task template instances, and importing them into NiFi for execution.
[0169] The schema for business semantics is defined as follows:
[0170] coding name Remark id Unique identifier of the form item paramCode Form item encoding paramName Form item name paramType Form item type boolean,string,number,date defaultValue Default value of form items paramsValue Form item parameter value showType Form item display type INPUT,SELECT require Is this field required? true,false description Form item description order Form item order
[0171] The specific implementation is as follows:
[0172] 1. Parsing business semantics in the task template: Use a JSON parsing tool to parse the NiFi task template, which contains the following attributes: flowContents, externalControllerServices, parameterContexts, flowEncodingVersion, parameterProviders, and latest. Extract the parameterContexts attribute and then take the first attribute of the parameterContexts attribute, parameters, as the defined business attribute parameters. Iterate through the description field of each business attribute in parameters to obtain the business semantics defined, and render them through the dynamic interface layer.
[0173] 2. Business semantic conversion task template: After the business personnel fill in the business attributes through the form, they will fill in the parameter value paramValue of the business attribute into the value field of the corresponding business parameter in the template parameters.
[0174] The service user side,
[0175] The dynamic interface layer is based on Vue.js dynamic rendering interface, such as Figure 13 shown.
[0176] An embodiment of the present invention also provides a device for simplifying NiFi data integration task configuration, comprising: at least one memory and at least one processor;
[0177] The at least one memory is configured to store a machine-readable program;
[0178] The at least one processor is configured to call the machine-readable program to implement the method for simplifying NiFi data integration task configuration described in the above embodiment.
[0179] Embodiments of the present invention further provide a computer-readable medium having computer instructions stored thereon. When executed by a processor, the computer instructions implement the method for simplifying NiFi data integration task configuration described in the above embodiments. Specifically, a system or device equipped with a storage medium can be provided, on which software program code implementing the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device can be configured to read and execute the program code stored in the storage medium.
[0180] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute part of the present invention.
[0181] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0182] In addition, it should be clear that the functions of any of the above embodiments can be achieved not only by executing the program code read by the computer, but also by enabling the operating system operating on the computer to complete part or all of the actual operations based on the instructions of the program code.
[0183] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU installed on the expansion board or expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above embodiments.
[0184] The present invention has been shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the scope of protection of the present invention.
Claims
1. A method for simplifying NiFi data integration task configuration, characterized in that, The implementation of this method includes technical implementation layer, business semantic layer and dynamic interface layer. The technical implementation layer develops a task synchronization template based on NiFi's visual interface, maps business attributes to NiFi's Parameter Context parameters, improves the business semantics of each business attribute, and fills it into the Description field in JSON string format; The business semantic layer defines the business semantic model, implements the semantic bidirectional conversion engine, and realizes the automatic conversion between NiFi parameter context and business semantic model; The dynamic interface layer is based on Vue.js dynamic rendering interface.
2. A method for simplifying NiFi data integration task configuration according to claim 1, characterized in that, The technology implementation layer, Use NiFi components, including Processor, Controller Service, and Process Group components, to develop a general task synchronization template, extract business-related attributes, and map them to parameters in NIFI's Parameter Context component for business semantic layer analysis.
3. A method for simplifying NiFi data integration task configuration according to claim 1 or 2, characterized in that, The technical implementation layer develops a MYSQL to MYSQL synchronization template. The specific implementation process is as follows: 1) Create a ProcessGroup named "MYSQL Synchronization MYSQL Template" and create and reference a Parameter Context named "MYSQL to MYSQL Synchronization Parameters"; 2) Create the QueryDatabaseTable Processor component for the source MYSQL acquisition. Since this component relies on a Controller Service component to establish a data connection, you need to first create a DBCPConnectionPoolController Service named sourceDatabase. 3) Create the target MYSQL import component PutDatabaseRecord Processor. Since this component relies on a Controller Service component to establish a data connection, you need to first create a DBCPConnectionPoolController Service named targetDatabase. 4) Through steps 1) to 3), the entire synchronization task is mapped to a ParameterContext containing only business attributes; 5) The parameters in the Parameter Context contain only four attributes: Name, Value, Sensitive Value, and Description. Using the Description attribute, complete the remaining attributes according to the business semantic layer schema. The required fields are added to the Description in JSON format, and the business semantic layer parses and processes them. 6) After the business semantics are supplemented and improved, download the task template through the Download Flow Definition of the Process Group and perform subsequent operations.
4. A method for simplifying NiFi data integration task configuration according to claim 3, characterized in that, The DBCPConnectionPool Controller Service component named sourceDatabase contains the following key properties: Database Connection URL: Database connection URL, which is a business attribute and needs to be mapped to ParameterContext. The component references the mapping parameter. Database Driver Class Name: Database driver class name, which is a business attribute and needs to be mapped to the Parameter Context. The component references the mapping parameter. Database User: Database user name, which is a business attribute and needs to be mapped to the Parameter Context. The component references the mapping parameter. Password: Database user name, which is a business attribute and needs to be mapped to the Parameter Context. The component references the mapping parameter. Max Wait Time: The maximum waiting time for establishing a connection, which is a technical attribute; Max Total Connections: Maximum number of connections, a technical attribute; The QueryDatabaseTable Processor component contains the following key properties: Database Connection Pooling Service: Database connection pool service, a ControllerService component, used to configure data source connections; Database Type: Database type, which is the default attribute and does not require mapping; Table Name: table name, which is a business attribute and needs to be mapped to the Parameter Context. The component references the mapping parameter. Columns to Return: Returns the column name, which is a business attribute and needs to be mapped to the Parameter Context. The component references the mapping parameter. Additional WHERE clause: query conditions, which are business attributes and need to be mapped to the Parameter Context. The component references the mapping parameters. Custom Query: Query SQL, which is a business attribute and needs to be mapped to Parameter Context. The component references the mapping parameter. Maximum-value Columns: Maximum value columns, used for incremental queries; Initial Load Strategy: Initial load strategy, used to define the first query method, starting from the beginning or starting from the maximum value; Max Wait Time: Maximum waiting time, a technical attribute; Fetch Size: The amount of data retrieved from the result set each time; Max Rows Per Flow File: The number of data rows per FlowFile, which is a technical attribute; Output Batch Size: The number of FlowFiles output each time, which is a technical attribute.
5. A method for simplifying NiFi data integration task configuration according to claim 3, characterized in that: The DBCPConnectionPool Controller Service component named targetDatabase contains the following key properties: Database Connection URL: Database connection URL, which is a business attribute and needs to be mapped to ParameterContext. The component references the mapping parameter. Database Driver Class Name: Database driver class name, which is a business attribute and needs to be mapped to the Parameter Context. The component references the mapping parameter. Database User: Database user name, which is a business attribute and needs to be mapped to the Parameter Context. The component references the mapping parameter. Password: Database user name, which is a business attribute and needs to be mapped to the Parameter Context. The component references the mapping parameter. Max Wait Time: The maximum waiting time for establishing a connection, which is a technical attribute; Max Total Connections: Maximum number of connections, a technical attribute; The PutDatabaseRecord Processor component contains the following key properties: Database Connection Pooling Service: A ControllerService component that configures data source connections. Record Reader: Record reader, a Controller Service component. The default is AvroReader and does not require mapping. Database Type: database type, default attribute, no mapping required; Table Name: table name, which is a business attribute and needs to be mapped to the Parameter Context. The component references the mapping parameter. Statement Type: Operation type, including INSERT, UPDATE, DELETE, and UPSERT, business attributes, which need to be mapped to the Parameter Context. The component references the mapping parameters. Max Wait Time: The maximum waiting time, which is a technical attribute. If it is set to 0, it will never time out and will wait forever. Rollback On Failure: error handling mechanism, which is a technical attribute; Table Schema Cache Size: The size of the cached table schema, a technical attribute. Maximum Batch Size: The maximum number of SQL statements executed per database execution. This is a technical attribute.
6. A method for simplifying NiFi data integration task configuration according to claim 1, characterized in that: The business semantic layer, Parsing technology is used to implement the exported NiFi task template, extract the Parameter Context attributes in the template, and convert them into business semantic layer patterns for dynamic interface layer rendering; At the same time, the form information filled in by the user in the dynamic interface layer is converted into Parameter Context attributes, which replaces the Parameter Context attributes of the NiFi template, and the task template instance is generated and imported into NiFi for execution.
7. A method for simplifying NiFi data integration task configuration according to claim 6, characterized in that: The specific implementation process of the business semantic layer is as follows: 1) Task template parsing business semantics: Use JSON parsing tools to parse NiFi task templates, which contain the following attributes: flowContents, externalControllerServices, parameterContexts, flowEncodingVersion, parameterProviders, latest; extract the parameterContexts attribute, and then take the parameters of the first attribute of the parameterContexts attribute as the defined business attribute parameters; traverse and obtain the business semantics defined in the description field of each business attribute in parameters, and render them through the dynamic interface layer; 2) Business semantic conversion task template: After the business personnel fill in the business attributes through the form, they will fill in the parameter value paramValue of the business attribute into the value field of the business parameter corresponding to the template parameters.
8. A system for simplifying NiFi data integration task configuration, characterized in that: include: On the technical expert side, based on NiFi's visual interface, develop a task synchronization template, map business attributes to NiFi's Parameter Context parameters, improve the business semantics of each business attribute, and fill it into the Description field in JSON string format; On the semantic engine side, define the business semantics model, implement the semantic bidirectional conversion engine, and realize the automatic conversion between NiFi parameter context and business semantic model; On the business user side, the interface is dynamically rendered based on Vue.js; The system simplifies NiFi data integration task configuration through the method described in any one of claims 1 to 7.
9. A device for simplifying NiFi data integration task configuration, characterized in that, include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to implement the method according to any one of claims 1 to 7.
10. A computer-readable medium, characterized in that The computer-readable medium stores computer instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 7.