Data integration method and server

By deploying data integration processes for multiple processes in the server, combining data acquisition and processing information, the problem of difficulty in dealing with semi-structured and unstructured data in the prior art is solved, and a simplified data integration process and efficient data processing are achieved.

CN120144537APending Publication Date: 2025-06-13HENAN QINWEI DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510128060.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing data integration methods are difficult to effectively process semi-structured or unstructured data, and users need to understand the functional characteristics of various components and have strong professional dependencies.

Method used

Provides a data integration method to obtain configuration information to execute the data integration process by deploying a data integration process of multiple processes in a server. Configuration information includes data acquisition information and data processing information, and supports the integration of structured, semi-structured and unstructured data.

Benefits of technology

It simplifies the data integration process, improves data processing efficiency, is suitable for multiple types of data integration, reduces the professional dependence of users, and allows non-professional personnel to easily complete data integration operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144537A_ABST
    Figure CN120144537A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data integration method and a server, according to the embodiment of the invention, a data integration process comprising a plurality of processes is deployed in the server, and when configuration information is acquired, the server can execute the data integration process based on the configuration information. Wherein the configuration information comprises data acquisition information and data processing information. Therefore, data can be read from the source component by configuring data acquisition and data processing information, and the data can be structured data, semi-structured data and non-structured data. By performing format conversion on the extracted data, various types of data can be converted into data matched with the format of the target component, and the converted data is written into the target component, so that data integration of various data such as structured data, semi-structured data and non-structured data is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of servers, and in particular, to a data integration method and a server. Background Art

[0002] Data integration is a technology that integrates data from different sources to provide a unified view. The data integration technology enables users to obtain data across multiple systems without having to concern themselves with the actual storage location of the data. Therefore, it has become increasingly widely used.

[0003] Current data integration methods directly copy the database tables of the source component to the destination component and determine that the fields at both ends correspond one by one. However, the application of this data integration method has limitations and is not suitable for the integration of complex types of data such as semi-structured data or unstructured data. Currently, for the data integration of complex data, users need to call various components of a data integration management tool (such as Nifi), such as processors, connectors, and controllers, to trigger an integration job before they can perform subsequent operations. However, using the various components of the data integration management tool requires users to understand the functional characteristics of each component in advance, and the user's professional dependence is strong. Summary of the Invention

[0004] Embodiments of this application provide a data integration method and a server, which can simplify the data integration process and improve data processing efficiency.

[0005] In a first aspect, an embodiment of this application provides a data integration method applied to a server. The method includes:

[0006] The server obtains configuration information; the configuration information includes data collection information and data processing information. The data collection information is used to configure the source component to obtain target data from the source component; the data processing information is used to convert the data into a format that matches the target component; the server determines a data integration process according to the configuration information; the data integration process includes obtaining target data from the source component, converting the format of the target data, and integrating the converted data into the target component; according to the data integration process, the target data is integrated into the target component.

[0007] Thus, the server can execute a data integration process including multiple processes by deploying the data integration process including multiple processes in the server. When obtaining the configuration information, the data integration process can be executed based on the configuration information. Among them, the multiple processes can include data collection, data conversion, and data loading. Among them, the configuration information includes data collection information and data processing information. The data collection information is used to configure the source component to obtain data from the source component; the data processing information is used to convert the data into a format matching the target component. Thus, by configuring the data collection and data processing information, data can be read from the source component, and the data can be structured data, semi-structured data, or unstructured data. By converting the extracted data, various types of data can be converted into data in a format matching the target component, and the converted data is written into the target component. Thus, data integration operations for various data such as structured data, semi-structured data, and unstructured data can be realized.

[0008] In a specific implementation, the data collection information further includes connection information of the source component, and the connection information is used to establish a connection with the source component. Different source component types have different corresponding connection information. Among them, if the source component is a relational database, the connection information includes: database uniform resource identifier, username, password, and database driver class name; if the source component is a non-relational database, the connection information includes: database uniform resource identifier, port number, database name, and keyspace; if the source component is a file system, the connection information includes: file path or directory path, file model; if the source component is an application programming interface, the connection information includes: interface address and HTTP method; if the source component is a message queue and the message queue is Kafka, the connection information includes: Kafka cluster address, consumer group identifier, and topic name; if the message queue is RabbitMQ, the connection information includes: RabbitMQ server address, port number, queue name, username, and password; if the source component is a big data platform and the big data platform is Hadoop HDFS, the connection information includes HDFS namenode address and uniform resource identifier, file path; if the big data platform is Hive, the connection information includes Hive uniform resource identifier, username, and password.

[0009] In another specific implementation, the configuration information further includes target processing rule information, and the server invokes the target processor corresponding to the target processing rule information; the target processor is used to invoke the target processing rule corresponding to the target processing rule for data format conversion; using the target processor, the target data is formatted according to the target processing rule, and the converted data format is the format of the target component. By invoking specialized data processing rules and target processors for format conversion, it helps to improve the efficiency and flexibility of format conversion.

[0010] Optionally, the target processing rules include one or more of the following: date processing rules, file replacement rules, null value processing rules, and ID card desensitization rules; among them, the date processing rules are used to convert dates in different formats into a date format that matches the target component; the file replacement rules are used to convert the extracted text information into text information that matches the target component; the null value processing rules are used to process null values in the extracted data; and the ID card desensitization rules are used to desensitize the ID card numbers in the extracted data.

[0011] In a specific implementation, the null value processing rules include one or more of the following: deleting null values, filling null values, and marking null values; deleting null values indicates deleting records or fields containing null values; filling null values includes filling null values with a specific value, filling null values with the previous non-null value, or filling null values with the next non-null value, and the specific value includes the mean, median, and mode of the extracted data.

[0012] In another specific implementation, the target processing rule is one of multiple pre-stored processing rules, and the server receives adjustment information; the adjustment information includes adding a processing rule, deleting one or more of the multiple processing rules, and modifying the processing rules among the multiple processing rules; according to the adjustment information, add a processing rule, delete, or modify the processing rules among the multiple processing rules. Thus, the target processing rules can be adjusted as needed to improve the flexibility and adaptability of format conversion.

[0013] In another specific implementation, the configuration information further includes data loading information, and the data loading information includes selected data loading strategy information. The data loading strategy includes one of only full-volume loading, only incremental loading, and comprehensive loading information of full-volume loading and incremental loading. The server loads the converted data into the target component based on the data loading strategy. Thus, by providing diverse loading strategies, users can select appropriate loading strategies according to their needs, thereby improving the user experience.

[0014] In yet another specific implementation, the configuration information further includes one or more of the following: first information, second information, third information, fourth information, and fifth information; among them, the first information is used to select whether to set a monitoring mechanism, the second information is used to determine whether to perform data verification, the third information is used to determine whether to perform permission control, the fourth information is used to determine whether to display data integration process information, and the fifth information is used to determine whether to display data integration results.

[0015] In a second aspect, an embodiment of the present application provides a data integration method applied to a display device, and the method includes:

[0016] Display a display page, and the display page includes configuration items of a data integration process;

[0017] Among them, the configuration item is used to configure a data integration process, and according to the data integration process, target data is integrated into a target component; the data integration process includes obtaining the target data from the source component, converting the format of the target data, and integrating the converted data into the target component;

[0018] In response to a configuration operation for the configuration item, configuration information is obtained; the configuration information includes data collection information and data processing information, the data collection information is used to configure the source component to obtain target data from the source component; the data processing information is used to convert the data into a format matching the target component;

[0019] The configuration information is output to configure the data integration process according to the configuration information.

[0020] In the embodiments of the present application, a user can configure data collection and data conversion through a display page to implement reading data from a source component, and the data can be structured data, or semi-structured data and unstructured data. By converting the format of the extracted data, various types of data can be converted into data in a format matching the target component, and the converted data is written into the target component. Therefore, compared with the method of directly copying the database table of the source component to the target component, the applicable range is wider. And the user only needs to perform simple configuration to realize the automation of the data integration process. The user does not need to understand the functional characteristics of various components, and even non-professionals can easily complete the data integration operation.

[0021] In a third aspect, embodiments of the present application provide a server, including:

[0022] A memory for storing a program;

[0023] A processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of the first aspect.

[0024] In a fourth aspect, the present application provides a computer storage medium for storing a computer program. When the computer program is executed, it is used to implement the method provided by any one of the implementation manners in the first aspect of the present application.

[0025] In a fifth aspect, the present application provides a computer program product containing instructions. When it runs on at least one computing device, at least one computing device is enabled to implement the method provided by any one of the implementation manners in the first aspect of the present application.

[0026] Any of the data processing methods provided above, corresponding computing devices, computer-readable storage media, computer program products, etc. are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 FIG. is an application scenario diagram of a data integration method provided by an embodiment of the present application;

[0028] Figure 2 FIG. is a schematic diagram of a data integration method provided by an embodiment of the present application;

[0029] Figure 3 FIG. is a schematic diagram of a display page provided by an embodiment of the present application;

[0030] Figure 4 FIG. is a schematic diagram of another display page provided by an embodiment of the present application;

[0031] Figure 5 FIG. is a schematic diagram of yet another display page provided by an embodiment of the present application;

[0032] Figure 6 FIG. is a schematic diagram of a display page provided by an embodiment of the present application;

[0033] Figure 7 FIG. is a schematic diagram of yet another display page provided by an embodiment of the present application;

[0034] Figure 8 FIG. is a schematic diagram of the structure of a data integration device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0036] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0037] In the description of the embodiments of the present application, the terms "first", "second", etc. in the specification and claims are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first target object and the second target object, etc. are used to distinguish different target objects, rather than to describe a specific order of the target objects.

[0038] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.

[0039] In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more. For example, a plurality of processing units means two or more processing units; a plurality of systems means two or more systems.

[0040] Data integration refers to the process of combining data from different sources, different formats, and different data structures to provide a unified view. Data integration often involves extracting data from multiple different source components and then transforming the extracted data so that the transformed data can be compatible with the target component. Among them, the source component is the starting point of the data stream, and the source component can be a file system, database, message queue, interface service, etc. that provides data. The target component is the end point of the data stream and is used to receive the data provided by the source component. The target component can be a file system, database, message queue, and interface service, etc. that receives data.

[0041] The data integration method provided by the embodiments of the present application can be applied to the simple data integration of structured data. Among them, structured data refers to data with a fixed schema and format, usually stored in a table form, containing rows and columns. Each row represents a record, and each column represents a field or attribute. Each field has a clear data type (such as integer, string, date, etc.), and the structure of the data is predefined.

[0042] It can also be applied to the data integration of unstructured data and semi-structured data. Unstructured data refers to data without a fixed schema or structure, such as text files, images, audio, and video, etc. Semi-structured data is a data format between structured data and unstructured data. It has a certain structure, but does not strictly follow the schema like structured data. For example, log files, XML files, HTML documents, JSON files, and YAML files, etc.

[0043] An embodiment of the present application provides a data integration method. By deploying a data integration process including multiple processes in a server, when obtaining configuration information, the data integration process can be executed based on the configuration information. Among them, the multiple processes can include data collection, data conversion, and data loading. Among them, the configuration information includes data collection information and data processing information. The data collection information is used to configure the source component to obtain data from the source component; the data processing information is used to convert the data into a format that matches the target component. Thus, by configuring the data collection and data processing information, it is possible to read data from the source component, and the data can be structured data, semi-structured data, or unstructured data. By performing format conversion on the extracted data, various types of data can be converted into data that matches the target component format, and the converted data is written into the target component. Thus, data integration operations for various data such as structured data, semi-structured data, and unstructured data are realized.

[0044] The data integration method provided by the embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0045] The data integration method provided by the embodiment of the present application is applied to a server. To enable those skilled in the art to better understand the embodiment of the present application, hereinafter, a data integration management tool Nifi is used as an example for illustrative purposes. Among them, Nifi is arranged in a server cluster and is used to execute the data integration method.

[0046] Attached Figure 1 is an application scenario diagram of a data integration method provided by an embodiment of the present application. The server 102 is used to communicate with the display device 101.

[0047] The display device 101 is used to provide a visual interface. Among them, the visual interface can be a Web interface. The visual interface includes multiple steps of pre-orchestrated data integration. Specifically, the visual interface includes data collection, data conversion, data loading, and data synchronization and update.

[0048] The user can independently orchestrate the data integration process through the visual interface. For example, the user selects the source component to be collected, fills in the connection information corresponding to the collected source component, selects the pre-orchestrated data processing rules to perform data conversion, selects the data loading strategy, and selects the data synchronization and update method. Among them, the connection information corresponding to the source component refers to the information for establishing a connection between the server 102 and the source component, including but not limited to database type, server address, port number, etc. The display device 101 receives the configuration information of the independently orchestrated data integration process and sends the received configuration information to the server 102.

[0049] Server 102, specifically Nifi in Server 102, can determine a data integration process based on configuration information. Among them, the data integration process includes obtaining target data from source components, converting the format of the target data, and integrating the converted data into target components; integrating the target data into the target components according to the data integration process.

[0050] The following will describe in detail the data integration method provided by the embodiments of the present application with reference to the accompanying drawings.

[0051] Attached Figure 2 It is a schematic diagram of a data integration method provided by an embodiment of the present application. This method can be applied to Server 102, and this method includes the following steps:

[0052] S210. The display device 101 displays a display page.

[0053] In the embodiments of the present application, the display page is used to display the process of data integration including data collection, data conversion, and data loading. Users can, according to their needs, perform process orchestration operations for data integration on the display page.

[0054] Data collection is used to select source components and extract data from the source components. Users can, through the display page, select source components and fill in the connection information of the source components. Nifi connects to the source components based on the source component information and connection information selected by the users.

[0055] For example, as Figure 3 It is a schematic diagram of a display page provided by an embodiment of the present application. This user interface 300 includes a navigation bar area 301 and a function display area 302. The navigation bar area 301 displays multiple processes. Figure 3 The multiple processes shown are data collection and data processing respectively. The function display area 302 is used to display the configuration items corresponding to the specified process. Users can configure the specified process through the configuration items.

[0056] As Figure 3 shown in (a) of, when the user clicks on data collection in the navigation bar area 301, the function display area 302 displays the information that needs to be configured in the data collection process. Specifically, the function display area 302 includes a source component selection item and a connection information input box item. The user selects the target source component as MySQL through the source component selection item, and the connection information filled in the connection information input box is specifically: the Uniform Resource Locator (URL) of MySQL, username, password, and database driver class name, etc., specifically as Figure 3As shown in (b) thereof. Thus, Nifi 100 can configure the processor based on the data URL, username, password, and database driver class name to establish a connection with the MySQL database and execute the data collection task.

[0057] It should be noted that different source component types may have different corresponding connection information. In the embodiments of the present application, different connection information is set for different types of source components, which will be specifically described below.

[0058] If the source component is a file system, the connection information includes but is not limited to the file path, directory path, and file mode. Among them, the file path is the complete path pointing to the specific location of a single file in the file system. For example, the file path can be "C:\Users\username\example.txt", which points to the example.txt file under the Username subfolder in the Users folder on drive C. A directory refers to a folder containing multiple files or subdirectories. The directory path is the path leading to this folder. The file mode is a set of rules or patterns used to match specific parts in the file name or path.

[0059] If the source component is a non-relational database, the corresponding connection information is set according to the non-relational database type. Specifically, if the non-relational database is MongoDB, the connection information includes but is not limited to the following: database URL, database name, and collection name. If the non-relational database is Cassandra, the connection information includes but is not limited to the following: hostname or IP address, port number, and keyspace. If the non-relational database is Redis, the connection information includes the hostname or IP address, port number, and password, etc.

[0060] If the source component is an Application Programming Interface (API), the connection information includes but is not limited to the following: API URL, HTTP method, request header information, request body content, and authentication information.

[0061] If the source component is a message queue, connection information is set according to the message queue type. Specifically, for Kafka, its connection information includes but is not limited to the following: Kafka cluster address, consumer group ID, and topic name. For RabbitMQ, its connection information includes but is not limited to the following: RabbitMQ server address, port number, queue name, username, and password. For a big data platform, corresponding connection information is set according to the platform type. For example, if the big data platform is Hadoop HDFS, the connection information includes but is not limited to the following: HDFS NameNode address or URI, and file path. If the big data platform is Hive, the connection information includes but is not limited to the following: Hive URL, username, and password.

[0062] Data conversion is used to convert data in different formats and structures into a unified format. In the embodiments of the present application, data conversion includes data type conversion, field mapping, unit conversion, etc. The user can configure data conversion information such as data type conversion, field mapping, and unit conversion through the display page. The server 102 converts the data extracted from the source component into a unified data format based on the data conversion configured by the user, and synchronizes the converted data to the target component. For example, the user can configure data type conversion through the display page, such as converting the data extracted from the source component into a Json file uniformly. Thus, the server 102 obtains the converted Json file based on the configured data type conversion.

[0063] For example, the data extracted from the source component includes CSV files, Json files, and database tables. Among them, the CSV file contains user information, and the specific fields are name, age, email, which represent name, age, and email respectively. By way of example, the following is an example of a CSV file:

[0064] Name, age, emai

[0065] A, 30, A@example.com

[0066] The Json file contains order information, and the specific fields are order Id, customer Name, amount, which are the order identification code, customer name, and order quantity respectively. Exemplarily, the following is an example of a Json file:

[0067]

[0068] The database table contains product information, and the specific fields are product Id, product Name, price, which are the product identification code, product name, and product price respectively. Exemplarily, the following is an example of database table data.

[0069] SELECT productId, productName, price FROM products WHERE productId = 1;

[0070] Result:

[0071] --productId: 1, productName: 'Product A', price: 200.00

[0072] Server 102 unifies the conversion of CSV files, Json files, and database tables into Json files for writing into the target component. The specific conversion method is as follows: Parse the data of CSV files, Json files, and database tables, and obtain the Json file through data type conversion, field mapping, and unit conversion. For example, the obtained Json file is as described below:

[0073]

[0074] In addition, considering that there are abnormal data in the data directly extracted from the source component, such as data errors, data missing, and duplicates, if the data conversion is directly performed, it will affect the conversion quality. Therefore, the data conversion also includes data cleaning. Users can select whether to configure data cleaning from the display page according to their needs.

[0075] Furthermore, to improve the subsequent data processing process, the data after data conversion can be aggregated. Exemplarily, Server 102 can perform summarization or grouping processing on the data after data conversion based on business requirements.

[0076] For example, as Figure 3 shown, if the user clicks on data processing in the navigation bar area 301, the function display area 302 displays data processing information. As Figure 4 shown, the function display area 302 shows data conversion information. Figure 4 Options for whether to perform data cleaning and whether to perform data aggregation are shown. Users can select whether to perform data cleaning and whether to perform data aggregation according to their needs. As Figure 4 shown, when the user selects "Yes" in the box for whether to perform data cleaning, data cleaning is performed. And when the user selects "Yes" in the box for whether to perform data aggregation, data aggregation is performed. Server 102 can perform data cleaning and data aggregation operations based on the user's selection.

[0077] In addition, Figure 4Data conversion rules are also shown. The data conversion rules are custom rules, and users can configure the data conversion rules according to their needs. For example, the data type conversion can be configured to be uniformly converted to the Json format, etc.

[0078] Exemplarily, the server 102 has built-in data conversion rules. Users can check whether the built-in data conversion rules of the server 102 meet the requirements, and then directly trigger the server 102 to perform data conversion operations using these data conversion rules. For example, the user can click the "View" button on the display page to view the built-in data conversion rules of the server 102. To meet the requirements, click "Import" to trigger the server 102 to directly perform conversion operations using the built-in data conversion rules.

[0079] In addition, if the requirements are not met, users can add, delete, or modify the built-in data conversion rules of the server 102. After the modification is completed, click "Import" to trigger the server 102 to perform conversion operations using the modified data conversion rules.

[0080] It should be noted that the above display page is only for illustrative purposes. In actual use, the navigation bar area 301 of the display page may also include other processes, such as source component management, task monitoring, data loading, data synchronization and update, etc. The embodiments of the present application do not specifically limit this.

[0081] It should be noted that the display page only shows one process at a time. For example, the display page first shows data collection information. After the user selects a source component, the content of the display page is updated, and the updated display page shows the connection information of the selected source component. After the user enters the connection information on the updated display page, the content of the display page is updated again, and the display page after the second update shows data conversion information. After the user selects the data conversion information, the content of the display page can be updated for the third time as needed, and the display page after the third update shows data loading information. After the user selects the data loading information, the content of the display page is updated for the fourth time, and the content of the display page after the fourth update is data synchronization and update, etc.

[0082] S220. The display device 101 obtains data integration configuration information based on the configuration operation.

[0083] The configuration operation is an interactive operation between the user and the display page. Exemplarily, as Figures 3 - 4 shown, the configuration operation is an operation of selecting a source component, filling in the connection information corresponding to the selected source component, and selecting data conversion information.

[0084] Based on the configuration operation, the display device 101 can obtain data integration configuration information. Among them, the data integration process refers to the information corresponding to the execution of the data integration task, such as the source component from which data needs to be extracted, the connection information of the source component, data conversion rules, etc.

[0085] S230. The display device 101 sends the configuration information to the server 102.

[0086] S240. The server 102 performs a data integration task based on the configuration information.

[0087] Specifically, after receiving the configuration information, the server 102 parses the source components included in the configuration information, as well as the connection information, conversion rules, etc. corresponding to the source components. Based on the selected source components and the connection information corresponding to the source components, the server 102 establishes a connection with the source components selected by the user and extracts the required data from the source components.

[0088] Then, the server 102 processes the raw data extracted from the selected source components based on the selected data conversion rules. By way of example, if the raw data includes error data, missing data, and duplicate data, and considering data types including structured, unstructured, and semi-structured data, and if the selected data conversion rules include performing data cleaning operations, operations for converting data of different formats and structures into a unified format, and data aggregation operations. Then the server 102 first cleans the raw data, removing the error data, missing data, and duplicate data in the raw data. Then the server 102 converts the cleaned data into a unified format. Then the data in the unified format is summarized or grouped according to business requirements. For example, the data in the unified format can be grouped into user information data, order information data, and product information data according to business requirements.

[0089] After that, the server 102 loads the data after data conversion according to a preset data loading strategy, and updates the data according to a preset data synchronization and update method.

[0090] Thus, in the embodiments of the present application, the user can configure data collection and data conversion through a display page, realize reading data from source components, and the data can be structured data, or semi-structured data and unstructured data. By converting the extracted data in format, various types of data can be converted into data matching the format of the target component, and the converted data is written into the target component. Therefore, compared with the method of directly copying the database table of the source component to the target component, the applicable range is wider. And the user only needs to perform simple configuration to realize the automation of the data integration process, and the user does not need to understand the functional characteristics of various components, and even non-professionals can easily complete the data integration operation.

[0091] Next, the data integration method provided by the embodiments of the present application will be introduced in combination with a specific application scenario. This scenario is described by taking the source component as a database as an example.

[0092] In the data integration method provided by the embodiments of the present application, users can configure the entire data integration process through a visual interface. Specifically, data collection, data conversion, data loading, data synchronization and update, data quality management, and security and compliance are displayed on the displayable page. For example, as Figure 5 Shown is another schematic diagram of a display page provided by the embodiments of the present application. The display page 300 includes a navigation bar area 301 including information such as data collection, data conversion, data loading, data synchronization and update, data quality management, and security and compliance.

[0093] In the embodiments of the present application, when the user clicks on "Data Conversion" in the navigation bar area 301, the function display area 302 displays the data conversion configuration. As Figure 6 Shown is a schematic diagram of a display page provided by the embodiments of the present application.

[0094] As Figure 6 Shown, the function display area 302 shows a variety of pre-embedded processing rules. By way of example, Figure 6 Four data processing rules are shown, namely the date processing rule, the text replacement rule, the null value processing rule, and the ID card desensitization rule.

[0095] Among them, the date processing rule is used to unify the date format. As Figure 6 Shown, formats such as the date DD / MM / YYYY are unified and converted to the format YYYY-MM-DD. In the embodiments of the present application, the date processing rule is based on a date processing operator. In one example, a date processing operator interface can be written in Java to extract the date field from the data extracted from the source component, and using regular expressions, based on the date processing operator, the date field is replaced and unified into the format YYYY-MM-DD. The converted date is added to the flow file and output to the target component.

[0096] In addition, the display page can also show the necessary information of the date processing rule, such as the type of processing parameters, the standard date format after processing, etc. In addition, the display page will also show whether this rule can be modified. Figure 6 The date processing rule shown is a built-in rule, that is, this rule cannot be changed, the parameter type is date, and the processed standard date format is YYYY-MM-DD.

[0097] It can be understood that the display method of the above display page is only for illustrative purposes. In actual use, the parameter type and the standard date format can be adjusted based on needs, and the display method of each field can also be adjusted based on actual needs. The embodiments of the present application do not specifically limit.

[0098] In addition, the user can also process the date processing rules, such as viewing the date processing rules and testing the date processing rules. As Figure 6 shown, when the user clicks the "View" button in the date processing rule box, the date processing rules can be viewed. When the user clicks the "Test" button in the date processing rule box, the date processing rules can be tested.

[0099] The text replacement rule is also called the data conversion rule, which is used to unify the data format of non-date data. As Figure 6 shown, formats such as Json, csv, and database tables are uniformly converted into the target format. In the embodiment of the present application, the text replacement rule is based on a text replacement operator. In one example, a text replacement operator interface is written in Java, and the text information input by the user is extracted from the file stream to be replaced using a regular expression, and then replaced with the replacement content input by the user, and then added to the file stream as an attribute.

[0100] In addition, the display page can also display the necessary information of the text replacement rule, such as the parameter type to be processed, the processing method, etc. In addition, the display page will also show whether the rule can be modified. Figure 4 The text replacement rule shown is a custom rule. The user can adjust the rule based on needs. The parameter type is a string, and the processing method is regular expression search and processing.

[0101] It can be understood that the display method of the above display page is only a schematic representation. In actual use, the parameter type and the processing method can be adjusted based on needs, and the display method of each field can also be adjusted based on actual needs. The embodiment of the present application does not specifically limit.

[0102] In addition, the user can also process the text replacement rule, such as viewing the text replacement rule and testing the text replacement rule. As Figure 6 shown, when the user clicks the "View" button in the text replacement rule box, the text replacement rule can be viewed. When the user clicks the "Test" button of the text replacement rule, the text replacement rule can be tested.

[0103] The null value processing rule is used for the null value problem in data processing and data analysis. In the embodiment of the present application, Nifi can process the null value based on a null value processing operator. In one example, the user can write a custom processing operator interface in Java, and the user can select a null value processing operator according to needs.

[0104] For example, if the null value handling operator is the delete null operator, Nifi will delete the records or fields containing null values. If the null value handling operator is the fill null operator, Nifi will fill the null values with specific values, where the specific values can be the mean, median, or mode, etc., and can be adjusted according to needs. If the null value handling operator is the forward fill operator, the null values will be filled with the previous non-null value. If the null value handling operator is the backward fill operator, the null values will be filled with the next non-null value. If the null value handling operator is the marking operator, the null values will be marked out.

[0105] In addition, the display page can also show the necessary information of the null value handling rule, such as the parameter type of the processing, the processing method, etc. In addition, the display page will also show whether the rule can be modified. Figure 6 The shown null value handling rule is a custom rule. Users can adjust the rule based on needs. The parameter type is date, and the processing method is null value conversion.

[0106] It can be understood that the display method of the above display page is only for illustrative purposes. In actual use, the parameter type and the processing method can be adjusted based on needs, and the display method of each field can also be adjusted based on actual needs. The embodiments of the present application do not specifically limit.

[0107] In addition, users can also process the null value handling rule, such as viewing the null value handling rule and testing the null value handling rule. As Figure 6 shown, when the user clicks the "View" button in the null value handling rule box, the null value handling rule can be viewed. When the user clicks the "Test" button of the null value handling rule, the null value handling rule can be tested.

[0108] The ID card desensitization rule is used to desensitize the ID card information to protect the ID card information. In the embodiments of the present application, Nifi embeds an ID card decryption operator, and the ID card desensitization rule is based on the ID card desensitization operator for ID card desensitization. For example, the ID card desensitization operator retains the first 6 digits and the last 4 digits of the ID card number based on the characteristics of the ID card number, and replaces the middle 8 digits with a preset symbol for desensitization. For example, desensitize with the "*" symbol. The desensitized data is added to the data stream for processing.

[0109] In addition, the display page can also show the necessary information of the ID card desensitization rule, such as the parameter type of the processing, the processing method, etc. In addition, the display page will also show whether the rule can be modified. Figure 6 The shown ID card desensitization rule is a built-in rule. Users can't modify this rule. The parameter type is date, and the processing method is ID card desensitization.

[0110] In addition, users can also process the ID card desensitization rule, such as viewing the ID card desensitization rule and testing the ID card desensitization rule. AsFigure 6 As shown, when the user clicks the "View" button in the ID card desensitization rule box, the ID card desensitization rule can be viewed. When the user clicks the "Test" button of the ID card desensitization rule, the ID card desensitization rule can be tested.

[0111] It can be understood that the display method of the above display page is only a schematic representation. In actual use, the parameter type and processing method can be adjusted according to needs, and the display method of each field can also be adjusted according to actual needs. The embodiments of the present application do not specifically limit this.

[0112] It should be noted that Figure 6 The function display area in also includes "Import", "New", "Delete" and search items. After the user selects a rule and clicks "Import", all the selected rules in the function display area will be sent to the server 102. Clicking "New" can create a new rule, and clicking "Delete" can delete the selected rule. The user can search for rules through the search item.

[0113] It should be noted that in actual use, each processing rule corresponds to a processor. For example, the date processing rule corresponds to the first processor, the text replacement rule corresponds to the second processor, the null value processing rule corresponds to the third processor, and the ID card desensitization rule corresponds to the fourth processor. After the Nifi in the server receives the processing rule information configured by the user, it will call the corresponding processor to execute the processing rule corresponding to the processing rule information. For example, if the processing rule information is date processing rule information, the server will call the first processor to perform the operation of converting the target data based on the date processing rule.

[0114] Data loading is used to load the extracted data into the target component. In the embodiments of the present application, data loading includes but is not limited to: target component selection and loading strategy selection. Among them, target component selection is used to configure the target component and output the data extracted from the source component to the target component. Loading strategy selection is used to configure the loading strategy.

[0115] In one example, the loading strategy includes full loading and incremental loading. The user can choose to load only fully, load only incrementally, or choose a comprehensive loading strategy that combines full loading and incremental loading, etc.

[0116] Full loading means that all data will be reloaded regardless of whether the data already exists in the target component. Full loading can ensure that the data set loaded each time is complete, which helps to ensure that the data in the target component is exactly the same as the data in the source component. However, full loading consumes a large amount of resources and takes a long time. Frequent full loading may affect the performance of the source component and the target component.

[0117] Incremental loading only loads changed data to the target component. The changed data refers to newly added, modified, or deleted data. Loading on demand consumes less resources and is fast.

[0118] The comprehensive loading strategy is a pre-set loading strategy, such as periodic full loading, and between two full loadings, an incremental loading strategy is adopted, such as the first full loading and then the incremental loading strategy, etc. The embodiments of the present application are not specifically limited.

[0119] For example, Figure 5 As shown, the user clicks on the data loading in the navigation bar area 301, and the function display area 302 displays the data loading configuration items, such as Figure 7 The figure shows another display page provided by the embodiment of the present application. The data loading configuration item includes a target component selection item and a loading strategy item. The user can select the target component through the target component selection item, for example, Figure 7 The drop-down menu of the target component selection item includes multiple target components, specifically relational database, non-relational data, message queue, big data platform, API interface, etc. Furthermore, each type of target component can further include multiple sub-target components, for example, message queue is further subdivided into Pulsar, Kafla, etc. Users can select the target component by clicking.

[0120] like Figure 7 As shown, users can configure loading strategies by loading measurement items. Figure 7 The drop-down menu shown includes multiple loading strategies, specifically, full load only, incremental load only, and periodic full load with intermediate incremental load. The user can select a loading strategy by clicking.

[0121] Data synchronization and update is used to update and synchronize the data of the source component to the target component in real time. For example, the data synchronization and update method is: a combination of real-time synchronization and batch synchronization, that is, real-time synchronization is performed for data that needs to be updated in real time, and regular batch updates are performed for data that does not need to be updated in real time.

[0122] For example, Figure 5 As shown, the user clicks on data synchronization and update in the navigation bar area 301, and the function display area 302 displays the synchronization update policy items, wherein the synchronization update policy items include multiple synchronization update policies, such as real-time synchronization policy only, batch synchronization policy only, and a comprehensive policy of real-time synchronization and batch synchronization. The user can select the synchronization update policy according to needs.

[0123] Data quality management is used to manage the quality of data to be integrated. For example, the data to be integrated is verified at each stage of data integration to ensure the correctness, consistency, and integrity of the data. For example, a monitoring mechanism is set up to monitor abnormal data. Security and compliance are used to manage permissions for data integration methods. For example, Nifi only allows authorized personnel to access and operate sensitive data, uses encryption technology to protect sensitive information, and complies with relevant laws, regulations, and industry standards, etc.

[0124] For example, as Figure 5 shown, when the user clicks on data quality management in the navigation bar area 301, the function display area 302 will display quality management information. Exemplarily, the function display area includes whether a monitoring mechanism is set up and whether data verification is performed. If the user selects the "Yes" item corresponding to setting up the monitoring mechanism, it means setting up the monitoring mechanism, and the data flow of data integration can be monitored. If the user selects "Yes" corresponding to data verification, it means selecting data verification, and it can be verified whether the source component data and the target component data are consistent.

[0125] As Figure 5 shown, when the user clicks on security and compliance in the navigation bar area 301, the function display area 302 will display security and compliance. Among them, the function display area 302 includes whether permission control is performed. If the user selects the corresponding "Yes" item, it means performing permission control, and only authorized personnel can access and operate.

[0126] Furthermore, the display page provided in the embodiment of the present application also includes task monitoring, source component management, and integration reconciliation. For example, as Figure 5 shown, the navigation bar area 301 also includes task monitoring, source component management, and integration reconciliation. When the user clicks on task monitoring, the display area of the function display area 302 can be used to configure the monitoring of the data integration process and display the data integration results. Furthermore, the display methods of the data integration process and results can also be configured. For example, the processing situation of data integration can be displayed and analyzed in real time in the form of statistical charts.

[0127] When the user clicks on source component management, source components can be added, modified, and deleted in the display area of the function display area 302. When the user clicks on integration reconciliation, integration reconciliation rules can be configured in the display area of the function display area 302, etc.

[0128] In an embodiment of the present application, after the user configures data integration configuration information on the display page, the server 102 embeds the rules of the data integration process. For example, the rules of the data integration process are to sequentially execute obtaining the selected source components to be collected and the corresponding connection information, processing the data based on the selected data processing rules, obtaining the target components, loading the data based on the configured data loading strategy, monitoring the data volume, verifying whether the data of the source components is consistent with the data of the target components, only allowing authorized personnel to access and operate, displaying the data integration process, and displaying the data integration result. The server 102 generates a data integration process and executes a data integration task based on the embedded data integration process rules and the obtained data integration configuration information.

[0129] It should be noted that in the above process, if one or several processes are deleted, the execution order remains unchanged. For example, if the processes of only allowing authorized personnel to access and operate and displaying the data integration process are deleted, the data integration process is to obtain the selected source components to be collected and the corresponding connection information, process the data based on the selected data processing rules, obtain the target components, load the data based on the configured data loading strategy, monitor the data volume, verify whether the data of the source components is consistent with the data of the target components, and display the data integration result.

[0130] In a specific implementation, a custom script can be written in the Java language to call the preset API of NiFi to dynamically create and configure a data collection process. The specific creation method includes:

[0131] Process 1: Define the API interface.

[0132] In an embodiment of the present application, by defining the API interface, the API endpoints and request methods to be called are described, and the functions and usage logics of the API are clarified.

[0133] Process 2: Create an HTTP connection.

[0134] Using the HTTP client library of Java, an HTTP request is sent to the NiFi server to establish an HTTP connection between the web interface and the source components.

[0135] Process 3: Read the response.

[0136] Obtain and parse the HTTP response from the NiFi server, and extract the required information.

[0137] Process 4: Create a new flow file.

[0138] Send a POST request to NiFi to create a new flow file, and record the ID of the flow file.

[0139] Process 5: Write the response content to the flow file.

[0140] Send a PUT request to the previous stream file to write the new configuration data into the stream file, completing the dynamic configuration of the data collection process.

[0141] In summary, the data integration method provided by the embodiments of this application aims to enable users to easily configure and execute complex data integration tasks. Users select source components (such as databases, API interfaces, etc.) through a simple interface, and input connection information (such as database URLs, usernames, passwords, etc.). Through preset data processing rules (such as text replacement, date conversion, handling null values, etc.), configure the data loading strategy (such as full or incremental loading), and display the entire process in a visual chart. Users can also decide whether to set up a monitoring mechanism to discover and solve data quality problems. Finally, the processed data will be stored in the target component. The visual interface integrating components such as source component selection, data processing rules, data loading strategy, and monitoring mechanism significantly improves the convenience and efficiency of data integration. Users can select appropriate tools and methods through simple operations without having to deal with underlying technical details. At the same time, a series of custom operators are provided for users to use, increasing the flexibility and reusability of the system and lowering the usage threshold for non-professional users. At the same time, this solution dynamically configures the data collection process through specific API interfaces and NiFi components, making data integration simpler and more convenient, and applicable to various application scenarios.

[0142] In addition, the embodiments of this application also provide a data integration device.

[0143] Appendix Figure 8 It is a schematic structural diagram of a data integration device provided by the embodiments of this application. The device 800 includes:

[0144] An acquisition unit 801 is used to acquire configuration information; the configuration information includes data collection information and data processing information. The data collection information is used to configure the source component to obtain target data from the source component; the data processing information is used to convert the data into a format that matches the target component.

[0145] A determination unit 802 is used to determine the data integration process according to the configuration information; the data integration process includes obtaining target data from the source component, converting the format of the target data, and integrating the converted data into the target component.

[0146] An integration unit 803 is used to integrate the target data into the target component according to the data integration process.

[0147] Among them, the integration unit 803 is specifically configured to: call a target processor corresponding to target processing rule information; the target processor is used to call a target processing rule corresponding to the target processing rule to perform data format conversion; and use the target processor to perform format conversion on target data according to the target processing rule, where the converted data format is the format of the target component.

[0148] In another specific implementation, the apparatus 800 further includes an adjustment unit, configured to receive adjustment information; the adjustment information includes adding a processing rule, deleting one or more of a plurality of processing rules, and modifying a processing rule among a plurality of processing rules; and adding a processing rule, deleting or modifying a processing rule among a plurality of processing rules according to the adjustment information.

[0149] The embodiments of the present application further provide a computer program product including instructions. The computer program product may be software or a program product including instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on the computing device, the computing device is caused to execute the above data integration method. The embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by the computing device or a data storage device such as a data center including one or more available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that instruct the computing device to execute the above data integration method.

[0150] The descriptions of the processes or structures corresponding to the above respective drawings each have their own focuses. For parts not detailed in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.

[0151] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data integration method, characterized in that: Applied to a server, the method comprises: Acquire configuration information; the configuration information includes data acquisition information and data processing information, the data acquisition information is used to configure the source component to acquire target data from the source component; the data processing information is used to convert the data into a format that matches the target component; Determine a data integration process according to the configuration information; the data integration process includes acquiring the target data from the source component, converting the format of the target data, and integrating the converted data into the target component; According to the data integration process, the target data is integrated into the target component.

2. The method according to claim 1, characterized in that The data collection information also includes connection information of the source component, and the connection information is used to establish a connection with the source component; Wherein, if the source component is relational data, the connection information includes: database uniform resource identifier, user name, password and database driver class name; If the source component is a non-relational database, the connection information includes: a database uniform resource identifier, a port number, a database name, and a key space; If the source component is a file system, the connection information includes: a file path or a directory path, and a file model; If the source component is an application program interface, the connection information includes: an interface address and an HTTP method; If the source component is a message queue, and the message queue is Kafka, the connection information includes: Kafka cluster address, consumer group identifier, and topic name; if the message queue is RabbitMQ, the connection information includes: RabbitMQ server address, port number, queue name, user name, and password; If the source component is a big data platform, if the big data platform is Hadoop HDFS, the connection information includes the HDFS name node address and uniform resource identifier, and the file path; if the big data platform is Hive, the connection information includes the Hive uniform resource identifier, user name and password.

3. The method according to claim 1, characterized in that The configuration information also includes target processing rule information, and the method further includes: Calling a target processor corresponding to the target processing rule information; the target processor is used to call a target processing rule corresponding to the target processing rule to perform data format conversion; The format conversion of the target data includes: The target processor is used to convert the target data into a format according to the target processing rule, wherein the converted data format is the format of the target component.

4. The method according to claim 3, characterized in that The target processing rules include one or more of the following: date processing rules, file replacement rules, null value processing rules and ID card desensitization rules; Among them, the date processing rule is used to convert dates in different formats into the date format that matches the target component; the file replacement rule is used to convert the extracted text information into text information that matches the target component; the null value processing rule is used to process null values ​​in the extracted data; and the ID card desensitization rule is used to desensitize the ID card number in the extracted data.

5. The method according to claim 4, characterized in that The null value processing rules include one or more of the following: deleting null values, filling null values ​​and marking null values; The deleting empty values ​​indicates deleting records or fields containing empty values; the filling empty values ​​includes one of filling empty spaces with specific values, filling empty values ​​with forward non-empty values, and filling empty values ​​with backward non-empty values, and the specific values ​​include the mean, median and mode of the extracted data.

6. The method according to claim 3, characterized in that The target processing rule is one of a plurality of pre-stored processing rules, and the method further includes: receiving adjustment information; the adjustment information includes adding a new processing rule, deleting one or more of the multiple processing rules, and modifying a processing rule of the multiple processing rules; According to the adjustment information, a new processing rule is added, or a processing rule among the multiple processing rules is deleted or modified.

7. The method according to claim 1, characterized in that The configuration information further includes data loading information, the data loading information includes selected data loading strategy information, the data loading strategy includes one of full loading only, incremental loading only, and comprehensive loading information of full loading and incremental loading; The method further includes: loading the converted data into the target component based on the data loading strategy.

8. The method according to claim 1, characterized in that The configuration information further includes one or more of the following: first information, second information, third information, fourth information and fifth information; Among them, the first information is used to select whether to set up a monitoring mechanism, the second information is used to determine whether to perform data verification, the third information is used to determine whether to perform authority control, the fourth information is used to determine whether to display data integration process information and the fifth information is used to determine whether to display data integration results.

9. A data integration method, characterized in that: Applied to a display device, the method comprises: Displaying a display page, wherein the display page includes configuration items of the data integration process; The configuration item is used to configure a data integration process, so as to integrate the target data into the target component according to the data integration process; the data integration process includes acquiring the target data from the source component, converting the format of the target data, and integrating the converted data into the target component; In response to the configuration operation for the configuration item, configuration information is acquired; the configuration information includes data acquisition information and data processing information, the data acquisition information is used to configure the source component to acquire target data from the source component; the data processing information is used to convert the data into a format matching the target component; The configuration information is output to configure the data integration process according to the configuration information.

10. A server, characterized in that: The invention comprises a memory and a processor, wherein the memory and the processor are coupled: The memory is used to store programs; The processor is configured to execute the method according to any one of claims 1 to 8 based on the program stored in the memory.