A multi-source heterogeneous data integration system based on distributed computing

Through the combination of distributed computing and formatting processing networks, the problems of efficiency and flexibility in multi-source heterogeneous data integration systems are solved, and the efficient and flexible processing of data integration systems is achieved to meet the diversified needs of enterprises.

CN118708585BActive Publication Date: 2025-07-04GUANLIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410737328.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-07
Publication Date
2025-07-04
Estimated Expiration
2044-06-07

AI Technical Summary

Technical Problem

The existing multi-source heterogeneous data integration system has cumbersome steps in the data integration process, which cannot comprehensively improve the data integration efficiency in enterprise accounts, and can only be processed in a fixed manner during the data integration process, which lacks flexibility.

Method used

A multi-source heterogeneous data integration system based on distributed computing is adopted, including data acquisition, storage, preprocessing, allocation and integration modules. By setting up a distributed formatting processing network and a distributed computing processing network, combining resource configuration tables and enterprise organizational structures, flexible data processing and efficient integration are achieved.

Benefits of technology

It improves the resource utilization efficiency and flexibility of the data integration system, ensures the efficiency and flexibility of the data integration process, and adapts to different enterprise needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118708585B_ABST
    Figure CN118708585B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-source heterogeneous data integration system based on distributed computing. The present invention relates to the field of data integration and includes a data integration platform, which comprises a data acquisition module, a data storage module, a data preprocessing module, a data distribution module, and a data integration module. The data acquisition module is used to acquire multi-source heterogeneous data and load data within an enterprise account. The data storage module is used to store historical multi-source heterogeneous data. The data preprocessing module is used to set up a distributed formatting processing network according to the historical multi-source heterogeneous data and select a distributed formatting processing node based on its load data to perform format-unifying processing on the multi-source heterogeneous data. The data distribution module sets up a distributed computing processing network and assigns tasks to it. The data integration module completes data processing according to the task assignment result, obtains the data to be integrated, and performs integration processing on it. The present invention improves the efficiency of data integration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data integration, and specifically to a multi-source heterogeneous data integration system based on distributed computing. Background Art

[0002] A multi-source institutional data integration system refers to integrating data from different sources, different types, and different structures into a unified platform or system, including functions such as data extraction, transformation, and loading, to ensure that it can be effectively integrated and managed. It is usually used in enterprises to help enterprises integrate data scattered in different systems, databases, or files, thereby supporting business requirements such as data analysis, reporting, and decision-making;

[0003] A multi-source heterogeneous data integration system with the publication number of CN116483840A discloses a multi-source heterogeneous data integration system based on distributed computing, which is used to solve the problems that the existing multi-source heterogeneous data integration system cannot reasonably store multi-source heterogeneous data and cannot ensure the storage stability and storage efficiency of data; this multi-source heterogeneous data integration system integrates the data of the data source according to the data source and stores it distributively, so that the stored data does not interfere with each other, ensuring the security of the data. At the same time, it avoids the data being in a mess and makes it easy to be searched. Then, the distributed storage area is supplemented to ensure the sufficiency of the data storage space, ensuring the stability and security of data storage and the storage efficiency. Then, the source data packet is transferred to further ensure the sufficiency of the storage space and be able to store more data;

[0004] However, the steps in the multi-source heterogeneous data integration process are cumbersome. Only performing distributed computing on its integration process cannot comprehensively improve the efficiency of data integration in enterprise accounts. Moreover, during the data integration process, only the fixed integration method can be used to integrate data information, and the data integration system cannot be flexibly utilized; therefore, how to improve the efficiency and flexibility of the data integration system is a problem that we need to solve. For this reason, a multi-source heterogeneous data integration system based on distributed computing is provided now. Summary of the Invention

[0005] In order to solve the above technical problems, the purpose of the present invention is to provide a multi-source heterogeneous data integration system based on distributed computing.

[0006] The purpose of the present invention can be achieved through the following technical solutions: A multi-source heterogeneous data integration system based on distributed computing, including a data integration platform, and the data integration platform is communicatively linked with a data collection module, a data storage module, a data preprocessing module, a data distribution module, a distributed computing module, and a data integration module;

[0007] The data acquisition module is used to acquire the multi-source heterogeneous data corresponding to the enterprise account and the load data of the distributed nodes, mark them according to the data sources of the multi-source heterogeneous data, and mark them according to the data sources.

[0008] The data storage module is used to store the multi-source heterogeneous data according to the data sources corresponding to the enterprise account, and obtain its historical multi-source heterogeneous data.

[0009] The data preprocessing module is used to set up a distributed formatting processing network according to the historical multi-source heterogeneous data corresponding to the enterprise account, select the corresponding distributed formatting processing node according to the load data corresponding to the distributed formatting processing network to uniformly process the format of the multi-source heterogeneous data, and obtain the characteristic data corresponding to the multi-source heterogeneous data.

[0010] The data distribution module is used to mark the source of the characteristic data corresponding to the multi-source heterogeneous data, set up a distributed computing processing network, and assign the characteristic data of the multi-source heterogeneous data according to the source mark and the load data.

[0011] The data integration module calculates the characteristic data according to the task assignment result to obtain the data to be integrated, sets up a fast integration node according to the relationship between the data to be integrated from different data sources for integration processing, and sets up a temporary integration node according to the enterprise account requirements to perform integration processing on the data to be integrated.

[0012] Further, the process of the data acquisition module acquiring the multi-source heterogeneous data corresponding to the enterprise account and the load data of the distributed nodes includes:

[0013] The data integration platform obtains enterprise verification information, which includes the corresponding organizational structure and enterprise qualification information. The organizational structure is the relationship between each organizational department within the enterprise, including multiple organizational departments, and sets up an enterprise account for the obtained enterprise verification information.

[0014] A data acquisition unit and a load acquisition unit are set in the data acquisition module.

[0015] The data acquisition unit is used to obtain the organizational structure corresponding to the corresponding enterprise account, generate an associated window set according to the organizational structure corresponding to the enterprise account; the associated window set includes associated sub-windows corresponding to each organizational department; the associated sub-windows obtain the multi-source heterogeneous data corresponding to the corresponding organizational department within the corresponding enterprise account, mark the obtained multi-source heterogeneous data according to its organizational department, and send it to the data preprocessing module.

[0016] The load acquisition unit is used to acquire the load data of distributed nodes within the platform, mark the acquired load data according to the types of distributed nodes, and send the marked results to the data preprocessing module and the data distribution module respectively.

[0017] Further, the data storage module stores the corresponding multi-source heterogeneous data in the enterprise account according to its data source. The process of obtaining its historical multi-source heterogeneous data includes:

[0018] The data storage module obtains the organizational structure of the corresponding enterprise account, sets a data storage space set according to the organizational structure of the enterprise account, marks the data storage space set according to the enterprise account, including data storage sub-spaces corresponding to the corresponding organizational departments within the organizational structure. The data storage sub-space is used to obtain the historical multi-source heterogeneous data and its format processing results obtained by the corresponding organizational department within the enterprise account, and record them as historical data pairs, and mark and store each historical data pair.

[0019] Further, the process of the data preprocessing module setting up a distributed formatting processing network according to the historical multi-source heterogeneous data corresponding to the enterprise account includes:

[0020] A formatting network construction unit is set in the data preprocessing module;

[0021] The formatting network construction unit is used to set up a corresponding distributed formatting processing network according to the enterprise account, obtain the historical data pairs stored in each data storage sub-space within the data storage space set corresponding to the enterprise account, and count their data volumes, obtain the data volumes stored in each data storage sub-space, and obtain the comprehensive total value of the data volumes stored in the corresponding data storage space set; preset a data monitoring period, obtain the period occupancy ratio of the data volume stored in the corresponding data storage sub-space within the data monitoring period, and obtain the average occupancy ratio corresponding to the period occupancy ratio within the data monitoring period; set a fluctuation occupancy ratio interval according to the period occupancy ratios corresponding to each data monitoring period, obtain the lower limit occupancy ratio and the average occupancy ratio of the fluctuation occupancy ratio interval of each data storage sub-space, preset a fixed node coefficient, and perform analysis and processing according to the lower limit occupancy ratio, the average occupancy ratio and the fixed node coefficient to obtain the fixed distributed processing node occupancy ratio corresponding to each data storage sub-space;

[0022] Allocate fixed distributed formatting processing nodes corresponding to the corresponding data volumes according to the fixed distributed node occupancy ratios corresponding to each storage sub-space. After the allocation is completed, mark the remaining distributed formatting processing nodes as elastic distributed formatting processing nodes; set the distribution of the fixed distributed formatting processing nodes and the elastic distributed formatting processing nodes as the distributed formatting processing network.

[0023] Further, the process of the data preprocessing module for uniformly processing multi-source heterogeneous data according to the distributed formatting processing network includes:

[0024] A formatting processing unit is set in the data preprocessing module;

[0025] The formatting processing unit obtains the formatting processing node load data corresponding to the fixed distributed formatting processing node corresponding to the corresponding data storage subspace in the distributed formatting processing network and the multi-source heterogeneous data collected by the corresponding organization department; obtains the amount of data to be processed of the multi-source heterogeneous data; obtains the idle data according to the formatting processing node load data; compares and analyzes the obtained idle data with the amount of data to be processed. When the idle data is greater than or equal to the amount of data to be processed, the corresponding fixed distributed formatting processing node performs uniform formatting processing on the multi-source heterogeneous data; when the idle data is less than the amount of data to be processed, the difference data is obtained, and the corresponding elastic distributed formatting processing node is obtained according to the difference data to perform uniform formatting processing on the multi-source heterogeneous data;

[0026] The formatting processing unit extracts features from the multi-source heterogeneous data that has completed the uniform formatting processing. Among them, a feature extraction algorithm related to enterprise accounts is preset, and the multi-source heterogeneous data is analyzed and processed based on the feature extraction algorithm to obtain the corresponding feature data.

[0027] Further, the process of the data distribution module setting up a distributed computing processing network and allocating tasks for the feature data of multi-source heterogeneous data includes:

[0028] A computing network construction unit and a data distribution unit are set in the data distribution module;

[0029] The computing network construction unit obtains the organizational structure corresponding to the enterprise account and the corresponding organization department, sets up distributed computing nodes according to the organization department; connects the set distributed computing nodes according to the organizational structure to construct a distributed computing processing network;

[0030] The data distribution unit is used to obtain the load data of each distributed computing node in the distributed computing processing network, set up a resource configuration table for the distributed computing processing network according to the load data, and update it in real time; obtain the amount of data of the feature data to be processed in each distributed computing node, and match it with the corresponding distributed computing processing network in the resource configuration table. If the match is successful, the task allocation is completed; if the match is not successful, the amount of data of other distributed computing nodes connected to this distributed computing node in the resource configuration table is obtained for matching until the match is successful and the task allocation is completed; the task allocation result is sent to the data integration module.

[0031] Further, the process of the data integration module calculating the feature data according to the task allocation result includes:

[0032] A data integration unit and a user management unit are set in the data integration module;

[0033] The data integration unit obtains the feature data of the corresponding tasks of each distributed computing node in the distributed computing processing network, and the corresponding distributed computing node analyzes and processes the corresponding feature data to obtain the data to be integrated;

[0034] The data integration unit performs multi-department integration according to the organizational departments corresponding to the data to be integrated to obtain a multi-department data set; the multi-department data set includes integration subsets of different organizational departments; the integration subsets include the data to be integrated of the corresponding organizational departments; the data to be integrated in the integration subsets is integrated to obtain its correlation data, and a correlation threshold is preset. When the correlation data is greater than or equal to the correlation threshold, the organizational department in the integration subset is saved, and a fast integration node is set; the obtained fast integration nodes are stored, and the data integration process is performed by the fast integration nodes.

[0035] Further, the user management unit is used for the temporary integration nodes of the staff in the corresponding enterprise account. Different organizational departments' data to be integrated are selected in the temporary integration nodes according to the needs of the staff, and the integration process of the data to be integrated is selected according to the needs of the staff.

[0036] Compared with the prior art, the beneficial effects of the present invention are:

[0037] 1. By setting up a distributed formatting processing network and a distributed computing processing network to respectively perform formatting processing and calculation and analysis processing on multi-source heterogeneous data, and by setting up a resource configuration table for the distributed computing nodes in the distributed computing processing network to analyze and process the multi-source heterogeneous data that has completed formatting processing, the resource utilization and efficiency in the data integration system are improved to the greatest extent;

[0038] 2. A data monitoring period is preset, and distributed formatting processing nodes are set according to the proportion of the data of the corresponding organizational departments in the corresponding enterprise account in this data integration platform to adjust the distributed formatting processing network, thereby improving the flexibility of the data integration system to a certain extent;

[0039] 3. By the enterprise setting the corresponding organizational structure and the corresponding organizational departments therein, setting the corresponding distributed formatting processing network and distributed computing processing network, and by setting up temporary integration nodes according to the needs of the staff in the enterprise account to analyze and process the data to be integrated, the flexibility of the data integration system is ensured to a certain extent. Description of the Drawings

[0040] Figure 1 This is the schematic diagram of a multi-source heterogeneous data integration system based on distributed computing according to an embodiment of the present application. Detailed implementation manners

[0041] As Figure 1 shown, a multi-source heterogeneous data integration system based on distributed computing includes a data integration platform, in which a data acquisition module, a data extraction module, a data classification module, a data distribution module, and a data integration module are communicatively linked;

[0042] The data integration platform is used to integratively and collaboratively manage relevant data information within an enterprise, so as to obtain data information related to the enterprise's decision-making. Its specific implementation process includes:

[0043] An enterprise verification window is provided in the data integration platform;

[0044] The enterprise verification window is used to obtain enterprise verification information, which includes enterprise name, registration information, organizational structure, and enterprise qualification information. The registration information of the corresponding enterprise verification information is verified to determine whether the corresponding enterprise meets the admission criteria, and a corresponding enterprise account is set for the enterprise verification information that has completed the verification process.

[0045] It should be further noted that in the specific implementation process, the data integration platform sets an associated window set according to the enterprise account and marks the associated window set according to the enterprise account; in addition, the organizational structure includes multiple types, and the multiple types include but are not limited to organizational structures such as financial organizational structure, sales organizational structure, market organizational structure, procurement organizational structure, product organizational structure, and technology trend organizational structure, which are set according to the organizational structure provided by the enterprise account.

[0046] The data acquisition module is used to obtain multi-source heterogeneous data within the corresponding enterprise, perform marking processing on the obtained multi-source heterogeneous data, and send it to the data extraction module. Its specific implementation process includes:

[0047] A data acquisition unit and a load acquisition unit are provided in the data acquisition module;

[0048] The data acquisition unit is used to obtain the organizational structure corresponding to the enterprise account, generate an associated window set according to the organizational structure corresponding to the enterprise account, and the associated window set includes associated sub-windows corresponding to multiple organizational structures; the associated sub-windows are respectively connected to the corresponding organizational structures in the enterprise account, obtain the enterprise data information corresponding to the organizational structure, and mark the associated sub-windows according to the organizational structure; the associated sub-windows obtain the multi-source heterogeneous data corresponding to the corresponding organizational department in the enterprise account, mark the obtained multi-source heterogeneous data according to its organizational department, and send it to the data preprocessing module; the data types of the multi-source heterogeneous data include but are not limited to structured data, semi-structured data, unstructured data, graph data, text data, time series data, etc.

[0049] The structured data is data composed of a fixed format, including corresponding tabular data, etc.

[0050] The semi-structured data is data information with partial structured features, but does not conform to the traditional relational database table form.

[0051] The unstructured data is data information without a clear structure, usually existing in a free form, including but not limited to data information such as text files, images, audio, and video.

[0052] The obtained multi-source heterogeneous data is marked and stored, and the mark includes the corresponding organizational structure type and collection time in the enterprise account.

[0053] The load acquisition unit is used to collect the load data of the distributed nodes in the platform, mark the collected load data according to the type of the distributed nodes, and send it to the data preprocessing module and the data distribution module respectively according to the marking result.

[0054] The data storage module stores it according to the data source of the multi-source heterogeneous data corresponding to the enterprise account. The process of obtaining its historical multi-source heterogeneous data includes:

[0055] The data storage module obtains the organizational structure of the corresponding enterprise account, sets a data storage space set according to the organizational structure of the enterprise account, and the data storage space set is marked according to the enterprise account, including data storage sub-spaces corresponding to the corresponding organizational departments in the organizational structure. The data storage sub-space is used to obtain the historical multi-source heterogeneous data and its format processing results obtained by the corresponding organizational department in the enterprise account, and record them as historical data pairs, and mark and store each historical data pair.

[0056] The data preprocessing module is used to set up a distributed formatting processing network according to the historical multi-source heterogeneous data corresponding to the enterprise account, select corresponding distributed formatting processing nodes according to the load data corresponding to the distributed formatting processing network to perform unified formatting processing on the multi-source heterogeneous data, and obtain the feature data corresponding to the multi-source heterogeneous data. The specific implementation process includes:

[0057] A formatting network construction unit is set in the data preprocessing module;

[0058] The formatting network construction unit is used to set up a corresponding distributed formatting processing network according to the enterprise account, obtain the historical data pairs stored in each data storage subspace in the centralized data storage space corresponding to the enterprise account, and count their data volumes to obtain the data volumes stored in each data storage subspace, and obtain the comprehensive total value of the data volumes stored in the corresponding data storage space set; preset a data monitoring period, obtain the period occupancy ratio of the data volume stored in the corresponding data storage subspace within the data monitoring period, and obtain the average occupancy ratio corresponding to the period occupancy ratio within the data monitoring period; set a fluctuation occupancy ratio interval according to the period occupancy ratio corresponding to each data monitoring period, obtain the lower limit occupancy ratio and the average occupancy ratio of the fluctuation occupancy ratio interval of each data storage subspace, preset a fixed node coefficient, and perform analysis and processing according to the lower limit occupancy ratio, the average occupancy ratio and the fixed node coefficient to obtain the fixed distributed processing node occupancy ratio corresponding to each data storage subspace;

[0059] Allocate fixed distributed formatting processing nodes with corresponding data volumes to each storage subspace according to the fixed distributed node occupancy ratio corresponding to each storage subspace. After the allocation is completed, mark the remaining distributed formatting processing nodes as elastic distributed formatting processing nodes; set the distribution of the fixed distributed formatting processing nodes and the elastic distributed formatting processing nodes as the distributed formatting processing network.

[0060] The process by which the data preprocessing module performs unified formatting processing on the multi-source heterogeneous data according to the distributed formatting processing network includes:

[0061] A formatting processing unit is set in the data preprocessing module;

[0062] The formatting processing unit obtains the formatting processing node load data corresponding to the fixed distributed formatting processing node corresponding to the corresponding data storage subspace in the distributed formatting processing network and the multi-source heterogeneous data collected by the corresponding organization department; obtains the amount of data to be processed in the multi-source heterogeneous data; obtains the idle data according to the formatting processing node load data; compares and analyzes the obtained idle data with the amount of data to be processed. When the idle data is greater than or equal to the amount of data to be processed, the corresponding fixed distributed formatting processing node performs format unification processing on the multi-source heterogeneous data; when the idle data is less than the amount of data to be processed, it obtains the difference data, and obtains the corresponding elastic distributed formatting processing node according to the difference data to perform format unification processing on the multi-source heterogeneous data;

[0063] The formatting processing unit extracts features from the multi-source heterogeneous data that has completed format unification processing. Among them, there is a preset feature extraction algorithm related to enterprise accounts, and based on the feature extraction algorithm, the multi-source heterogeneous data is analyzed and processed to obtain the corresponding feature data.

[0064] The process of the data distribution module setting up a distributed computing processing network and distributing the feature data of the multi-source heterogeneous data includes:

[0065] The data distribution module is provided with a computing network construction unit and a data distribution unit;

[0066] The computing network construction unit obtains the organizational structure corresponding to the enterprise account and the corresponding organization department, sets up distributed computing nodes according to the organization department; connects the set distributed computing nodes according to the organizational structure to construct a distributed computing processing network;

[0067] The data distribution unit is used to obtain the load data of each distributed computing node in the distributed computing processing network, set up a resource configuration table for the distributed computing processing network according to the load data, and update it in real time; obtain the amount of feature data to be processed in each distributed computing node, and match it with the corresponding distributed computing processing network in the resource configuration table. If the match is successful, the task assignment is completed; if the match is not successful, obtain the amount of data of other distributed computing nodes connected to this distributed computing node in the resource configuration table and match it with the feature data until the match is successful and the task assignment is completed; send the task assignment result to the data integration module.

[0068] The process of the data integration module calculating the feature data according to the task assignment result includes:

[0069] The data integration module is provided with a data integration unit and a user management unit;

[0070] The data integration unit obtains the characteristic data of the tasks corresponding to each distributed computing node in the distributed computing processing network, and the corresponding distributed computing node analyzes and processes the corresponding characteristic data to obtain the data to be integrated;

[0071] The data integration unit performs multi-department integration according to the organizational departments corresponding to the data to be integrated to obtain a multi-department data set; the multi-department data set includes integration subsets that are integrated from different organizational departments; the integration subsets include the data to be integrated corresponding to the organizational departments; the data to be integrated in the integration subsets is integrated to obtain its correlation data, and a correlation threshold is preset. When the correlation data is greater than or equal to the correlation threshold, the organizational department in the integration subset is saved, and a fast integration node is set; the obtained fast integration nodes are stored, and the data integration process is performed by the fast integration nodes;

[0072] The user management unit is used to correspond to the temporary integration nodes of the staff in the enterprise account. Different organizational departments' data to be integrated are selected in the temporary integration nodes according to the needs of the staff, and the data to be integrated is integrated according to the needs of the staff.

[0073] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A multi-source heterogeneous data integration system based on distributed computing, including a data integration platform, characterized in that, The communication link of the data integration platform includes a data acquisition module, a data storage module, a data preprocessing module, a data distribution module, and a data integration module; The data acquisition module is used to acquire multi-source heterogeneous data corresponding to the enterprise account and the load data of the distributed nodes, mark them according to the data sources of the multi-source heterogeneous data; The data storage module is used to store the multi-source heterogeneous data according to the data sources corresponding to the enterprise account, and obtain its historical multi-source heterogeneous data; The data preprocessing module is used to set up a distributed formatting processing network according to the historical multi-source heterogeneous data corresponding to the enterprise account, select the corresponding distributed formatting processing node according to the load data corresponding to the distributed formatting processing network to perform format unified processing on the multi-source heterogeneous data, and obtain the characteristic data corresponding to the multi-source heterogeneous data; The data distribution module is used to mark the source of the characteristic data corresponding to the multi-source heterogeneous data, set up a distributed computing processing network, and allocate the characteristic data of the multi-source heterogeneous data according to the source mark and the load data; The data integration module calculates the characteristic data according to the task allocation result to obtain the data to be integrated, sets up a fast integration node for integration processing according to the relationship between the data to be integrated from different data sources, and sets up a temporary integration node according to the enterprise account requirements to perform integration processing on the data to be integrated; The process of the data preprocessing module setting up a distributed formatting processing network according to the historical multi-source heterogeneous data corresponding to the enterprise account includes: A formatting network construction unit is provided in the data preprocessing module; The formatting network construction unit is used to set up the corresponding distributed formatting processing network according to the enterprise account, obtain the historical data pairs stored in each data storage subspace in the data storage space corresponding to the enterprise account, count the data volume thereof, obtain the data volume stored in each data storage subspace, and obtain the comprehensive total value of the data volume stored in the corresponding data storage space set; preset a data monitoring period, obtain the period occupancy ratio of the data volume stored in the corresponding data storage subspace within the data monitoring period, and obtain the average occupancy ratio corresponding to the period occupancy ratio within the data monitoring period; set a fluctuation occupancy ratio interval according to the period occupancy ratio corresponding to each data monitoring period, obtain the lower limit occupancy ratio and the average occupancy ratio of the fluctuation occupancy ratio interval of each data storage subspace, preset a fixed node coefficient, and perform analysis and processing according to the lower limit occupancy ratio, the average occupancy ratio and the fixed node coefficient to obtain the fixed distributed processing node occupancy ratio corresponding to each data storage subspace; Allocate the corresponding fixed distributed formatting processing nodes with the corresponding data volume according to the fixed distributed node occupancy ratio corresponding to each storage subspace. After the allocation is completed, mark the remaining distributed formatting processing nodes as elastic distributed formatting processing nodes; set the distribution of the fixed distributed formatting processing nodes and the elastic distributed formatting processing nodes as the distributed formatting processing network.

2. The multi-source heterogeneous data integration system based on distributed computing according to claim 1, characterized in that The process of the data acquisition module acquiring the multi-source heterogeneous data corresponding to the enterprise account and the load data of the distributed nodes includes: The data integration platform obtains enterprise verification information, which includes the corresponding organizational structure and enterprise qualification information. The organizational structure is the relationship between various organizational departments within the enterprise, including multiple organizational departments, and sets up an enterprise account for the obtained enterprise verification information; A data acquisition unit and a load acquisition unit are set in the data acquisition module; The data acquisition unit is used to obtain the organizational structure corresponding to the enterprise account, generate an associated window set according to the organizational structure corresponding to the enterprise account; the associated window set includes associated sub-windows corresponding to each organizational department; the associated sub-windows obtain the multi-source heterogeneous data corresponding to the corresponding organizational department within the enterprise account, mark the obtained multi-source heterogeneous data according to its organizational department, and send it to the data preprocessing module; The load acquisition unit is used to collect the load data of the distributed nodes in the platform, mark the collected load data according to the type of the distributed nodes, and send it to the data preprocessing module and the data distribution module respectively according to the marking result.

3. The multi-source heterogeneous data integration system based on distributed computing according to claim 2, wherein, The data storage module stores the multi-source heterogeneous data corresponding to the enterprise account according to its data source. The process of obtaining its historical multi-source heterogeneous data includes: The data storage module obtains the organizational structure of the corresponding enterprise account, sets up a data storage space set according to the organizational structure of the enterprise account. The data storage space set is marked according to the enterprise account, and includes data storage sub-spaces corresponding to the corresponding organizational departments within the organizational structure. The data storage sub-spaces are used to obtain the historical multi-source heterogeneous data and its format processing results obtained by the corresponding organizational department within the enterprise account, and record them as historical data pairs, and mark and store each historical data pair.

4. A multi-source heterogeneous data integration system based on distributed computing according to claim 3, characterized in that, The process of the data preprocessing module performing format unified processing on the multi-source heterogeneous data according to the distributed formatting processing network includes: A formatting processing unit is set in the data preprocessing module; The formatting processing unit obtains the formatting processing node load data corresponding to the fixed distributed formatting processing node corresponding to the data storage sub-space in the distributed formatting processing network and the multi-source heterogeneous data collected by the corresponding organizational department; obtains the amount of data to be processed of the multi-source heterogeneous data; obtains the idle data according to the formatting processing node load data; compares and analyzes the obtained idle data with the amount of data to be processed. When the idle data is greater than or equal to the amount of data to be processed, the corresponding fixed distributed formatting processing node performs format unified processing on the multi-source heterogeneous data; when the idle data is less than the amount of data to be processed, it obtains the difference data, and obtains the corresponding elastic distributed formatting processing node according to the difference data to perform format unified processing on the multi-source heterogeneous data; The formatting processing unit performs feature extraction on the multi-source heterogeneous data that has completed format unified processing. There is a preset feature extraction algorithm related to the enterprise account, and the multi-source heterogeneous data is analyzed and processed based on the feature extraction algorithm to obtain its corresponding feature data.

5. A multi-source heterogeneous data integration system based on distributed computing according to claim 4, characterized in that, The data distribution module sets up a distributed computing processing network. The process of task allocation for the feature data of multi-source heterogeneous data includes: The data distribution module is provided with a computing network construction unit and a data allocation unit; The computing network construction unit obtains the organizational structure corresponding to the enterprise account and the corresponding organizational departments, sets up distributed computing nodes according to the organizational departments; connects the set distributed computing nodes according to the organizational structure to construct a distributed computing processing network; The data allocation unit is used to obtain the load data of each distributed computing node in the distributed computing processing network, set up a resource configuration table for the distributed computing processing network according to the load data, and update it in real time; obtain the data volume of the feature data to be processed in each distributed computing node, and match it with the corresponding distributed computing processing network in the resource configuration table. If the match is successful, the task allocation is completed; if the match is not successful, obtain the data volume of the other distributed computing nodes connected to this distributed computing node in the resource configuration table and match it with the feature data until the match is successful and the task allocation is completed; send the task allocation result to the data integration module.

6. A multi-source heterogeneous data integration system based on distributed computing according to claim 5, characterized in that, The process of the data integration module calculating the feature data according to the task allocation result includes: The data integration module is provided with a data integration unit and a user management unit; The data integration unit obtains the feature data corresponding to the tasks of each distributed computing node in the distributed computing processing network, and the corresponding distributed computing node analyzes and processes the corresponding feature data to obtain the data to be integrated; The data integration unit performs multi-department fusion according to the organizational department corresponding to the data to be integrated to obtain a multi-department data set; the multi-department data set includes fusion subsets of different organizational departments; the fusion subsets include the data to be integrated of the corresponding organizational departments; perform integration processing on the data to be integrated in the fusion subsets to obtain its correlation data, preset a correlation threshold, and when the correlation data is greater than or equal to the correlation threshold, save the organizational department in the fusion subset and set up a fast integration node; store the obtained fast integration nodes, and perform data integration processing by the fast integration nodes.

7. A multi-source heterogeneous data integration system based on distributed computing according to claim 6, characterized in that, The user management unit is used to correspond to the temporary integration nodes of the staff in the enterprise account. In the temporary integration nodes, the data to be integrated of different organizational departments are selected according to the needs of the staff, and the integration processing of the data to be integrated is selected according to the needs of the staff.

Citation Information

Patent Citations

  • Multi-source heterogeneous data integration system based on distributed computing

    CN116483840A

  • Big data-based bank-enterprise docking service system and method for science and technology type medium and small-sized enterprises

    CN117992241A