Distributed computing and comparing system for heterogeneous source data
Through the heterogeneous source data distributed computing and comparison system, the problem of data silos and data interaction difficulties is solved, data integration and efficient comparison and analysis are realized, and powerful data analysis support is provided.
Patent Information
- Application Number
- CN202411972897.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, data island phenomenon is serious, and data interaction is difficult, making it difficult to conduct unified query and comparison analysis.
It provides a distributed computing comparison system for heterogeneous source data, including a big data distributed query engine, a business rule management module, a distributed computing comparison module and a data analysis module, through which data integration, comparison and analysis are realized.
It realizes seamless integration and efficient processing of data from different sources, formats, and semantics, provides efficient and accurate data comparison and analysis capabilities, supports complex data analysis needs, and provides users with a unified query interface and efficient data comparison services.
Smart Images

Figure CN120011419A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular to a distributed computing and comparison system for heterogeneous source data. Background Art
[0002] In the medical field, with the advent of the big data era, the diversity of data sources and the surge in data volume have put forward higher requirements for data processing and analysis. Due to the diversity of data sources, the data structure, query method, storage format, access interface and semantic differences between different data sources have led to serious data silos, making it difficult to interact with data and conduct unified queries and comparative analysis. Summary of the invention
[0003] The technical problem to be solved by the present disclosure is to overcome the defects of the prior art, such as the serious data island phenomenon, the difficulty in data interaction, and the difficulty in unified query and comparison analysis, and to provide a distributed computing comparison system for heterogeneous source data.
[0004] The present invention solves the above technical problems through the following technical solutions:
[0005] The present disclosure provides a heterogeneous source data distributed computing comparison system, which comprises: a big data distributed query engine, a business rule management module, a distributed computing comparison module and a data analysis module;
[0006] The big data distributed query engine is used to obtain data by querying large data sets distributed on one or more external heterogeneous data sources;
[0007] The business rule management module is used to manage business rules through an integrated rule engine;
[0008] The distributed computing comparison module is used to integrate the rule engine and the engineering project, define the business logic rules of data calculation and comparison in the rule engine, and execute the defined rules in the engineering project;
[0009] The data analysis module is used to integrate data from different data sources according to business needs, provide a data basis, and generate data comparison results based on a data comparison and analysis algorithm corresponding to the business rules.
[0010] Optionally, the heterogeneous source data distributed computing comparison system further includes: a heterogeneous data source access module;
[0011] The heterogeneous data source access module is used to configure the data source required for the project in advance in the catalog directory before the rule engine is started, and edit it into the properties file of the corresponding data source so that the connection information of the data source required for the project can be obtained when the rule engine is started.
[0012] Optionally, the heterogeneous source data distributed computing comparison system further includes: a data preprocessing module;
[0013] The data preprocessing module is used to perform at least one preprocessing operation of data cleaning, data integration, data transformation and data reduction on the acquired data, and store the preprocessed data in a preset format in an internal database.
[0014] Optionally, the heterogeneous source data distributed computing comparison system further includes: a data visualization module;
[0015] The data visualization module is used to display the data comparison results in an intuitive graphical manner.
[0016] Optionally, the distributed computing comparison module is also used to automatically generate an execution plan for the engineering project according to the data computing comparison logic configured by the rule engine and the relevant configuration of the required data source.
[0017] Optionally, the distributed computing comparison module is further used to synchronize data to the engineering project through the rule engine in batch processing or streaming processing for the execution plan according to the configuration information of the relevant rules;
[0018] The distributed computing comparison module is also used to perform data computing comparison analysis on the engineering project in accordance with the rule logic of the rule engine and generate corresponding results.
[0019] Optionally, the big data distributed query engine adopts Trino;
[0020] And / or, the rule engine adopts Drools.
[0021] Optionally, the data analysis module is also used to complete setting the data comparison and analysis algorithm according to the business rules through a machine learning algorithm library.
[0022] Optionally, the machine learning algorithm library adopts the Weka library in Java.
[0023] Optionally, the data preprocessing module integrates the deep learning library Deeplearning4j and the Apache Commons series library in Java.
[0024] On the basis of being in accordance with the common sense in the art, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present disclosure.
[0025] The positive progressive effects of the present disclosure are: through the integration of microservice architecture, business rule management, big data processing, distributed query and heterogeneous data source integration, seamless integration and efficient processing of data from different sources, different formats and different semantics are achieved, efficient and accurate data comparison and analysis capabilities are achieved, and diverse data sources can be seamlessly integrated and processed, providing strong support for complex data analysis needs, and providing users with a unified query interface and efficient data comparison services; at the same time, it deeply integrates data analysis and cutting-edge visualization technology to present the full picture of data in an intuitive and dynamic way, enabling decision makers to quickly gain insight into business dynamics and accurately capture market trends, thereby formulating more efficient and practical strategies and accelerating the intelligent process of business decision-making and execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 A module diagram of a heterogeneous source data distributed computing and comparison system provided as an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] The present disclosure is further described below by way of examples, but the present disclosure is not limited to the scope of the examples.
[0028] Prefixes such as "first" and "second" are used in the embodiments of the present disclosure only to distinguish different description objects, and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of prefixes such as ordinal numbers to distinguish description objects in the embodiments of the present disclosure does not constitute a limitation on the described objects. For the statement of the described objects, please refer to the description in the context of the embodiments, and no unnecessary limitation should be constituted due to the use of such prefixes. In addition, in the description of the present embodiment, unless otherwise specified, the meaning of "plurality" is two or more.
[0029] Figure 1 A module diagram of a distributed computing and comparison system for heterogeneous source data is provided as an exemplary embodiment of the present disclosure. The distributed computing and comparison system for heterogeneous source data includes: a big data distributed query engine 1, a business rule management module 2, a distributed computing comparison module 3 and a data analysis module 4.
[0030] The big data distributed query engine 1 is used to obtain data from large data sets distributed on one or more external heterogeneous data sources by querying.
[0031] The business rule management module 2 is used to manage business rules through an integrated rule engine.
[0032] The distributed computing comparison module 3 is used to integrate the rule engine and the engineering project, define the business logic rules of data computing and comparison in the rule engine, and execute the defined rules in the engineering project.
[0033] The data analysis module 4 is used to integrate data from different data sources according to business requirements, provide a data basis, and generate data comparison results based on a data comparison analysis algorithm corresponding to business rules.
[0034] Among them, this system integrates the SpringCloud framework to achieve the flexibility and scalability of the microservice architecture, integrates the big data distributed query engine to accelerate the performance of complex data queries; integrates the rule engine to support flexible configuration and automated decision-making of business logic; at the same time, by introducing machine learning algorithm libraries, deep learning libraries and some third-party libraries, data processing and in-depth trend comparison analysis of data can be realized; in addition, the data comparison and analysis results can be visualized through the front-end technical solution, providing users with high-precision data comparison results display.
[0035] In this embodiment, through the integration of microservice architecture, business rule management, big data processing, distributed query and heterogeneous data source integration, seamless integration and efficient processing of data from different sources, different formats and different semantics are achieved, and efficient and accurate data comparison and analysis capabilities are achieved. It can seamlessly integrate and process diverse data sources, provide strong support for complex data analysis needs, and provide users with a unified query interface and efficient data comparison services; at the same time, it deeply integrates data analysis and cutting-edge visualization technology to present the full picture of data in an intuitive and dynamic way, so that decision makers can quickly understand business dynamics and accurately capture market trends, thereby formulating more efficient and practical strategies and accelerating the intelligent process of business decision-making and execution.
[0036] In one embodiment, the heterogeneous source data distributed computing and comparison system further includes: a heterogeneous data source access module 5 .
[0037] The heterogeneous data source access module 5 is used to configure the data source required by the project in advance in the catalog directory before the rule engine is started, and edit it into the properties file of the corresponding data source so that the connection information of the data source required by the project can be obtained when the rule engine is started.
[0038] Among them, when using a big data distributed query engine (such as Trino), it is necessary to configure the data source required for the project in advance in the catalog directory before starting, and edit it into the properties file of the corresponding data source, so that when the big data distributed query engine is started, the connection information of the data source required for the project can be obtained. However, in the process of actual application, due to various factors, the data source information is not static, but needs to be changed according to the current actual scenario, so it is required to dynamically maintain the data source information in the big data distributed query engine. In this embodiment, the big data distributed query engine is secondary developed, the catalog directory is uniformly maintained and stored through the Database, the Restful API interface is added to uniformly manage the catalog directory, and the CoordinatorDynamicCatalogManager file is modified at the same time to realize the scheduled loading and updating of the catalog configuration. Finally, the dynamic update of the catalog configuration file is completed without restarting the big data distributed query engine.
[0039] In one embodiment, the heterogeneous source data distributed computing and comparison system further includes: a data preprocessing module 6 .
[0040] The data preprocessing module 6 is used to perform at least one preprocessing operation of data cleaning, data integration, data transformation and data reduction on the acquired data, and store the preprocessed data in an internal database in a preset format.
[0041] Given the diversity and complexity of data sources, data from different sources often have significant differences in storage format, data semantics, and effective data. In order to ensure data consistency and availability, in actual application scenarios, a series of preprocessing steps are performed on the original data, mainly including data cleaning, data integration, data transformation, and data reduction.
[0042] Specifically, this system can introduce deep learning libraries in the Java field (such as Deeplearning4j) and reusable, open source Java code component libraries (such as the Apache Commons series of libraries), which can efficiently handle various common data processing needs such as strings, mathematical operations, and file operations.
[0043] In addition, in order to further improve the flexibility and performance of data processing, this system integrates a big data distributed query engine (such as Trino), which allows us to quickly integrate and transform data from different data sources in the form of SQL queries, greatly simplifying the data preprocessing process. Through the big data distributed query engine, this system can easily achieve data access, query optimization, and data integration across data sources.
[0044] This system achieves efficient preprocessing of complex data sources by comprehensively using data processing tools and libraries such as Java libraries (such as Deeplearning4j, Apache Commons) and big data distributed query engines, ensuring the accuracy, consistency and availability of data, and laying a solid foundation for subsequent data analysis.
[0045] In one embodiment, the heterogeneous source data distributed computing and comparison system further includes: a data visualization module 7 .
[0046] The data visualization module 7 is used to display the data comparison results in an intuitive graphical manner.
[0047] Among them, the data visualization module 7 can use the front-end technical solutions of Vue and ECharts to realize the visualization of data comparison and analysis results, and provide users with high-precision data comparison results display.
[0048] In one embodiment, the distributed computing comparison module 3 is also used to automatically generate an execution plan for an engineering project according to the data computing comparison logic configured by the rule engine and the relevant configuration of the required data source.
[0049] Among them, by integrating the rule engine (such as Drools) and the engineering project, the business logic rules for data calculation and comparison can be accurately defined in the rule engine, and the specific rule implementation is completed by the engineering project. The engineering project automatically generates an execution plan based on the data calculation and comparison logic configured by the rule engine and the relevant configuration of the required data source.
[0050] These execution plans can be configured according to the rules, and can synchronize data to the engineering project in batch processing or streaming mode using a big data distributed query engine (such as Trino). Subsequently, the engineering project follows the rule logic in the rule engine, performs data calculation comparison and analysis, and generates corresponding results. This process aims to optimize the data processing flow and improve the flexibility and efficiency of business logic processing.
[0051] In one embodiment, the distributed computing comparison module 3 is also used to synchronize data to the engineering project through the rule engine in batch processing or streaming processing according to the configuration information of the relevant rules for the execution plan.
[0052] The distributed computing comparison module 3 is also used to perform data computing comparison analysis on engineering projects in accordance with the rule logic of the rule engine and generate corresponding results.
[0053] In one embodiment, the big data distributed query engine 1 adopts Trino.
[0054] Among them, Trino (formerly PrestoSQL) is used as a big data distributed query engine to accelerate complex data query performance.
[0055] The Trino big data distributed query engine technology solves various complex data query requirements for large data sets distributed on one or more heterogeneous data sources. This technology has the following advantages:
[0056] High performance: Trino uses full in-memory computing, an internal pipeline-based model, and parallel data query technology, making the query speed very fast.
[0057] Distributed architecture: Trino uses a Master-Slave architecture, where a coordinator is responsible for task generation and scheduling, and multiple workers are responsible for task execution. This architecture enables Trino to process massive amounts of data in parallel and improve query efficiency.
[0058] Flexible data source connection: Trino supports various data source connections, including Oracle, MySQL, Elasticsearch, Kafka, Redis, etc., and can perform associated queries between different data sources, simplifying the data analysis process.
[0059] The big data distributed query engine can not only query and obtain the required data from external data sources, but also query and obtain data from the internal database of the heterogeneous source data distributed computing and comparison system.
[0060] For the distributed computing and comparison system of heterogeneous source data, Trino is an in-memory computing engine that processes data in memory and can quickly process large amounts of data to meet the needs of real-time query and analysis.
[0061] Trino uses SQL language for data query, which reduces the learning cost and usage threshold.
[0062] Trino can process multiple data sources, including Hadoop, relational databases, NoSQL databases, etc., making it easier for users to obtain data from different sources.
[0063] Trino significantly enhances the ability to query associations across different data sources, greatly simplifies the process of complex data analysis, and not only improves data processing efficiency, but also significantly reduces the workload of R&D personnel in the data integration and preprocessing stages. Through its flexible architecture design, Trino makes it more efficient and intuitive to seamlessly extract, integrate and analyze data from multiple heterogeneous data sources.
[0064] In one embodiment, the rule engine employs Drools.
[0065] Among them, the integrated Drools rule engine supports flexible configuration and automated decision-making of business logic.
[0066] The integration of the Drools rule engine has achieved the separation of business decision logic from application code, and flexibly managed through rules. This transformation has greatly simplified the maintenance of the system, allowing developers to focus on optimizing the core performance of data analysis algorithms, while business analysts can focus more on accurately defining and maintaining business rules, thereby significantly improving the efficiency and focus of team collaboration.
[0067] In addition, the Drools engine has built-in advanced optimization algorithms and efficient data structures. These design advantages ensure that the rule matching process can be executed quickly and accurately in large-scale data analysis tasks, greatly improving the efficiency of data analysis and the accuracy of results. Whether facing complex and changing business scenarios or processing massive amounts of data, Drools can demonstrate its excellent performance and stability, providing solid technical support for the efficient operation of the system and accurate decision-making.
[0068] In one embodiment, the data analysis module 4 is also used to complete the setting of the data comparison and analysis algorithm according to the business rules through the machine learning algorithm library.
[0069] Among them, when the big data distributed query engine integrates data from different data sources and provides a data foundation according to business needs, the machine learning algorithm library can be used to complete the setting of data comparison and analysis algorithms according to different business rules.
[0070] Based on comparing two or more data sets. By measuring the overall size of the data, the sizes of different data sets or data items are compared to understand their quantitative differences; by measuring the overall fluctuation of the data, the volatility of the data is analyzed, including cyclical fluctuations, fluctuations affected by internal business factors, fluctuations affected by external factors, etc., to reveal the laws and trends behind the data. By measuring the trend of data changes, the trend of data changes is compared and analyzed from the two dimensions of time and space; in this way, the differences, trends and correlations are revealed, and efficient data comparison services are provided to users.
[0071] A business data comparison and analysis algorithm system can be built based on machine learning algorithm libraries (such as Weka) and deep learning libraries (such as Deeplearning4j). While supporting simple data comparison and analysis, it can further explore and analyze the future development trends of the data, the correlation between the data, and the hidden information in the data through the use of algorithms such as neural networks and multivariate regression analysis.
[0072] In one embodiment, the machine learning algorithm library uses the Weka library in Java.
[0073] Among them, by introducing the Weka machine learning algorithm library, data processing and in-depth trend comparative analysis of data are realized.
[0074] In one embodiment, the data preprocessing module 6 integrates the deep learning library Deeplearning4j and the Apache Commons series library in Java.
[0075] Among them, by introducing the Deeplearning4j deep learning library and some third-party libraries, data processing and in-depth trend comparative analysis of data are further realized.
[0076] In one embodiment, the data visualization module 7 uses the Vue3 framework and the ECharts chart library as the core of the front end.
[0077] In order to provide users with an intuitive and clear display of data comparison results, the data visualization module uses the Vue3 framework and the ECharts chart library as the core of the front-end technical solution. Vue.js, with its responsive data binding and component-based development advantages, provides strong support for the dynamic display of data analysis results; while ECharts, with its rich chart types and highly customizable features, ensures the accuracy and beauty of data visualization. The perfect combination of the two enables this system to efficiently present complex data analysis results, making it easier for users to find errors and anomalies in the data, and providing strong support for subsequent data analysis and decision-making.
[0078] The following is an example of implementing a distributed computing comparison system for heterogeneous source data, in which the big data distributed query engine uses Trino, the rule engine uses Drools, the machine learning algorithm library uses the Weka library in Java, the data preprocessing module integrates the deep learning library Deeplearning4j in Java and the Apache Commons series library, and the data visualization module uses the Vue3 framework and the ECharts chart library as the core of the front end.
[0079] Step 1: Perform secondary development on Trino, maintain and store the catalog directory uniformly through the Database, add a Restful API interface to uniformly manage the catalog directory, dynamically update the configuration file, and modify the CoordinatorDynamicCatalogManager class to implement scheduled loading and updating of the catalog configuration.
[0080] Step 2: Deploy and build Trino, import the Trino driver into the SpringCloud framework, set the Trinoca certificate and trust the certificate, create a yml file, configure the Trino data source connection information, and complete the SpringCloud integration of Trino.
[0081] Step 3. Implement the data source management function in the system. Users can change the data source through the interface. The system will synchronize the data source change information to the Trino server-side Restful API interface in real time. The interface will dynamically update the change information to the configuration file in the catalog directory and call the method in the CoordinatorDynamicCatalogManager class to implement scheduled loading and updating of the catalog configuration, and complete the dynamic update of the Trino data source configuration.
[0082] Step 4: Integrate the Drools rule engine into the system to separate business decisions from code. The code is mainly responsible for the data processing and data analysis algorithm implementation required in the business rules. The specific rule logic is stored in the rule library, and the corresponding rule logic is executed by writing LHS and RHS to complete the data comparison and analysis process, and the results are persistently stored.
[0083] Step 5. Through the Vue3 framework and ECharts, provide users with a rule selection and configuration page. Through the business rules and data processing options provided by the system, complete the configuration of business rules and their execution plans, and visualize the final data comparison and analysis results.
[0084] In general, the heterogeneous source data distributed computing comparison system in this embodiment can have the following advantages:
[0085] (1) High performance: In this study, compared with data loading through traditional IO streams, using Trino for data processing undoubtedly has significant performance advantages. Trino performs data processing based on memory and can quickly process large amounts of data in a short period of time, saving time costs.
[0086] (2) Distributed query: Different from the traditional single-node processing method, Trino distributes the memory pressure to each sub-node through distributed deployment, which greatly reduces the pressure on the main service. In addition, Trino can process massive data in parallel, which greatly improves the query efficiency.
[0087] (3) Seamless connection with various data sources: In previous projects, due to different actual environments, the data sources provided often differed. Usually, secondary development was passively carried out on unsupported data sources in a demand-driven manner to adapt to the data sources provided in the actual environment, which cost a lot of money. In this study, Trino's own advantage of supporting connections to various data sources greatly saved the development time of R&D personnel. At the same time, to a certain extent, the system's adaptability to different actual scenarios was further improved.
[0088] (4) Cross-data source association query: In the past, when processing data between different data sources, the usual practice was to load the required data from different data sources into memory and process it step by step through the program. This approach is time-consuming and consumes a lot of resources. In this study, by using Trino to support association queries between different data sources, the data processing process is greatly simplified. When querying different data sources, a certain degree of data analysis and processing can be performed directly through query statements, which effectively improves data processing efficiency and saves a lot of costs.
[0089] (5) Dynamic update of data source: Since Trino does not support dynamic management of data sources, it needs to be restarted each time after configuring the data source for the data source to take effect. To solve this problem, in this study, Trino was redeveloped to achieve dynamic management of Trino data sources by dynamically managing the catalog, which greatly simplified the configuration of Trino data sources.
[0090] (6) Business rule management: In most projects, the project code contains both business logic and its specific functional implementation, which makes the project code bloated and difficult to maintain. In this study, Drools was introduced to separate the business decision logic from the application code, allowing developers to focus only on the implementation of code functions and the improvement of code performance, while business personnel can flexibly manage the business decision process in the form of rules.
[0091] (7) Data analysis: In this study, a business data comparison and analysis algorithm system was built based on the Weka machine learning algorithm library and the Deeplearning4j deep learning library. While supporting simple data comparison and analysis, it further explored and analyzed the future development trends of the data, the correlation between the data, and the hidden information in the data by using algorithms such as neural networks and multivariate regression analysis.
[0092] Although the specific embodiments of the present disclosure are described above, those skilled in the art should understand that this is only an example, and the protection scope of the present disclosure is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present disclosure, but these changes and modifications all fall within the protection scope of the present disclosure.
Claims
1. A distributed computing comparison system for heterogeneous source data, characterized in that: The heterogeneous source data distributed computing comparison system includes: a big data distributed query engine, a business rule management module, a distributed computing comparison module and a data analysis module; The big data distributed query engine is used to obtain data by querying large data sets distributed on one or more external heterogeneous data sources; The business rule management module is used to manage business rules through an integrated rule engine; The distributed computing comparison module is used to integrate the rule engine and the engineering project, define the business logic rules of data calculation and comparison in the rule engine, and execute the defined rules in the engineering project; The data analysis module is used to integrate data from different data sources according to business needs, provide a data basis, and generate data comparison results based on a data comparison and analysis algorithm corresponding to the business rules.
2. The heterogeneous source data distributed computing comparison system according to claim 1, characterized in that: The heterogeneous source data distributed computing comparison system also includes: a heterogeneous data source access module; The heterogeneous data source access module is used to configure the data source required for the project in advance in the catalog directory before the rule engine is started, and edit it into the properties file of the corresponding data source so that the connection information of the data source required for the project can be obtained when the rule engine is started.
3. The heterogeneous source data distributed computing comparison system according to claim 1, characterized in that: The heterogeneous source data distributed computing comparison system also includes: a data preprocessing module; The data preprocessing module is used to perform at least one preprocessing operation of data cleaning, data integration, data transformation and data reduction on the acquired data, and store the preprocessed data in a preset format in an internal database.
4. The heterogeneous source data distributed computing comparison system according to claim 1, characterized in that: The heterogeneous source data distributed computing comparison system also includes: a data visualization module; The data visualization module is used to display the data comparison results in an intuitive graphical manner.
5. The heterogeneous source data distributed computing comparison system according to claim 1, characterized in that: The distributed computing comparison module is also used to automatically generate an execution plan for the engineering project according to the data computing comparison logic configured by the rule engine and the relevant configuration of the required data source.
6. The heterogeneous source data distributed computing comparison system according to claim 5, characterized in that: The distributed computing comparison module is also used to synchronize data to the engineering project through the rule engine in batch processing or streaming processing for the execution plan according to the configuration information of the relevant rules; The distributed computing comparison module is also used to perform data computing comparison analysis on the engineering project in accordance with the rule logic of the rule engine and generate corresponding results.
7. The heterogeneous source data distributed computing comparison system according to claim 1, characterized in that: The big data distributed query engine adopts Trino; And / or, the rule engine adopts Drools.
8. The heterogeneous source data distributed computing comparison system according to claim 1, characterized in that: The data analysis module is also used to complete the setting of the data comparison and analysis algorithm according to the business rules through a machine learning algorithm library.
9. The heterogeneous source data distributed computing comparison system according to claim 8, characterized in that: The machine learning algorithm library uses the Weka library in Java.
10. The heterogeneous source data distributed computing comparison system according to claim 3, characterized in that: The data preprocessing module integrates the deep learning library Deeplearning4j and the Apache Commons series library in Java.