A Data Sharing Method, Device, and Storage Medium for an Enterprise Data Middle Platform Based on Batch Processing

The enterprise data middle-ware method integrates data aggregation, governance, and modeling using Flink and Spark to address data silos and inefficiencies, enhancing data accuracy and efficiency for precise sharing and utilization.

CN115658658BActive Publication Date: 2025-07-15XIAMEN MEIYA PICO INFORMATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211364274.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-02
Publication Date
2025-07-15
Estimated Expiration
2042-11-02

AI Technical Summary

Technical Problem

In the existing technology, there are problems in the process of enterprise digital transformation, such as classification definition of data middle platform construction, redundant technology reissue development, data storage, separation of data and business relationships, data silos and repeated development, resulting in insufficient accuracy of data analysis and low R&D efficiency.

Method used

Using enterprise data middle platform technology based on Flink and Spark, through aggregation, governance and modeling development, the data from multiple source databases are generated and aggregated databases and original databases, and data development is carried out based on the data resource catalog, thematic databases and business databases are built, and data services are provided using APIs to achieve standardization and efficient utilization of data.

Benefits of technology

It improves data accuracy and generation efficiency, simplifies data access and governance processes, improves data utilization efficiency, avoids data packet loss problems, and supports unified analysis and services across business fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115658658B_ABST
    Figure CN115658658B_ABST
Patent Text Reader

Abstract

The present invention provides a data sharing method, device and storage medium for an enterprise data middle platform based on batch processing. The method includes: a convergence step of generating a convergence database for multiple data sources from multiple source databases based on a data resource catalog, and generating corresponding original databases after governing the data sources in the source databases; a processing step of performing data development on the data in the multiple original databases based on the data resource catalog in the convergence database and storing the result data in a subject database or a business database according to business types; and a sharing step of constructing corresponding APIs based on the data table structures in the subject database or the business database, and the enterprise data middle platform using the APIs to provide external data services. The data in the subject database and the business database generated by the present invention is more accurate, and the generation efficiency is high. Moreover, the APIs are constructed based on the data table structures, improving the data utilization efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer databases and big data, and particularly relates to a data sharing method, device and storage medium based on batch processing for an enterprise data middle platform. Background Art

[0002] In the prior art, digital transformation has attracted state-owned enterprises and large enterprises to immediately join this wave, and they have joined the ranks of building data development, hoping to accelerate the transformation of enterprises, subvert the traditional business model and the bottleneck of high-quality sustainable development, and speed up the response speed of the data development of the enterprise R & D team. The main current situation is as follows:

[0003] First, the mainstream technologies such as hadoop, kafka, and ETL are used, and blindly follow the trend to join the ranks of building a data middle platform. Most enterprises understand it as just building a platform, a set of software systems or a standard product;

[0004] Second, the definition of the data middle platform is ambiguous and the classification is chaotic. The concepts and values have not been clarified and defined. In order to pursue the real-time nature of data, the data stream processing method is used, and usually a data warehouse is built;

[0005] Third, blindly pursue the construction of data (set) reports that are disconnected from the business. The analysis results of the data are incorrect, resulting in disagreement with the business decision-making, and the accuracy of the data is insufficient;

[0006] Fourth, the status quo of data islands, business fragmentation and difficult resource allocation cannot be changed, and duplicate development occurs in the R & D team, which is not conducive to improving quality and efficiency, resulting in ineffective investment in development costs.

[0007] Fifth, the data access and governance processes on the enterprise side based on big data components are long and cumbersome, and the access and governance are decentralized and independent, hindering the efficiency of the R & D team.

[0008] That is, in the prior art, through technical implementation methods such as data warehouses and data platforms in the digital transformation practices of each enterprise, there are problems such as chaotic classification and definition, redundant technology and duplicate development, chaotic data storage, disconnection from the business relationship, and data packet loss affecting decision-making analysis. Summary of the Invention

[0009] In view of one or more of the above technical defects in the prior art, the present invention proposes the following technical solutions.

[0010] A data sharing method based on batch processing for an enterprise data middle platform, the method includes:

[0011] A convergence step of generating a convergence database based on a data resource directory for multiple data sources from multiple source databases, and generating corresponding original databases after governing the data sources in the source databases;

[0012] Processing step: After developing the data in the multiple original databases based on the data resource catalog in the aggregation database, the result data is stored in the subject database or business database according to the business type.

[0013] Sharing step: Based on the data table structure in the subject database or business database, the corresponding API is constructed, and the enterprise data middle platform uses the API to provide external data services.

[0014] Furthermore, the operation of the aggregation step is as follows:

[0015] Register the data resource catalog: Configure the data resource metadata; Data item registration. If the data items in the aggregation library are the same as those in the original database, they are automatically generated and do not need to be registered separately. If the data items of other data organizations are used as standard fields, registration is required. The data items that need to be registered include identifiers, field codes, field names, field types, field classifications, and field hierarchies.

[0016] Data governance: Use Flink to simplify task creation through a canvas method, integrate access governance, periodically access data into the aggregation database, and write it into the corresponding original database after governance.

[0017] Furthermore, the operation of the data governance is as follows:

[0018] Select the Flink running environment to create a data aggregation task and configure the running jar package; Obtain the source databases that need to be aggregated. According to business needs, the source databases include: Oracle, MySql, SqlServer, and MongoDB; Orchestrate nodes on the canvas, and the nodes are associated with each other through connections: Read the source database as the starting node, and perform the following configurations on this starting node: Dataset exploration: Configure the data update method: full volume and incremental; When it is full volume, the operation types of data are Insert, Delete, Update (IDU); When it is incremental, there are Insert (I), Insert and Update (IU); Scheduling configuration: Configure the timing strategy of the task; Field exploration: Configure the Chinese names of the primary key fields and comment fields; Read the target database PostgreSql as the aggregation database, configure the table name, register the node to the data resource catalog, and organize the data as the aggregation database; Read the "dataset mapping" operator, select the data resource catalog, and map the fields of the source database to the standard fields of the aggregation database one by one; Read the "format conversion" operator, select the fields to be converted, configure the UDF function, generate FlinkSql according to the configured fields, and call the UdfDateFormat function to process; Read the target database HDFS as the corresponding original database, configure the data table name, and register the node to the data resource catalog. After data governance is completed, the corresponding original database is generated.

[0019] Further, the operation of the processing step is as follows: model and develop the data in the original database, using Spark as the computing engine, providing data source management, model orchestration, and operation scheduling functions; for data source management, it supports selection from the original database and knowledge base database in the resource catalog; for model orchestration, it adopts a visual drag-and-drop method, where users can directly complete the model construction without coding; use Mongo as the data storage node and map it to the subject database or business database in the data resource catalog.

[0020] Further, when constructing the corresponding API, set the permission control field, and when making an API call, judge whether the caller has the permission to call based on this field.

[0021] The present invention also proposes a data sharing device for an enterprise data middle platform based on batch processing, which includes:

[0022] A convergence unit that generates a convergence database based on multiple data sources from multiple source databases based on the data resource catalog, and generates a corresponding original database after governing the data sources in the source databases;

[0023] A processing unit that performs data development on the data in the multiple original databases based on the data resource catalog in the convergence database and stores the result data in the subject database or business database according to the business type;

[0024] A sharing unit that constructs a corresponding API based on the data table structure in the subject database or business database, and the enterprise data middle platform uses the API to provide external data services.

[0025] Further, the operation of the convergence unit is as follows:

[0026] Register the data resource catalog: configure the data resource metadata; for data item registration, if the data items in the convergence library are the same as those in the original database, they will be automatically generated and do not need to be registered separately. If the data items of other data organizations are used as standard fields, registration is required. The data items that need to be registered include identifiers, field codes, field names, field types, field classifications, and field hierarchies;

[0027] Data governance: Adopt Flink, simplify task creation through a canvas method, integrate access governance, periodically access data into the convergence database, and write it into the corresponding original database after governance.

[0028] Further, the operation of the data governance is as follows:

[0029] Select the Flink runtime environment to create a data aggregation task and configure the running jar package; obtain the source databases that need to be aggregated. According to business needs, the source databases include: Oracle, MySql, SqlServer, and MongoDB; arrange nodes on the canvas, and connect the nodes through lines: read the source database as the starting node, and perform the following configurations on this starting node: Dataset exploration: Configure the data update method: full volume and incremental; when it is full volume, the operation types of data are insert, update, and delete (IDU); when it is incremental, there are insert (I), insert and update (IU); Scheduling configuration: Configure the timing strategy of the task; Field exploration: Configure the Chinese names of the primary key field and the comment field; Read the target database PostgreSql as the aggregation database, configure the table name, register the node to the data resource directory, and organize the data into the aggregation database; Read the "dataset mapping" operator, select the data resource directory, and map the fields of the source database to the standard fields of the aggregation database one by one; Read the "grid conversion" operator, select the fields to be converted, configure the UDF function, generate FlinkSql according to the configured fields, and call the UdfDateFormat function for processing; Read the target database HDFS as the corresponding original database, configure the data table name, and register the node to the data resource directory. After data governance, the corresponding original database is generated.

[0030] Furthermore, the operation of the processing unit is: perform modeling and development on the data in the original database. Spark is used as the computing engine for modeling and development, providing data source management, model orchestration, and runtime scheduling functions; Data source management supports selection from the original database and knowledge base database in the resource directory; Model orchestration uses a visual drag-and-drop method, allowing users to directly complete the model construction without coding; Use Mongo as the data storage node and map it to the subject database or business database in the data resource directory.

[0031] Furthermore, when constructing the corresponding API, set the permission control field, and when making an API call, judge whether the caller has the permission to call based on this field.

[0032] The present invention also proposes a computer-readable storage medium, on which computer program code is stored. When the computer program code is executed by a computer, it executes any one of the above methods.

[0033] The technical effect of the present invention is as follows: A method, device, and storage medium for data sharing based on batch processing in an enterprise data middle platform of the present invention. The method includes: a convergence step of generating a convergence database based on a data resource catalog for multiple data sources from multiple source databases, and generating corresponding original databases after governing the data sources in the source databases; a processing step of performing data development on the data in the multiple original databases based on the data resource catalog in the convergence database, and storing the result data in a subject database or a business database according to the business type; a sharing step of constructing corresponding APIs based on the data table structures in the subject database or the business database, and the enterprise data middle platform providing external data services using the APIs. In the present invention, first, multiple data sources from multiple source databases are used to generate a convergence database based on the data resource catalog, and corresponding original databases are generated after governing the data sources in the source databases. Then, data development is performed on the data in the multiple original databases based on the data resource catalog in the convergence database, and the result data is stored in a subject database or a business database according to the business type. Then, corresponding APIs are constructed according to the structures of the data tables to provide services externally. That is, in the present invention, when constructing the subject database or the business database, it is generated based on the original databases that have been data-governed according to the data resource catalog in the convergence database, and the original databases are generated through data governance based on the source databases. Such a generation method makes the data in the generated subject database and business database more accurate and has a high generation efficiency. In the present invention, through the above-mentioned convergence and data governance, it changes the current situation in the prior art where the data access and governance processes are long, cumbersome, and scattered and independent. The technical bottom layer design of the enterprise-side data middle platform is based on the Flink framework. A visual and simplified operation canvas is designed and developed to quickly create tasks, integrating access and governance into one, distributing data to the data resource library on demand and periodically, and introducing the timed scheduling technology of batch processing to avoid the data packet loss problem in stream processing, thereby quickly constructing the data resource catalog. The APIs in the present invention are constructed based on the structures of the data tables, thereby shielding the data sources and the data fetching logic. By API-ifying the data tables, users only need to focus on the query logic of the APIs themselves and do not need to care about infrastructure such as the operating environment, and can provide data APIs and data services for business metric scenarios, improving the efficiency of data utilization. Brief Description of the Drawings

[0034] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, purposes, and advantages of the present application will become more obvious.

[0035] Figure 1 It is a flowchart of a method for data sharing based on batch processing in an enterprise data middle platform according to an embodiment of the present invention.

[0036] Figure 2 It is a structural diagram of a data sharing device based on batch processing in an enterprise data middle platform according to an embodiment of the present invention. Detailed implementation manners

[0037] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the convenience of description, only parts related to the relevant invention are shown in the drawings.

[0038] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0039] Figure 1 A data sharing method based on batch processing in an enterprise data middle platform of the present invention is shown. The method includes:

[0040] A convergence step S101, generating a convergence database for multiple data sources from multiple source databases based on a data resource catalog, and generating corresponding original databases after governing the data sources in the source databases;

[0041] A processing step S102, performing data development on the data in the multiple original databases based on the data resource catalog in the convergence database, and storing the result data in a subject database or a business database according to the business type;

[0042] A sharing step S103, constructing corresponding APIs based on the data table structures in the subject database or the business database, and the enterprise data middle platform using the APIs to provide external data services.

[0043] In the present invention, first, multiple data sources from multiple source databases are used to generate a convergence database based on a data resource catalog, and corresponding original databases are generated after governing the data sources in the source databases. Then, data development is performed on the data in the multiple original databases based on the data resource catalog in the convergence database, and the result data is stored in a subject database or a business database according to the business type. Subsequently, corresponding APIs are constructed according to the data table structures to provide external services. That is, in the present invention, when constructing the subject database or the business database, it is generated from the original databases that have been data-governed based on the data resource catalog in the convergence database. The original databases are generated by data governance based on the source databases. Such a generation method makes the data in the generated subject database and business database more accurate and has a high generation efficiency, which is one of the important inventive points of the present invention.

[0044] The converged database in the present invention can provide data input to the data processing loop when accessing data, while maintaining the originality of the accessed data, the original data, convenient for checking and verifying the quality of the data, and can trace the original appearance; the original database is a data set that retains the original data and can reflect the original business scene, and on this basis, supplements the standardized data generated after a series of processing of various source data. The original library realizes the standardization and value-added of data, provides basic data support for various applications, completes data preparation for data fusion, data abstraction and further value-added, and supports business needs such as information tracing and original scene backtracking; the subject database is established to facilitate work and accurately and quickly reflect the overall picture of the work object. It integrates various types of original data and resource data, and is a public data set of multiple dimensions accumulated over a long period of time around the subject objects that can identify people, finances, materials, organizations, etc., including personnel subject libraries, organization subject libraries, etc. The subject database abstracts the subject objects from a higher level, forming a unified view across business fields, providing a basis for unified analysis and unified services of data, thereby making higher-level use of data value; the knowledge database refers to the knowledge data and rule method collection of the enterprise, including the knowledge data required for data access, processing, governance, organization and service, various rules, methods, process collections, and knowledge data and general algorithms required by various general models in the field. It mainly includes basic knowledge base, basic algorithm base, rule base, etc.; the resource database is a public data collection of key elements (various identification attributes, such as citizen ID number, license plate number, mobile phone number, MAC, etc.) established by integrating various data resources and the association and relationship between elements. It mainly includes: the spatiotemporal distribution of the behavior and content (speech) of elements and elements, the spatiotemporal distribution of the association between elements of the same subject, and the spatiotemporal distribution of the relationship between elements of different subjects. The resource library is public data, which supports various business tasks, can exist independently from any business, and is also related to each business; the business database is a database of business in various professional fields, supporting data of business in various professional fields, recording business processes, and providing data support for various business activities.

[0045] In a further embodiment, the operation of the aggregation step S101 is:

[0046] Register the data resource directory: configure the data resource metadata; register data items. If the data items in the aggregation library are consistent with the original database, they are automatically generated and do not need to be registered separately. If data items of other data organizations are used as standard fields, they need to be registered. The data items that need to be registered include identifiers, field codes, field names, field types, field categories, and field levels.

[0047] Data Governance: Using Flink, simplify task creation through the canvas method, integrate access governance, periodically access and converge data into a database, and write the data into the corresponding original database after governance.

[0048] In a further embodiment, the operation of the data governance is as follows:

[0049] Select the Flink runtime environment to create a data aggregation task and configure the running jar package; obtain the source databases that need to be aggregated. According to business needs, the source databases include: Oracle, MySql, SqlServer, and MongoDB; arrange nodes on the canvas, and connect the nodes through lines: read the source database as the starting node, and perform the following configurations on this starting node: Dataset exploration: Configure the data update method: full volume and incremental; when it is full volume, the operation types of data are Insert, Delete, Update (IDU); when it is incremental, there are Insert (I), Insert and Update (IU); Scheduling configuration: Configure the task's timing policy; Field exploration: Configure the Chinese names of the primary key field and the remarks field; read the target database PostgreSql as the aggregation database, configure the table name, register the node in the data resource directory, and organize the data into the aggregation database; read the "Dataset Mapping" operator, select the data resource directory, and map the fields of the source database to the standard fields of the aggregation database one by one; read the "Grid Conversion" operator, select the fields to be converted, configure the UDF function, generate FlinkSql according to the configured fields, and call the UdfDateFormat function for processing; read the target database HDFS as the corresponding original database, configure the data table name, and register the node in the data resource directory. After the data governance is completed, the corresponding original database is generated.

[0050] In the present invention, through the above-mentioned aggregation and data governance, it changes the current situation in the prior art where the data access and governance processes are long, cumbersome, scattered, and independent. The technical underlying design of the enterprise-side data middle platform is based on the Flink framework. Design and develop a visual and simplified operation canvas that can quickly create tasks, integrate access and governance into one, distribute data to the data resource library on demand and periodically, and introduce the batch processing timing scheduling technology to avoid the data packet loss problem in stream processing, thereby quickly building a data resource directory. Simplify and facilitate the research and development of data access and governance through the drag-and-drop method, and the efficiency can be increased by 60% compared with the current situation. This is another important inventive point of the present invention.

[0051] In a further embodiment, the operation of the processing step S102 is as follows: perform modeling and development on the data in the original database, using Spark as the computing engine, and providing data source management, model orchestration, and operation scheduling functions; for data source management, support selection from the original database and knowledge base database in the resource catalog; for model orchestration, adopt a visual drag-and-drop method, where users do not need to code and can directly complete the model construction by dragging and dropping; use Mongo as the data storage node and map it to the subject database or business database in the data resource catalog. The Spark modeling platform integrates a rich set of operators. During the model orchestration process, users can directly select and configure to meet the computing requirements during the model orchestration process; the operators include filtering, aggregation, intersection, union, difference, join, self-join, deduplication, column calculation, type conversion, column to row conversion, row to column conversion, time processing, value mapping, missing value processing, custom SQL, table structure processing, output, etc.

[0052] The present invention strengthens the construction of the data service system for data tagging and data models through the operators, algorithms, and computing power of the infrastructure, and further abstracts the data model and encapsulates the data service through the close integration with the business department, improving the efficiency of data processing. This is another important inventive point of the present invention.

[0053] In a further embodiment, when constructing the corresponding API, set the permission control field. When making an API call, determine whether the caller has the permission to call based on this field. By developing data APIs, data permission control can be performed according to data classification; through the data interaction platform, provide the data services of the middle platform externally, and managers can monitor the call situations of all APIs to achieve visibility and manageability. Since the APIs in the present invention are constructed based on the structure of the data table, the data source and data fetching logic are shielded. By API-ifying the data table, users only need to focus on the query logic of the API itself and do not need to care about the infrastructure such as the running environment, and can provide data APIs and data services for business metric scenarios. This improves the efficiency of data utilization. This is another important inventive point of the present invention.

[0054] Figure 2 A data sharing device based on batch processing of an enterprise data middle platform of the present invention is shown. The device includes:

[0055] An aggregation unit 201 generates an aggregated database based on multiple data sources from multiple source databases based on the data resource catalog, and generates a corresponding original database after governing the data sources in the source database;

[0056] The processing unit 202 develops the data in the multiple original databases based on the data resource catalog in the convergence database, and stores the result data in the subject database or the business database according to the business type.

[0057] The sharing unit 203 constructs corresponding APIs based on the data table structures in the subject database or the business database, and the enterprise data middle platform uses the APIs to provide external data services.

[0058] In the present invention, multiple data sources from multiple source databases are first used to generate a convergence database based on the data resource catalog, and the data sources in the source databases are governed to generate corresponding original databases. Then, the data in the multiple original databases is developed based on the data resource catalog in the convergence database, and the result data is stored in the subject database or the business database according to the business type. Subsequently, corresponding APIs are constructed according to the data table structures to provide external services. That is, in the present invention, when constructing the subject database or the business database, it is generated from the original databases that have been data-governed based on the data resource catalog in the convergence database, and the original databases are generated by data governance based on the source databases. Such a generation method makes the data in the generated subject database and business database more accurate and has high generation efficiency, which is one of the important inventive points of the present invention.

[0059] The converged database in the present invention can provide data input to the data processing loop when accessing data, while maintaining the originality of the accessed data, the original data, convenient for checking and verifying the quality of the data, and can trace the original appearance; the original database is a data set that retains the original data and can reflect the original business scene, and on this basis, supplements the standardized data generated after a series of processing of various source data. The original library realizes the standardization and value-added of data, provides basic data support for various applications, completes data preparation for data fusion, data abstraction and further value-added, and supports business needs such as information tracing and original scene backtracking; the subject database is established to facilitate work and accurately and quickly reflect the overall picture of the work object. It integrates various types of original data and resource data, and is a public data set of multiple dimensions accumulated over a long period of time around the subject objects that can identify people, finances, materials, organizations, etc., including personnel subject libraries, organization subject libraries, etc. The subject database abstracts the subject objects from a higher level, forming a unified view across business fields, providing a basis for unified analysis and unified services of data, thereby making higher-level use of data value; the knowledge database refers to the knowledge data and rule method collection of the enterprise, including the knowledge data required for data access, processing, governance, organization and service, various rules, methods, process collections, and knowledge data and general algorithms required by various general models in the field. It mainly includes basic knowledge base, basic algorithm base, rule base, etc.; the resource database is a public data collection of key elements (various identification attributes, such as citizen ID number, license plate number, mobile phone number, MAC, etc.) established by integrating various data resources and the association and relationship between elements. It mainly includes: the spatiotemporal distribution of the behavior and content (speech) of elements and elements, the spatiotemporal distribution of the association between elements of the same subject, and the spatiotemporal distribution of the relationship between elements of different subjects. The resource library is public data, which supports various business tasks, can exist independently from any business, and is also related to each business; the business database is a database of business in various professional fields, supporting data of business in various professional fields, recording business processes, and providing data support for various business activities.

[0060] In a further embodiment, the operation of the convergence unit 201 is:

[0061] Register the data resource directory: configure the data resource metadata; register data items. If the data items in the aggregation library are consistent with the original database, they are automatically generated and do not need to be registered separately. If data items of other data organizations are used as standard fields, they need to be registered. The data items that need to be registered include identifiers, field codes, field names, field types, field categories, and field levels.

[0062] Data Governance: Using Flink, simplify task creation through a canvas approach, integrating access and governance. Periodically connect data to a convergence database, and write the governed data to the corresponding original database after governance.

[0063] In a further embodiment, the operation of the data governance is as follows:

[0064] Select a Flink operating environment to create a data aggregation task and configure the running jar package; obtain the source databases that need to be aggregated. According to business needs, the source databases include: Oracle, MySql, SqlServer, and MongoDB; arrange nodes on the canvas, and connect the nodes through lines: read the source database as the starting node, and perform the following configurations on this starting node: Dataset exploration: Configure the data update method: full volume and incremental; when it is full volume, the operation types of data are Insert, Delete, Update (IDU); when it is incremental, there are Insert (I), Insert and Update (IU); Scheduling configuration: Configure the task's timing strategy; Field exploration: Configure the Chinese names of the primary key fields and remarks fields; read the target database PostgreSql as the convergence database, configure the table name, register the node in the data resource directory, and organize the data into the convergence database; read the "Dataset Mapping" operator, select the data resource directory, and map the fields of the source database to the standard fields of the convergence database one by one; read the "Format Conversion" operator, select the fields to be converted, configure the UDF function, generate FlinkSql according to the configured fields, and call the UdfDateFormat function for processing; read the target database HDFS as the corresponding original database, configure the data table name, and register the node in the data resource directory. After the data governance is completed, the corresponding original database is generated.

[0065] In the present invention, through the above-mentioned aggregation and data governance, it has changed the current situation in the prior art where the data access and governance processes are long, cumbersome, scattered, and independent. The technical underlying design of the enterprise-side data middle platform is based on the Flink framework. A visual and simplified operation canvas is designed and developed to quickly create tasks, integrate access and governance, distribute data to the data resource library on demand and periodically, and introduce batch processing's timing scheduling technology to avoid data packet loss problems in stream processing, thereby quickly building a data resource directory. Simplify and facilitate the research and development of data access and governance through a drag-and-drop method, and the efficiency can be increased by 60% compared with the current situation. This is another important inventive point of the present invention.

[0066] In a further embodiment, the operation of the processing unit 202 is as follows: perform modeling and development on the data in the original database, using Spark as the computing engine to provide data source management, model orchestration, and operation scheduling functions; for data source management, support selection from the original database and knowledge base database in the resource catalog; for model orchestration, use a visual drag-and-drop method, where users do not need to code and can directly complete the model construction by dragging and dropping; use Mongo as the data storage node and map it to the subject database or business database in the data resource catalog. The Spark modeling platform integrates a rich set of operators. During the model orchestration process, users can directly select and configure to meet the computing requirements during the model orchestration process; the operators include filtering, aggregation, intersection, union, difference, join, self-join, deduplication, column calculation, type conversion, column to row conversion, row to column conversion, time processing, value mapping, missing value processing, custom SQL, table structure processing, output, etc.

[0067] The present invention strengthens the construction of the data service system for data tagging and data models through the operators, algorithms, and computing power of the infrastructure, and further abstracts the data model and encapsulates the data service through the close integration with the business department, improving the efficiency of data processing. This is another important inventive point of the present invention.

[0068] In a further embodiment, when constructing the corresponding API, set the permission control field, and when making an API call, judge whether the caller has the permission to call based on this field. By developing data APIs, data permission control can be performed according to data classification; through the data interaction platform, provide the data services of the middle platform externally, and the manager can monitor the call situation of all APIs to achieve visibility and manageability. Since the APIs in the present invention are constructed based on the structure of the data table, the data source and data fetching logic are shielded. By API-ifying the data table, users only need to focus on the query logic of the API itself and do not need to care about the infrastructure such as the operating environment, and can provide data APIs and data services for business metric scenarios. This improves the efficiency of data utilization. This is another important inventive point of the present invention.

[0069] For the convenience of description, the above device is described by dividing it into various units according to functions. Of course, when implementing the present application, the functions of each unit can be realized in the same or multiple software and / or hardware.

[0070] From the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the device described in each embodiment or some parts of the embodiments of this application.

[0071] Finally, it should be noted that the above embodiments are only used to illustrate rather than limit the technical solutions of the present invention. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still modifications or equivalent replacements can be made to the present invention, and any modification or partial replacement without departing from the spirit and scope of the present invention shall be covered by the scope of the claims of the present invention.

Claims

1. A data sharing method based on batch processing for an enterprise data middle platform, characterized in that The method includes: A convergence step of generating a converged database based on a data resource catalog from multiple data sources of multiple source databases, and generating corresponding original databases after governing the data sources in the source databases; A processing step of performing data development on the data in the multiple original databases based on the data resource catalog in the converged database, and storing the result data in a subject database or a business database according to business types; The operation of the convergence step is: Registering the data resource catalog: configuring data resource metadata; data item registration. If the data items in the converged database are the same as those in the original database, they are automatically generated and do not require separate registration. If data items from other data organizations are used as standard fields, registration is required. The data items that need to be registered include identifiers, field codes, field names, field types, field classifications, and field hierarchies; Data governance: Using Flink, simplifying task creation through a canvas method, integrating access governance, periodically accessing data into the converged database, and writing it into the corresponding original database after governance; The operation of the data governance is: Selecting a Flink running environment to create a data aggregation task and configuring the running jar package; Obtaining the source databases that need to be converged. According to business needs, the source databases include: Oracle, MySql, SqlServer, and MongoDB; Orchestrating nodes on the canvas, with connections between nodes: Reading the source database as the starting node, and performing the following configurations on this starting node: Dataset exploration: Configuring the data update method: full volume and incremental; When it is full volume, the operation types of the data are insert, update, and delete (IDU); When it is incremental, there are insert (I), insert and update (IU); Scheduling configuration: Configuring the timing policy of the task; Field exploration: Configuring the Chinese names of the primary key fields and comment fields; Reading the target database PostgreSql as the converged database, configuring the table name, registering the node in the data resource catalog, and organizing the data as the converged database; Reading the "dataset mapping" operator, selecting the data resource catalog, and corresponding the fields of the source database with the standard fields of the converged database one by one; Reading the "format conversion" operator, selecting the fields to be converted, configuring the UDF function, generating FlinkSql according to the configured fields, and calling the UdfDateFormat function for processing; Reading the target database HDFS as the corresponding original database, configuring the data table name, and registering the node in the data resource catalog. After the data governance is completed, the corresponding original database is generated; A sharing step of constructing corresponding APIs based on the data table structures in the subject database or the business database, and the enterprise data middle platform using the APIs to provide external data services.

2. The method according to claim 1, wherein The operation of the processing step is: Performing modeling and development on the data in the original database. Spark is used as the computing engine for the modeling and development, providing data source management, model orchestration, and operation scheduling functions; Data source management supports selection from the original database and knowledge base database in the resource catalog; model orchestration uses a visual drag-and-drop method, enabling users to directly drag and drop to complete model construction without coding; MongoDB is used as the data storage node and mapped to the subject database or business database in the data resource catalog.

3. The method according to claim 2, characterized in that When constructing the corresponding API, set the permission control field. When making an API call, determine whether the caller has the permission to call based on this field.

4. An enterprise data middle platform data sharing device based on batch processing, characterized in that, The device includes: A convergence unit that generates a convergence database for multiple data sources from multiple source databases based on the data resource catalog, and generates a corresponding original database after governing the data sources in the source database; the operations of the convergence unit are: Register the data resource catalog: Configure data resource metadata; data item registration. If the data items in the convergence library are the same as those in the original database, they are automatically generated and do not require separate registration. If data items from other data organizations are used as standard fields, registration is required. The data items that need to be registered include identifiers, field codes, field names, field types, field classifications, and field hierarchies. Data governance: Use Flink to simplify task creation through a canvas method, integrating access governance. Periodically access data into the convergence database and write the governed data into the corresponding original database. The operations of the data governance are: Select the Flink operating environment to create a data aggregation task and configure the running jar package; obtain the source databases that need to be aggregated. According to business needs, the source databases include Oracle, MySql, SqlServer, and MongoDB; orchestrate nodes on the canvas, and connect the nodes through lines: Read the source database as the starting node and perform the following configurations on this starting node: Dataset exploration: Configure the data update method: full amount and increment; when it is full amount, the data operation types are insert, update, and delete (IDU); when it is increment, there are insert (I) and insert update (IU); Scheduling configuration: Configure the task timing strategy; Field exploration: Configure the Chinese names of the primary key field and the comment field; Read the target database PostgreSql as the convergence database, configure the table name, register the node to the data resource catalog, and organize the data as the convergence database; Read the "dataset mapping" operator, select the data resource catalog, and map the fields of the source database to the standard fields of the convergence database one by one; Read the "format conversion" operator, select the fields to be converted, configure the UDF function, generate FlinkSql according to the configured fields, and call the UdfDateFormat function for processing; Read the target database HDFS as the corresponding original database, configure the data table name, and register the node to the data resource catalog. After data governance, generate the corresponding original database; A processing unit that performs data development on the data in the multiple original databases based on the data resource catalog in the convergence database and stores the result data in the subject database or business database according to the business type. A shared unit constructs corresponding APIs based on the data table structures in the theme database or business database, and the enterprise data middle platform uses the APIs to provide external data services.

5. The device according to claim 4, characterized in that, The operation of the processing unit is to perform modeling and development on the data in the original database. Spark is used as the computing engine for modeling and development, providing data source management, model orchestration, and operation scheduling functions. Data source management supports selection from the original database and knowledge base database in the resource catalog; for model orchestration, a visual drag-and-drop method is adopted, allowing users to directly complete the model construction by dragging without coding. Mongo is used as the data storage node and is mapped to the theme database or business database in the data resource catalog.

6. A computer-readable storage medium stores computer program code thereon, and when the computer program code is executed by a computer, it executes the method according to any one of claims 1 - 3 above.

Citation Information

Patent Citations

  • Emergency resource pool system suitable for emergency management and implementation method thereof

    CN114637799A

  • Generalized data warehouse

    CN114647716A