Stream-batch integrated master data management method and device based on multi-source and multi-database
Through the multi-source and multi-database integrated batch and stream master data governance method, the data island problem is solved, the standardization and consistency of data are achieved, the data processing efficiency and organizational collaboration are improved, and the interaction difficulties between cross-departmental collaborative business systems and the inconsistent management of multiple data are solved.
Patent Information
- Application Number
- CN202311136684.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-05
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2043-09-05
AI Technical Summary
In existing technologies, the data silo phenomenon in enterprise business systems leads to difficulties in data search, selection and application, and cross-departmental collaborative business systems are difficult to interact with. The management of multiple data is inconsistent, making it difficult to meet the real-time, accuracy and consistency requirements of massive data processing.
A multi-source and multi-database integrated stream and batch master data governance method is adopted. Through a hub-and-spoke data collection model, combined with streaming and batch processing technologies, data cleaning and cross-validation are performed, a master data analysis matrix is established, and conceptual, logical and physical data models are constructed to ensure data consistency and availability.
It has achieved standardized, stable and easy-to-use data, improved data relevance and availability, eliminated data redundancy, ensured the consistency and efficiency of master data processing, and enhanced the strategic synergy of the organization.
Smart Images

Figure CN117312268B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data, and in particular to a method and device for managing master data based on multi-source and multi-database integration of stream and batch. Background Art
[0002] With the rapid development of the internet and mobile internet, enterprise businesses are expanding, and more and more business systems are being built. Data volumes are rapidly expanding, but are being stored in the form of "data silos." At the same time, with the intersection of business systems, some common data will appear and be called upon by different business systems. The calling relationships are becoming increasingly complex, leading to problems such as "difficulty in finding, selecting, and applying data." To address this issue, the concept of master data has gradually emerged. Master data refers to the basic information of an organization that reflects the status attributes of core business entities and meets the needs of cross-departmental business collaboration. Master data transcends departments, processes, themes, systems, and technologies, providing a basis and standard to follow when integrating and sharing disparate information systems.
[0003] Industries across the globe are accumulating vast amounts of business data, and the demand for data is increasing. This has led to a continuous evolution of big data processing architectures, from traditional offline architectures to Lambda, Kappa, and integrated batch and stream architectures to address data link redundancy and inconsistent data caliber. However, difficulties persist in interoperating between multi-departmental and cross-departmental collaborative business systems, inconsistent data management across multiple sources, and the need for real-time, accurate, and consistent processing of massive amounts of data. Therefore, a master data governance approach is urgently needed to address these technical challenges. Summary of the Invention
[0004] In response to the technical problems mentioned above, the embodiment of the present application aims to propose a multi-source and multi-database-based stream-batch integrated master data management method and device to solve the technical problems mentioned in the above background technology section.
[0005] In a first aspect, the present invention provides a master data management method based on multi-source and multi-database integrated stream and batch, comprising the following steps:
[0006] Obtaining raw data from at least one data source among the multiple data sources;
[0007] Perform data cleaning on the original data to obtain standardized data;
[0008] Build a master data analysis matrix based on standardized data, and identify and obtain basic data based on the master data analysis matrix;
[0009] Define the master data table structure according to data requirements and data standards, build a conceptual data model based on basic data, expand the logical data model based on the conceptual data model and the master data table structure, establish a physical data model based on the logical data model, cross-validate the data generated by the physical data model to obtain the master data.
[0010] Preferably, obtaining original data from at least one of the multiple data sources specifically includes:
[0011] The original data of at least one data source from different data sources is collected and aggregated through a hub-and-spoke data collection mode. The different data sources include relational databases, message queue databases, column storage databases, object-oriented databases and file systems. Among them, relational databases, column storage databases and object-oriented databases adopt batch processing or streaming processing for collection, message queue databases adopt streaming processing for collection, and file systems adopt batch processing for collection.
[0012] Preferably, the data source is evaluated in the following ways, including:
[0013] Define the dimensions and indicators of data source quality analysis based on the business rules and data conditions of the data source. The dimensions of data source quality analysis include technical indicators and business indicators. Technical indicators include data integrity, data uniqueness, data validity, data timeliness, and data rationality; business indicators include data consistency, data authenticity, data accuracy, data readability, and data availability.
[0014] Perform quality analysis on the data source to determine whether the data in the data source meets the master data requirements.
[0015] Preferably, the raw data is cleaned to obtain standardized data, specifically including:
[0016] The original data is filtered, deduplicated, and format converted to obtain standardized data.
[0017] Preferably, a master data analysis matrix is constructed based on the standardized data, and basic data is obtained by analyzing and identifying the master data analysis matrix, specifically including:
[0018] Based on the five indicators of business impact, data sharing, master data management maturity, master data unification difficulty, and demand urgency, the hierarchical analysis method is used to establish a master data analysis matrix for standardized data;
[0019] The weight of each indicator is determined using the geometric mean method based on the main data analysis matrix;
[0020] The weight is compared with the threshold to determine whether the standardized data is basic data within the scope of master data management.
[0021] As a preferred method, the master data table structure is defined according to data requirements and data standards, a conceptual data model is constructed based on basic data, a logical data model is obtained by expanding the conceptual data model and the master data table structure, a physical data model is established based on the logical data model, and the data generated by the physical data model is cross-validated to obtain the master data, specifically including:
[0022] Define the attributes of the fields in the master data table structure, including field type, field dictionary, value range, and data rules;
[0023] Establish a conceptual data model based on the relationship between entities and entities;
[0024] The conceptual data model is formed by merging similar data, adding related entities, adjusting attributes, designing paradigms, removing data duplication, making primary keys unique, and calculating credibility to form a logical data model;
[0025] Abstract logic into entities through logical data models and define the time and cycle of model execution;
[0026] Through system verification, duplication checking, manual comparison, screening and verification, the integrity, standardization, uniqueness, consistency, accuracy and validity of the data generated by the physical data model are cross-validated in multiple dimensions to obtain the master data.
[0027] As an option, it also includes:
[0028] Determine the distribution method of master data, which includes passive distribution, active distribution, and interactive distribution.
[0029] In a second aspect, the present invention provides a master data management device based on multi-source and multi-database integration of stream and batch, comprising:
[0030] a data acquisition module, configured to acquire raw data from at least one data source among a plurality of data sources;
[0031] A data cleaning module is configured to clean the original data to obtain standardized data;
[0032] A master data analysis module is configured to construct a master data analysis matrix based on standardized data, and obtain basic data through analysis and identification based on the master data analysis matrix;
[0033] The data modeling module is configured to define the master data table structure according to data requirements and data standards, build a conceptual data model based on basic data, expand the logical data model based on the conceptual data model and the master data table structure, establish a physical data model based on the logical data model, cross-validate the data generated by the physical data model, and obtain the master data.
[0034] In a third aspect, the present invention provides an electronic device comprising one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0035] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] (1) The master data governance method based on multi-source and multi-database stream-batch integration proposed in the present invention combines the characteristics of multiple sources, multiple databases and the authority, globality and scalability of master data, selects a hub-and-spoke collection mode, and uses a data collection and processing technology framework that integrates big data streaming and batch computing. With reference to data governance specifications, data resources from different sources are associated and integrated using trusted modeling calculations. After cross-validation and approval, a "standardized, stable and easy-to-use" master data resource is formed, solving the problems of "difficult to find, difficult to select, and difficult to apply" of data.
[0038] (2) The master data governance method based on multi-source and multi-database stream-batch integration proposed in the present invention can adopt a stream-batch integration architecture from massive data, follow data governance specifications, and form master data with high sharing, uniqueness, long-term stability, and business criticality.
[0039] (3) The master data governance method based on multi-source and multi-database integration of stream and batch processing proposed in the present invention establishes a master data governance architecture for stream and batch processing, uses the same set of APIs and the same set of development paradigms to achieve master data governance of massive data, thereby ensuring the consistency of the master data processing process and results, eliminating data redundancy, greatly improving the relevance and availability of data, improving data processing efficiency, and enhancing the strategic synergy of the organization. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0041] Figure 1 is a diagram of an exemplary device architecture to which an embodiment of the present application may be applied;
[0042] Figure 2This is a flow chart of a master data management method based on multi-source and multi-database stream-batch integration according to an embodiment of the present application;
[0043] Figure 3 This is a module flow chart of a master data management method based on multi-source and multi-database stream-batch integration according to an embodiment of the present application;
[0044] Figure 4 Schematic diagram of a master data management device for integrated streaming and batch processing based on multiple sources and multiple databases according to an embodiment of the present application;
[0045] Figure 5 It is a structural diagram of a computer device suitable for implementing the electronic device of the embodiment of the present application. DETAILED DESCRIPTION
[0046] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It is apparent that the embodiments described are only some, not all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are intended to fall within the scope of protection of the present invention.
[0047] Figure 1 An exemplary device architecture 100 is shown to which a multi-source and multi-repository-based stream-batch integrated master data management method or a multi-source and multi-repository-based stream-batch integrated master data management device according to an embodiment of the present application can be applied.
[0048] like Figure 1 As shown, the device architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0049] Users can use terminal devices 101, 102, 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications, such as data processing applications and file processing applications, can be installed on terminal devices 101, 102, 103.
[0050] Terminal devices 101, 102, and 103 can be hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules (for example, software or software modules used to provide distributed services), or they can be implemented as a single software or software module. No specific limitations are given here.
[0051] The server 105 may be a server that provides various services, such as a background data processing server that processes files or data uploaded by the terminal devices 101, 102, and 103. The background data processing server may process the acquired files or data and generate processing results.
[0052] It should be noted that the master data management method for integrated batch and stream processing based on multiple sources and multiple databases provided in the embodiment of the present application can be executed by the server 105, or by the terminal devices 101, 102, and 103. Accordingly, the master data management device for integrated batch and stream processing based on multiple sources and multiple databases can be set in the server 105, or in the terminal devices 101, 102, and 103.
[0053] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above description is merely illustrative. Any number of terminal devices, networks, and servers may be provided as needed. If the processed data does not need to be acquired remotely, the above-described apparatus architecture may not include a network, but only require servers or terminal devices.
[0054] Figure 2 The embodiment of the present application provides a method for managing master data based on multi-source and multi-database stream and batch integration, including the following steps:
[0055] S1, obtaining original data from at least one data source among multiple data sources.
[0056] In a specific embodiment, step S1 specifically includes:
[0057] The original data of at least one data source from different data sources is collected and aggregated through a hub-and-spoke data collection mode. The different data sources include relational databases, message queue databases, column storage databases, object-oriented databases and file systems. Among them, relational databases, column storage databases and object-oriented databases adopt batch processing or streaming processing for collection, message queue databases adopt streaming processing for collection, and file systems adopt batch processing for collection.
[0058] Specifically, refer to Figure 3 Master data mostly originates from different systems and businesses. During the data aggregation phase, raw data is aggregated. Using a hub-and-spoke data collection model, raw data is collected and aggregated based on data needs. For example, for log data with high real-time and latency requirements, streaming data access is adopted. For dictionary data with low real-time requirements, or data that is not updated or is updated infrequently, batch data access is adopted.
[0059] In a specific embodiment, the data source is evaluated in the following manner, specifically including:
[0060] Define the dimensions and indicators of data source quality analysis based on the business rules and data conditions of the data source. The dimensions of data source quality analysis include technical indicators and business indicators. Technical indicators include data integrity, data uniqueness, data validity, data timeliness, and data rationality; business indicators include data consistency, data authenticity, data accuracy, data readability, and data availability.
[0061] Perform quality analysis on the data source to determine whether the data in the data source meets the master data requirements.
[0062] Specifically, before identifying master data, a demand research and analysis must be conducted. The results of the demand research must be analyzed to determine the master data required between systems, business systems, databases, business processes, and departments. This includes the following aspects:
[0063] 1) Identify drivers and requirements. This means confirming that multiple business areas within the organization require access to the same data resources and ensuring that these data sets are complete, up-to-date, and consistent master data. For example, if the HR system and the finance system require personnel information, then this personnel information constitutes a required master data requirement.
[0064] 2) Evaluate the data source. Before analyzing and identifying the master data, it is necessary to conduct a quality assessment and analysis of the data source corresponding to the original data. If the quality of the data source is very poor, it does not need to be included in the master data analysis. Define the dimensions and indicators of data quality analysis based on the business rules and data conditions of the data source. The dimensions of data source quality analysis and evaluation are analyzed from the technical dimension and the business dimension, and the indicators of data quality are divided into technical indicators and business indicators. Among them, technical indicators are analyzed from the following five perspectives: data integrity, data uniqueness, data validity, data timeliness, and data rationality. Business indicators are analyzed from the following five perspectives: data consistency, data authenticity, data accuracy, data readability, and data availability.
[0065] The steps for data source quality analysis are as follows:
[0066] 1. Determine the detection data source;
[0067] 2. Selection of key data items: Select based on the main business;
[0068] 3. Develop rules for testing indicators: For example, verify the length of the ID number to see if it is 18 digits, or verify the ID card format to see if it starts with the area code.
[0069] Each rule defines a criterion, target, or threshold for comparison. It is typically calculated using the following formula:
[0070]
[0071] Where r is the rule being tested, TestExecutions is the total number of tests, and ExceptionsFound is the number of exceptions. For example, if 560 exceptions are found in 10,000 tests of the business rule rule (r), then in this example, the result of valid data quality (ValidDQI) is 9440 / 10,000 = 94.4%, and the result of invalid data quality (lnvalidDQI) is 560 / 10,000 = 5.6%.
[0072] 4. Develop quality inspection plan;
[0073] 5. Carry out quality indicator verification;
[0074] 6. Prepare a quality assessment report.
[0075] 3) Define data access methods, data synchronization methods, and master data sharing methods. Data access methods include direct database connection, file access, and data download services. Data synchronization methods include incremental access and full access. Master data sharing methods include direct database connection to master data resources and master data services.
[0076] 4) Confirm management responsibilities and maintenance processes, including management of master data resources, management of master data quality, and management of the master data lifecycle.
[0077] S2, clean the original data to obtain standardized data.
[0078] In a specific embodiment, step S2 specifically includes:
[0079] The original data is filtered, deduplicated, and format converted to obtain standardized data.
[0080] Specifically, the duplicate, missing, scattered, and erroneous original data are cleaned according to the standards, and data governance operations such as filtering, deduplication, and format conversion are performed on the original data to standardize the original data and improve the data availability.
[0081] The data collected by the data collector is pushed to the distributed message system to provide high-throughput, scalable distributed message queue services, and through real-time streaming computing, it can meet the timeliness of data and other issues.
[0082] S3, build a master data analysis matrix based on standardized data, and analyze and identify the basic data according to the master data analysis matrix.
[0083] In a specific embodiment, step S3 specifically includes:
[0084] Based on the five indicators of business impact, data sharing, master data management maturity, master data unification difficulty, and demand urgency, the hierarchical analysis method is used to establish a master data analysis matrix for standardized data;
[0085] The weight of each indicator is determined using the geometric mean method based on the main data analysis matrix;
[0086] The weight is compared with the threshold to determine whether the standardized data is basic data within the scope of master data management.
[0087] Specifically, we analyze five dimensions: business impact, data sharing, master data management maturity, difficulty in master data unification, and urgency of need, to form a master data analysis matrix. Ultimately, the matrix scores determine whether to include each item in the master data management scope. We used the Analytic Hierarchy Process (AHP) to construct the master data analysis matrix, as shown in Table 1.
[0088] Table 1
[0089]
[0090]
[0091] Each dimension is judged and scored according to the importance of the resource indicators to obtain the score of each dimension.
[0092] Determine the weight of each indicator: Determining the weight of indicators is a very critical step in the process of master data identification. The weight is calculated by geometric mean method:
[0093] Step 1: Multiply the elements of the main data analysis matrix row by row to obtain a new column vector;
[0094] A1=1*(1 / A21)*(1 / A31)*(1 / A41)*(1 / A51);
[0095] A2=A21*1*(1 / A32)*(1 / A42)*(1 / A52);
[0096] A3=A31*A32*1*(1 / A43)*(1 / A53);
[0097] A4=A4 A41*A42*A43*1*(1 / A54);
[0098] A5=A51*A52*A53*A54*1;
[0099] Step 2: Raise each component of the new column vector to the nth power to obtain the weight vector;
[0100]
[0101] Step 3: Normalize the weight vector to get the weight.
[0102] SUM=W1+W2+W3+W4+W5;
[0103] The weight of business impact is: W1 / SUM;
[0104] The weight of data sharing degree is: W2 / SUM;
[0105] The master data management maturity weight is: W3 / SUM;
[0106] The weight of difficulty of master data unification is: W4 / SUM;
[0107] The weight of demand urgency is: W5 / SUM.
[0108] The above process can be expressed by the following formula:
[0109]
[0110] Multiply the score of each dimension by the corresponding weight, and then add up all the multiplication results to obtain the comprehensive score of the data source. Taking the threshold of 80 as an example, if the comprehensive score is above 80, it will be included in the master data management scope.
[0111] S4, define the master data table structure according to data requirements and data standards, build a conceptual data model based on basic data, expand the logical data model based on the conceptual data model and the master data table structure, establish a physical data model based on the logical data model, cross-validate the data generated by the physical data model to obtain the master data.
[0112] In a specific embodiment, step S4 specifically includes:
[0113] Define the attributes of the fields in the master data table structure, including field type, field dictionary, value range, and data rules;
[0114] Establish a conceptual data model based on the relationship between entities and entities;
[0115] The conceptual data model is formed by merging similar data, adding related entities, adjusting attributes, designing paradigms, removing data duplication, making primary keys unique, and calculating credibility to form a logical data model;
[0116] Abstract logic into entities through logical data models and define the time and cycle of model execution;
[0117] Through system verification, duplication checking, manual comparison, screening and verification, the integrity, standardization, uniqueness, consistency, accuracy and validity of the data generated by the physical data model are cross-validated in multiple dimensions to obtain the master data.
[0118] Specifically, reference international, national, industry, and corporate standards to ensure that subsequent master data planning complies with relevant regulations, reflects industry characteristics, and meets current business needs and future expansion requirements. Based on requirements and various standards, define the master data table structure and the field attributes within it, including field types, field dictionaries, value ranges, and data rules.
[0119] Based on master data planning and design, all master data-related business data resources are listed and attributed, with data with large business and attribute correlations serving as foundational data. The first step is to establish a conceptual data model to document the scope and key terms of the existing system. The second step is to establish a logical data model to document the existing system's solutions for master data governance. The third step is to establish a physical data model to understand the technical design of the existing system. The specific steps are as follows:
[0120] 1. Establish a conceptual data model
[0121] Based on the underlying data, a relational conceptual data model is depicted using information engineering syntax. This model contains only descriptions of business usage and the relationships between entities. The entity definition is the core metadata, which serves as the business rules that help the model accurately manage the relationships between entities. An entity can be a specific person or a specific thing, specifically categorized as people, things, places, objects, and organizations. Entity relationships are associations between entities. Relationship names vary across models, and the conceptual data model clearly describes the cardinality of the relationship, such as one-to-one, one-to-many, or many-to-many.
[0122] 2. Establish a logical data model
[0123] The conceptual data model is expanded by adding associated entities, attributes, specified data domains, specified keys, etc., supplementing the required details of the conceptual model to form a logical data model.
[0124] a. Merge similar data: Massive data may contain data of the same type from different sources. This data needs to be merged based on the business. For example, for personnel master data, there are different types of personnel such as corporate employees, students, and retirees. All data of the same type needs to be merged.
[0125] b. Add associated entities: Associate related attribute entities based on the relationship between entities in the conceptual data model to describe 1-to-1, 1-to-many, and many-to-many relationships.
[0126] c. Inductive mapping, attribute extension, and data type definition.
[0127] 1) Data attribute reduction: According to the master data table structure, data attributes are reduced accordingly, such as uncommon attributes and business-sensitive attributes, to increase the usability and versatility of master data.
[0128] 2) Data attribute supplementation: Based on the master data table structure, relevant attributes are associated in the trusted database, such as trusted mobile phone numbers and trusted network accounts, to increase the authority and integrity of the master data.
[0129] d. Plan and design according to business logic. The purpose of paradigm design is to reduce data redundancy and improve data sharing. Before establishing a physical data model, it is necessary to divide the sorted business entities into subject domains. For example, the three major subject domains of corporate finance, supply chain, and human resources should be further subdivided to analyze the core business categories. At this time, it is necessary to use project practical experience and industry standards to ensure that the subject domain attribute division complies with the specifications. Create a comprehensive view of the candidate data sources of master data entities within the subject domain to facilitate data modeling. Perform entity analysis and entity relationship analysis on the master data based on master data planning and design, and formulate rules for accurately matching entities and entity relationships.
[0130] e. De-duplicate data and uniquely identify primary keys. De-duplicate key fields based on business attributes and uniquely identify primary keys. One code can only represent one object entity, avoiding multiple codes for one entity and making master data more unique.
[0131] f. Calculate the credibility of data: Calculate the credibility of the data source and related data, filter out data with low credibility, and form reliable and authoritative master data.
[0132] 3. Establish a physical data model
[0133] Based on the logical data model, a physical data model is created, integrating system hardware, software, and network tools to form a detailed technical solution. The logical data model transforms logically abstract entities into independent objects in the physical database. In addition to defining standard table field attributes, reference data objects and designated proxy fields are added, employing denormalization to reduce storage and processing costs.
[0134] Furthermore, the above model adopts centralized data governance to centrally maintain master data for different data requests.
[0135] Define the time and cycle for model execution and select the data model type based on the main data type.
[0136] (1) Full model
[0137] There will be data that is updated or updated frequently, such as enterprise employee master data. Select the full mode and execute it regularly according to the defined model operation cycle.
[0138] (2) Incremental model
[0139] There will be no updated master data, such as travel trajectory master data. The travel records will not change, and the data volume is large. Incremental mode can be selected.
[0140] 4. Data cross-validation
[0141] The data generated after the physical data model is cross-validated in different dimensions such as data integrity, standardization, uniqueness, consistency, accuracy and validity through various means such as system verification, duplicate checking, manual comparison, screening, and verification to ensure that the data is accurate, complete, consistent and valid, and can meet the requirements and applications of the organization's various application systems for master data.
[0142] 5. Data Modeling Methods
[0143] The data modeling method uses batch computing, that is, offline computing, unified data layering and storage to ensure data consistency.
[0144] In a specific embodiment, it also includes:
[0145] Determine the distribution method of master data, which includes passive distribution, active distribution, and interactive distribution.
[0146] Specifically, passive distribution refers to an authorized query mode for database master data resources, whereby the demander periodically queries or extracts data for use, resulting in passive distribution of master data. Active distribution refers to providing master data to target systems such as partners, project management, customer management, and financial management through distribution interface adapters based on release strategies such as manual synchronization, automatic triggering, scheduled triggering, and passive triggering, thus achieving active distribution of master data. Interactive distribution refers to the interactive distribution method used for large amounts of real-time data, such as passenger boarding information. Passenger status may change in real time as they undergo booking, check-in, security check, boarding, and departure, ensuring the accuracy and timeliness of master data.
[0147] Further references Figure 4 As an implementation of the methods shown in the above figures, this application provides an embodiment of a master data management device based on multi-source and multi-database integrated flow and batch. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0148] The present application embodiment provides a multi-source and multi-database integrated stream and batch master data management device, including:
[0149] A data acquisition module 1 is configured to acquire raw data from at least one data source among a plurality of data sources;
[0150] Data cleaning module 2 is configured to clean the original data to obtain standardized data;
[0151] The master data analysis module 3 is configured to construct a master data analysis matrix based on the standardized data, and obtain basic data through analysis and identification according to the master data analysis matrix;
[0152] The data modeling module 4 is configured to define the master data table structure according to data requirements and data standards, build a conceptual data model based on basic data, expand the conceptual data model and the master data table structure to obtain a logical data model, establish a physical data model based on the logical data model, cross-validate the data generated by the physical data model, and obtain the master data.
[0153] Reference below Figure 5 , which shows an electronic device (eg Figure 1 A structural diagram of a computer device 500 (a server or terminal device as shown). Figure 5 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0154] like Figure 5 As shown, the computer device 500 includes a central processing unit (CPU) 501 and a graphics processing unit (GPU) 502, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 503 or the program loaded from the storage part 509 to the random access memory (RAM) 504. Various programs and data required for the operation of the device 500 are also stored in the RAM 504. The CPU 501, GPU 502, ROM 503 and RAM 504 are connected to each other via a bus 505. An input / output (I / O) interface 506 is also connected to the bus 505.
[0155] The following components are connected to the I / O interface 506: an input section 507 including a keyboard, a mouse, etc.; an output section 508 including a display such as a liquid crystal display (LCD), a speaker, etc.; a storage section 509 including a hard disk, etc.; and a communication section 510 including a network interface card such as a LAN card, a modem, etc. The communication section 510 performs communication processing via a network such as the Internet. A drive 511 may also be connected to the I / O interface 506 as needed. A removable medium 512, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed in the drive 511 as needed, so that a computer program read therefrom can be installed into the storage section 509 as needed.
[0156] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 510, and / or installed from a removable medium 512. When the computer program is executed by the central processing unit (CPU) 501 and the graphics processing unit (GPU) 502, the above-mentioned functions defined in the method of the present application are performed.
[0157] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable medium, or any combination thereof. Computer-readable media may be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, apparatuses, or components, or any combination thereof. More specific examples of computer-readable media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution device, apparatus, or component. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution apparatus, device, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical cable, RF, or any suitable combination thereof.
[0158] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based device that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0160] The modules involved in the embodiments described in this application may be implemented in software or hardware, and may also be set in a processor.
[0161] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiment; or may exist independently and not be assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device: obtains raw data from at least one data source among multiple data sources; performs data cleaning on the raw data to obtain standardized data; constructs a master data analysis matrix based on the standardized data, and obtains basic data by analyzing and identifying the master data analysis matrix; defines the master data table structure according to data requirements and data standards, constructs a conceptual data model based on the basic data, obtains a logical data model based on the conceptual data model and the master data table structure, establishes a physical data model based on the logical data model, and cross-validates the data generated by the physical data model to obtain master data.
[0162] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A master data management method based on multi-source and multi-database integrated flow and batch, characterized by: The following steps are involved: Obtaining raw data from at least one data source among the multiple data sources; Performing data cleaning on the raw data to obtain standardized data; Constructing a master data analysis matrix based on the standardized data, and obtaining basic data through analysis and identification according to the master data analysis matrix; The master data table structure is defined according to data requirements and data standards, a conceptual data model is constructed based on the basic data, a logical data model is obtained by expanding the conceptual data model and the master data table structure, a physical data model is established based on the logical data model, and the data generated by the physical data model is cross-validated to obtain the master data. Specifically, the following steps are performed: Defining attributes of fields in the master data table structure, including field type, field dictionary, value range, and data rules; Establishing a conceptual data model based on the basic data through the relationship between entities; The conceptual data model is subjected to the steps of merging similar data, adding related entities, adjusting attributes, designing paradigms, removing data duplication, making primary keys unique, and calculating credibility, to form the logical data model; Abstract logic into entities through logical data models and define the time and cycle of model execution; The master data is obtained by multi-dimensional cross-validation of the integrity, standardization, uniqueness, consistency, accuracy and validity of the data generated by the physical data model through system verification, duplicate checking, manual comparison, screening and verification.
2. The master data management method based on multi-source and multi-database stream and batch integration according to claim 1 is characterized in that: The obtaining of original data from at least one of the multiple data sources specifically includes: The raw data of at least one data source of different data sources are collected and aggregated through a hub-and-spoke data collection mode, wherein the different data sources include a relational database, a message queue database, a column storage database, an object-oriented database and a file system, wherein the relational database, the column storage database and the object-oriented database are collected using batch processing or streaming processing, the message queue database is collected using streaming processing, and the file system is collected using batch processing.
3. The master data management method based on multi-source and multi-database stream and batch integration according to claim 1 is characterized in that: The data sources are evaluated in the following ways, including: Define the dimensions and indicators of data source quality analysis based on the business rules and data conditions of the data source. The dimensions of data source quality analysis include technical indicators and business indicators. The technical indicators include data integrity, data uniqueness, data validity, data timeliness, and data rationality; the business indicators include data consistency, data authenticity, data accuracy, data readability, and data availability. Perform a quality analysis on the data source to determine whether the data in the data source meets the master data requirements.
4. The master data management method based on multi-source and multi-database stream and batch integration according to claim 1 is characterized in that: The data cleaning of the original data to obtain standardized data specifically includes: The original data is filtered, deduplicated, and format converted to obtain the standardized data.
5. The master data management method based on multi-source and multi-database stream and batch integration according to claim 1 is characterized in that: The step of constructing a master data analysis matrix based on the standardized data and obtaining basic data through analysis and identification according to the master data analysis matrix specifically includes: Based on the five indicators of business impact, data sharing, master data management maturity, master data unification difficulty, and demand urgency, the master data analysis matrix is established using the hierarchical analysis method for the standardized data; Determine the weight of each indicator using the geometric mean method based on the master data analysis matrix; The weight is compared with a threshold to determine whether the standardized data is basic data within the scope of master data management.
6. The master data management method based on multi-source and multi-database stream and batch integration according to claim 1 is characterized in that: Also includes: A distribution mode of the master data is determined, where the distribution mode includes passive distribution, active distribution, and interactive distribution.
7. A master data management device based on multi-source and multi-database integrated flow and batch, characterized by: include: a data acquisition module, configured to acquire raw data from at least one data source among a plurality of data sources; A data cleaning module is configured to clean the raw data to obtain standardized data; A master data analysis module is configured to construct a master data analysis matrix based on the standardized data, and obtain basic data through analysis and identification according to the master data analysis matrix; The data modeling module is configured to define the master data table structure according to data requirements and data standards, build a conceptual data model based on the basic data, expand the conceptual data model and the master data table structure to obtain a logical data model, establish a physical data model based on the logical data model, and cross-validate the data generated by the physical data model to obtain master data. Specifically, it includes: Defining attributes of fields in the master data table structure, including field type, field dictionary, value range, and data rules; Establishing a conceptual data model based on the basic data through the relationship between entities; The conceptual data model is subjected to the steps of merging similar data, adding related entities, adjusting attributes, designing paradigms, removing data duplication, making primary keys unique, and calculating credibility, to form the logical data model; Abstract logic into entities through logical data models and define the time and cycle of model execution; The master data is obtained by multi-dimensional cross-validation of the integrity, standardization, uniqueness, consistency, accuracy and validity of the data generated by the physical data model through system verification, duplicate checking, manual comparison, screening and verification.
8. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
An information processing system driven by a model and domain knowledge
CN108984761A
Data middle station building method and system and storage medium
CN112527774A