A database modeling method, device and equipment and computer storage medium

CN114138913BActive Publication Date: 2026-09-22CHINA CONSTRUCTION BANK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111491783.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-08
Publication Date
2026-09-22
Estimated Expiration
2041-12-08

AI Technical Summary

Benefits of technology

[0020]本申请实施例的数据库的建模方法、装置、设备及计算机存储介质,能够根据通过对业务系统的业务和数据进行分析,确定数据库的主题域,对主题域进行细化确定次级数据域,根据数据粒度确定数据域的重要实体,并将数据粒度域实体进行绑定,设计数据粒度的属性从而生成数据库的粒度模型。通过对数据粒度的含义和属性进行定义,将数据粒度进行归类划分,消除了维度属性和度量指标的定义;同时将数据粒度和数据实体对应,保证同一数据粒度的数据项属于同一实体,从数据粒度的角度进行数据模型建模,生成支持多种业务场景的服务模式的模型,提高了数据仓库的数据模型的稳定性与兼容性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114138913B_ABST
    Figure CN114138913B_ABST
Patent Text Reader

Abstract

The application discloses a database modeling method and device, equipment and computer storage medium, wherein the method comprises: analyzing business requirements and data requirements in a business system to obtain a subject domain of the business system and a data range corresponding to the subject domain; performing step-by-step refinement on the subject domain to obtain at least one secondary data domain located below the subject domain; for each secondary data domain, respectively performing: clustering data granularity of the secondary data domain to obtain a target entity of the secondary data domain; refining the target entity to obtain a logical entity corresponding to the target entity; and defining attributes of the logical entity according to data items in the data range. According to the embodiment of the application, the problem of poor data model stability and compatibility of the data warehouse in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data processing technology, and in particular relates to a database modeling method, apparatus, device, and computer storage medium. Background Technology

[0002] The relationship between enterprise business systems and big data processing is inseparable. With the development of business systems, supported by massive amounts of data and powerful technical capabilities, many data applications and mining models have emerged, such as fixed reports, data mining, multidimensional analysis, and autonomous data use and querying. In various data usage scenarios, normalized models and dimensional models are typically used to model data warehouses. However, normalized models perform poorly with multi-table joins and numerous queries, while dimensional models require redefining dimensions when business changes occur, often necessitating reprocessing of dimensional data, leading to data redundancy. Therefore, existing data warehouse models suffer from poor stability and compatibility. Summary of the Invention

[0003] This application provides a database modeling method, apparatus, device, and computer storage medium, which can solve the problem of poor stability and compatibility of data models in existing data warehouses.

[0004] In a first aspect, embodiments of this application provide a database modeling method, including:

[0005] Analyze the business requirements and data requirements in the business system to obtain the subject domain of the business system and the data range corresponding to the subject domain;

[0006] The subject domain is refined level by level to obtain at least one secondary data domain at a level below the subject domain;

[0007] For each of the sub-data domains, the following steps are performed: clustering is performed on the data granularity of the sub-data domain to obtain the target entity of the sub-data domain;

[0008] The target entity is refined to obtain the logical entity corresponding to the target entity;

[0009] Define the attributes of the logical entity based on the data items within the data range.

[0010] On the other hand, embodiments of this application provide a database modeling apparatus, the apparatus comprising:

[0011] The analysis module is used to analyze the business requirements and data requirements in the business system to obtain the subject domain of the business system and the data range corresponding to the subject domain;

[0012] The first refinement module is used to refine the subject domain step by step to obtain at least one secondary data domain at a level below the subject domain.

[0013] The clustering module is used to perform the following for each of the secondary data domains: clustering the data granularity of the secondary data domains to obtain the target entities of the secondary data domains;

[0014] The second refinement module is used to refine the target entity to obtain the logical entity corresponding to the target entity;

[0015] The definition module is used to define the attributes of the logical entity based on the data items within the data range.

[0016] In another aspect, embodiments of this application provide an electronic device, the device comprising:

[0017] Processor and memory storing computer program instructions;

[0018] When the processor executes the computer program instructions, it implements the database modeling method as described in any one of the claims of this application.

[0019] Furthermore, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the database modeling method as described in any claim of this application.

[0020] The database modeling method, apparatus, device, and computer storage medium in this application embodiment can determine the subject domain of the database by analyzing the business and data of the business system, refine the subject domain to determine the secondary data domain, determine the important entities of the data domain according to the data granularity, bind the entities of the data granularity domain, and design the attributes of the data granularity to generate a granular model of the database. By defining the meaning and attributes of data granularity, the data granularity is classified and divided, eliminating the definition of dimensional attributes and metrics; at the same time, the data granularity is mapped to data entities, ensuring that data items of the same data granularity belong to the same entity. Data modeling is performed from the perspective of data granularity, generating a model that supports service modes for multiple business scenarios, thereby improving the stability and compatibility of the data warehouse data model. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1This is a flowchart illustrating an embodiment of the database modeling method involved in this application;

[0023] Figure 2 This is a schematic diagram of one embodiment involved in this application;

[0024] Figure 3 This is a schematic diagram of an embodiment of the database modeling apparatus involved in this application;

[0025] Figure 4 This is a schematic diagram of an embodiment of the electronic device involved in this application. Detailed Implementation

[0026] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0027] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0028] All embodiments of this application comply with the relevant provisions of national laws and regulations regarding the acquisition, storage, use, and processing of data.

[0029] To address the problems of the prior art, embodiments of this application provide a database modeling method, apparatus, device, and computer storage medium. The database modeling method provided in this application embodiment will be described first below.

[0030] The database modeling method of this application can be executed by a computer or server with data processing capabilities. For example, database modeling in banking operations can be executed by the bank's server.

[0031] Figure 1 A schematic flowchart of a database modeling method according to an embodiment of this application is shown. Figure 1 As shown, the database modeling method of this application includes:

[0032] Step 101: Analyze the business requirements and data requirements in the business system to obtain the subject domain of the business system and the data range corresponding to the subject domain;

[0033] Step 102: Refine the subject domain step by step to obtain at least one secondary data domain at a level below the subject domain;

[0034] Step 103: For each of the sub-data domains, perform the following: cluster the data granularity of the sub-data domains to obtain the target entities of the sub-data domains;

[0035] Step 104: Refine the target entity to obtain the logical entity corresponding to the target entity;

[0036] Step 105: Define the attributes of the logical entity based on the data items within the data range.

[0037] The database modeling method in this application can determine the subject domain of the database by analyzing the business and data of the business system, refine the subject domain to determine the secondary data domain, determine the important entities of the data domain according to the data granularity, bind the entities of the data granularity domain, and design the attributes of the data granularity to generate the granular model of the database. This application defines the meaning and attributes of data granularity, classifies and divides data granularity, and eliminates the definition of dimensional attributes and metrics; at the same time, it maps data granularity to data entities, ensuring that data items of the same data granularity belong to the same entity. By modeling the data model from the perspective of data granularity, it generates a model that supports service modes for multiple business scenarios, improving the stability and compatibility of the data warehouse data model.

[0038] In step 101, the server analyzes the business requirements and data requirements of the business system to obtain the subject domain of the business system and the data range corresponding to the subject domain.

[0039] In the embodiments of this application, a business system refers to the business activities required for an enterprise to achieve its positioning, as well as the relationships and architecture of these business activities. The business system includes the enterprise's business activity processes, data structures, and data flows.

[0040] Business requirements may include: business projects and processes for building a data warehouse; information flow and structure within the business processes; business domains, business objects, business activities, and business priorities of the business system.

[0041] Data requirements can include: data related to the business needs of the data warehouse, data involved in the information flow and information structure in the business process, data sources, data rules and data quality of the data warehouse, etc.

[0042] Business and data requirements can be obtained by conducting preliminary data research on the business systems.

[0043] A subject area is a collection of closely related data topics and forms the boundary of a data warehouse model. A data warehouse's subject areas can be created by servers analyzing and categorizing data from business systems and their data flows, resulting in a collection of data topics such as participants, products, and contracts.

[0044] The division of subject areas can be determined based on the content contained in the data subject area and the definition of the data subject area.

[0045] The data scope can be the scope of the data warehouse subject, or it can be the data scope of data related to business processes, such as the data scope of source data and the data scope of data requirements.

[0046] The scope of data in a data warehouse can be determined based on its resources (personnel, systems, budget, etc.), schedule (business time requirements and phase requirements), and functions (the functions that the data warehouse needs to achieve).

[0047] For example, in the process of modeling a bank's data warehouse, the bank server can analyze the business needs and data requirements of the banking business system based on the survey results of the banking business system, determine the theme of the data warehouse, determine the subject area based on the summarization and classification of the theme, and determine the data range of the subject area based on the data range of the business needs and data requirements.

[0048] In step 102, the server refines the subject domain by subdividing it into at least one secondary data domain at a lower level than the subject domain.

[0049] In embodiments of this application, the refinement process of a subject domain may involve further refining the subject domain based on the business entities it contains and the relationships between those entities, generating a data domain with a data range lower than that of the subject domain. The business entities contained in the subject domain and the relationships between them are determined by the business meaning and business processes of the subject domain.

[0050] The data domain level is used to indicate the size of the data range of the data domain; the larger the data range of the data domain, the higher the level.

[0051] Secondary data fields are data fields that are at a lower level than the subject field. A subject field contains at least one secondary data field. A secondary data field can be the lowest level data field, or it can contain one or more data fields that are at a lower level than the secondary data field.

[0052] For example, in the process of modeling a bank's data warehouse, the bank server can determine the secondary data domains whose data scope is smaller than that of the subject domain, based on the business entities contained in the subject domain and the relationships between those entities. The number of secondary data domains is at least one. For example, the banking business subject domain includes secondary data domains such as deposits, loans, and agency services.

[0053] In step 103, the server performs the following for each sub-data domain: clustering the data granularity of each sub-data domain, and (based on the clustering results) obtaining the target entity of the sub-data domain.

[0054] Data granularity refers to the degree of refinement or integration of data aggregated and stored in a data warehouse.

[0055] In this embodiment, by naming and defining data granularity, data granularity can represent the degree of refinement of data, or it can represent the business meaning contained in its definition. The determination of data granularity needs to be within the same data range.

[0056] Clustering is a process of grouping data granularities based on their characteristic parameters to generate a set of data granularities.

[0057] The granularity set generated after clustering data granularity can be divided into single granularity and combined granularity. Single granularity represents a uniquely identified data granularity, while combined granularity is a combination of related single granularities.

[0058] The target entity of the secondary data domain is the conceptual entity of the data model, used to represent the business meaning of the set of data granularities generated by clustering in the secondary data domain.

[0059] The target entity is used to represent the conceptual entity of important business in the secondary data domain. The target entity can represent the business and business meaning contained in the entity of the secondary data domain through the business meaning at the data granularity.

[0060] For example, such as Figure 2 As shown, in the process of data warehouse modeling in a bank, the bank server can cluster the data granularity in each sub-data domain according to the definition of data granularity and feature parameters, and use the set of data granularity obtained by clustering to represent the target entity of the sub-data domain.

[0061] Optionally, clustering the data granularity of the secondary data domain to obtain the target entity of the secondary data domain includes:

[0062] Clustering is performed based on the feature parameters of the data granularity to obtain the target entities of the secondary data domain, wherein the feature parameters of the data granularity within the same target entity satisfy a preset similarity condition.

[0063] In the embodiments of this application, the feature parameters are parameters used as criteria for data granularity clustering, which are business information or attribute information contained in the data granularity definition. The selection of feature parameters can be based on information such as the business meaning of the target entity in the secondary data domain.

[0064] Feature parameters can be preset by the model-building administrator or set by the server based on the parameters of the target entity in the secondary data domain.

[0065] Similarity criteria are used to represent the conditions that must be met to cluster data granularities into the same set of data granularities. Similarity criteria can be similar or identical feature parameters of the data granularities, or the data granularities can have the same name and meaning, etc.

[0066] In this embodiment, the server can classify data granularity based on its characteristic parameters to identify important business entities in the secondary data domain. Categorizing data granularity facilitates user retrieval and use of data, and guides the design of the physical model and subsequent model building work.

[0067] Optionally, the feature parameters include at least one of the following:

[0068] The degree of refinement of the data granularity;

[0069] The source of the data granularity;

[0070] The application scenarios for the data granularity.

[0071] The characteristic parameters of data granularity can include the degree of refinement of data granularity, the source of data granularity, and the application scenarios of data granularity.

[0072] The higher the level of data granularity, the larger the data granularity. Data granularity is clustered according to the requirements for data granularity size in the data granularity set.

[0073] The source of data granularity can be determined based on the data range to which the data granularity belongs.

[0074] The application scenarios of data granularity can be based on the business application scenarios represented by the business meaning of data granularity, or they can represent the application scenarios corresponding to different levels of data granularity.

[0075] In step 104, the server refines each target entity to obtain the logical entity corresponding to the target entity.

[0076] In the embodiments of this application, refining the target entity means refining the target entity from a conceptual entity to a data entity, and using the data granularity set obtained by data granularity clustering as the logical entity of the target entity.

[0077] Logical entities and data granularities have a one-to-one relationship at the logical level; that is, a logical entity is a set of data granularities obtained by clustering the data granularities corresponding to the target entity.

[0078] For example, in the process of data warehouse modeling in a bank, the bank server can refine the target entity from a conceptual entity to a logical entity, and map the set of data granularities obtained by data granularity clustering to the target entity.

[0079] Optionally, refining the target entity to obtain the logical entity corresponding to the target entity includes:

[0080] Based on the business projects of the business system, define the data granularity within the data range and describe the business meaning contained in the data granularity;

[0081] The data granularity is attached to the secondary data field corresponding to the target entity to obtain the logical entity corresponding to the target entity.

[0082] In the embodiments of this application, the business items of the business system are the businesses included in the business system, and the data range to which the data granularity belongs is the data range of the subject domain and each sub-data domain to which the data granularity belongs.

[0083] Define data granularity by naming it according to the naming conventions. Data granularity can be named with reference to the business meaning expressed by the business primary key. The naming of data granularity should concisely express the business meaning.

[0084] The defined data granularity is attached to the set of data granularities generated during the data granularity clustering process, thereby attaching the data granularity to the target entity corresponding to the data granularity and obtaining the logical entity corresponding to the target entity.

[0085] In this embodiment, by defining data granularity, the set of data granularities generated by clustering is attached to the target entity, thereby refining the target entity into logical entities and achieving a correspondence between data granularity and the structure of the data warehouse model. By defining data granularity and eliminating dimensions and metrics, logical modeling based on data granularity is realized.

[0086] In step 105, the server defines the attributes of logical entities based on the data items within the data range of each data domain.

[0087] In the embodiments of this application, a data item can be data related to the business requirements of the data warehouse in the business system, or data involved in information flow and information structure in the business process, as well as data related to data flow in the business system. A data item can be associated with a data granularity to represent the attributes of that data granularity.

[0088] The content of a data item may include: attribute naming, attribute definition, and attribute business definition.

[0089] Attributes of a logical entity are all the characteristics of that logical entity. For example, a user might have a name, gender, address, and contact information. Attributes of a logical entity can be represented by data items linked to data granularity.

[0090] For example, in the process of modeling a bank's data warehouse, the bank server can attach data items of each data domain to data granularity, and use data items of data granularity to represent the attributes of the corresponding logical entities.

[0091] Optionally, defining the attributes of the logical entity based on the data items within the data range includes:

[0092] Select a target data granularity from the data granularities included in the data range; the target data granularity is any data granularity included in the data range.

[0093] Data items belonging to the target data granularity are processed so that different processed data items have different names and meanings;

[0094] The processed data items are then attached to the logical entity.

[0095] Define the attributes of a logical entity to which data items are attached, to describe the detailed information contained in the attributes of the logical entity.

[0096] In the embodiments of this application, the target data granularity may be a specific data granularity determined during the model design process.

[0097] Processing data items within the target data granularity can involve analyzing all data items at the same granularity, including their attribute names, definitions, and processing criteria. This process distinguishes data items at the same granularity and clarifies the name and meaning of each granularity.

[0098] The process of attaching data items to logical entities involves attaching the processed data items to the same data granularity. Based on the one-to-one relationship between the data granularity and the logical entity at the logical model level, the data items are attached to the logical entity.

[0099] After attaching data items to the same data granularity, the attributes of logical entities are defined in detail, explicitly describing the detailed information contained in the attributes of logical entities.

[0100] The attributes of logical entities can be pre-defined by the modelers or by the server based on pre-defined naming conventions and definition information.

[0101] In this embodiment, the server can preprocess data items at the target data granularity and attach the processed data items to the corresponding logical entities. By attaching data items to the corresponding logical entities, the attributes of the logical entities can be defined, the information of the logical entities can be improved, and thus better guide the modeling of the physical model.

[0102] Optionally, the processing of data items belonging to the target data granularity includes at least one of the following:

[0103] For data items with the same name or meaning, deduplication is performed;

[0104] For data items with the same meaning but different names, merge them.

[0105] For data items with the same name but different meanings, split them up.

[0106] The process of processing data items at the target data granularity may include analyzing the data items to determine whether they are homonyms, homonyms with different names, or homonyms with different meanings.

[0107] Data item deduplication: For data items with the same or similar names, perform deduplication processing and delete duplicate data items with the same or similar names;

[0108] Data item merging: For data items with the same meaning but different names, merge them into the same data item.

[0109] Data item splitting: For data items with the same name but different meanings, splitting is performed to separate the data items with the same name but different meanings into different data items.

[0110] In this embodiment, by performing deduplication, merging, and splitting on data items, the uniqueness, consistency, and accuracy of the data can be guaranteed.

[0111] Optionally, the definition of attributes of the logical entity to which data items are attached describes the detailed information contained in the attributes of the logical entity, including:

[0112] Name the attributes contained in the logical entity in accordance with the attribute naming convention;

[0113] The attributes are described and defined in detail;

[0114] Determine the value range of the attribute value package;

[0115] Clearly define the source component, source table, and source field information for the attribute;

[0116] Define the business rules and business definition information for the attributes.

[0117] In the embodiments of this application, the detailed information of the logical entity's attributes includes: whether it is a primary key, whether it is a foreign key, referencing entity, domestic / overseas flag, derived flag, data region to which it belongs, attribute definition, source component, Chinese name of the component source table, English name of the component source table, Chinese name of the component source field, English name of the component source field, definition, statistical frequency, registration information, etc.

[0118] The process of describing the attributes of logical entities can be defined and described by the server based on the attribute information pre-configured by the modeler, or it can be defined and described by the server based on the naming conventions or attribute information conventions of the business system.

[0119] In this embodiment, by defining the attributes of logical entities in detail, the data structure of the database model can be refined, making it easier for users to find and use data, and improving the accuracy of data retrieval.

[0120] Figure 3 The database modeling apparatus 300 provided in the embodiments of this application is shown, such as Figure 3 As shown, the database modeling apparatus 300 of this application includes:

[0121] Analysis module 301 is used to analyze the business requirements and data requirements in the business system to obtain the subject domain of the business system and the data range corresponding to the subject domain;

[0122] The first refinement module 302 is used to refine the subject domain step by step to obtain at least one secondary data domain at a level below the subject domain.

[0123] Clustering module 303 is used to perform the following for each of the secondary data domains: clustering the data granularity of the secondary data domains to obtain the target entities of the secondary data domains;

[0124] The second refinement module 304 is used to refine the target entity to obtain the logical entity corresponding to the target entity;

[0125] The definition module 305 is used to define the attributes of the logical entity based on the data items within the data range.

[0126] Optionally, the clustering module 303 is specifically used to perform clustering based on the feature parameters of the data granularity to obtain the target entity of the secondary data domain, wherein the feature parameters of the data granularity within the same target entity satisfy a preset similarity condition.

[0127] Optionally, the feature parameters include at least one of the following:

[0128] The degree of refinement of the data granularity;

[0129] The source of the data granularity;

[0130] The application scenarios for the data granularity.

[0131] Optionally, the second refinement module 304 includes:

[0132] The data granularity definition unit is used to define the data granularity within the data range and describe the business meaning contained in the data granularity according to the business items of the business system.

[0133] The first attaching unit is used to attach the data granularity to the secondary data field corresponding to the target entity to obtain the logical entity corresponding to the target entity.

[0134] Optionally, module 305 includes:

[0135] A data granularity selection unit is used to select a target data granularity from the data granularities included in the data range; the target data granularity is any data granularity included in the data range.

[0136] The data granularity processing unit is used to process data items belonging to the target data granularity so that the processed data items have different names and meanings;

[0137] The second attach unit is used to attach the processed data items to the logical entity;

[0138] An attribute definition unit is used to define the attributes of a logical entity to which data items are attached, in order to describe the detailed information contained in the attributes of the logical entity.

[0139] Optionally, the data granularity processing unit is specifically used for:

[0140] For data items with the same name or meaning, deduplication is performed;

[0141] For data items with the same meaning but different names, merge them.

[0142] For data items with the same name but different meanings, split them up.

[0143] Optionally, the attribute definition unit is specifically used for:

[0144] Name the attributes contained in the logical entity in accordance with the attribute naming convention;

[0145] The attributes are described and defined in detail;

[0146] Determine the value range of the attribute value package;

[0147] Clearly define the source component, source table, and source field information for the attribute;

[0148] Define the business rules and business definition information for the attributes.

[0149] The database modeling apparatus of this application embodiment can determine the subject domain of the database by analyzing the business and data of the business system, refine the subject domain to determine the secondary data domain, determine the important entities of the data domain according to the data granularity, bind the entities of the data granularity domain, and design the attributes of the data granularity to generate the granular model of the database. By defining the meaning and attributes of data granularity, the data granularity is classified and divided, eliminating the definition of dimensional attributes and metrics; at the same time, the data granularity is mapped to data entities, ensuring that data items of the same data granularity belong to the same entity. Data modeling is performed from the perspective of data granularity, generating a model that supports service modes for multiple business scenarios, thereby improving the stability and compatibility of the data warehouse data model.

[0150] Figure 4 A schematic diagram of the hardware structure of the electronic device 400 provided in an embodiment of this application is shown.

[0151] Electronic device 400 may include processor 401 and memory 402 storing computer program instructions.

[0152] Specifically, the processor 401 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0153] Memory 402 may include mass storage for data or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 402 is non-volatile solid-state memory.

[0154] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of this disclosure.

[0155] The processor 401 reads and executes computer program instructions stored in the memory 402 to implement any of the database modeling methods in the above embodiments.

[0156] In one example, electronic device 400 may also include communication interface 403 and bus 410. For example, Figure 4 As shown, the processor 401, memory 402, and communication interface 403 are connected through bus 410 and complete communication with each other.

[0157] The communication interface 403 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0158] Bus 410 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 410 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.

[0159] The electronic device 400 can execute the database modeling method in the embodiments of this application, thereby achieving a combination Figure 1 and Figure 3 The described database modeling methods and apparatus.

[0160] Furthermore, in conjunction with the database modeling methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. This computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the database modeling methods in the above embodiments.

[0161] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0162] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0163] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0164] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0165] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A database modeling method, characterized in that, include: Analyze the business requirements and data requirements in the business system to obtain the subject domain of the business system and the data range corresponding to the subject domain; The subject domain is refined level by level to obtain at least one secondary data domain at a level below the subject domain; For each of the sub-data domains, the following steps are performed: clustering is performed on the data granularity of the sub-data domain to obtain the target entity of the sub-data domain; The target entity is refined to obtain the logical entity corresponding to the target entity; Define the attributes of the logical entity based on the data items within the data range; The process of clustering the data granularity of the secondary data domain to obtain the target entities of the secondary data domain includes: Clustering is performed based on the feature parameters of the data granularity to obtain the target entities of the secondary data domain, wherein the feature parameters of the data granularity within the same target entity satisfy a preset similarity condition; the feature parameters include at least one of the following: The degree of refinement of the data granularity; The source of the data granularity; The application scenarios of the data granularity; The granularity set generated after clustering the data granularity includes single granularity and combined granularity. Single granularity represents a uniquely identified data granularity, and combined granularity is a combination of related single granularities. The refinement of the target entity to obtain the logical entity corresponding to the target entity includes: Based on the business projects of the business system, define the data granularity within the data range and describe the business meaning contained in the data granularity; The data granularity is attached to the secondary data field corresponding to the target entity to obtain the logical entity corresponding to the target entity; The step of defining the attributes of the logical entity based on the data items within the data range includes: Select a target data granularity from the data granularities included in the data range, wherein the target data granularity is any data granularity included in the data range; Data items belonging to the target data granularity are processed so that different processed data items have different names and meanings; The processed data items are attached to the logical entity. The process of attaching data items to the logical entity is to attach the processed data items to a set of data granularities. Based on the one-to-one relationship between the set of data granularities and the logical entity at the logical model level, the data items are attached to the logical entity. Define the attributes of a logical entity to which data items are attached, to describe the detailed information contained in the attributes of the logical entity, so that the logical entity can be used to guide the modeling of the physical model.

2. The database modeling method according to claim 1, characterized in that, The processing of data items belonging to the target data granularity includes at least one of the following: For data items with the same name or meaning, deduplication is performed; For data items with the same meaning but different names, merge them. For data items with the same name but different meanings, split them up.

3. The database modeling method according to claim 1, characterized in that, The definition defines the attributes of a logical entity to which data items are attached, to describe the detailed information contained in the attributes of the logical entity, including: Name the attributes contained in the logical entity in accordance with the attribute naming convention; The attributes are described and defined in detail; Determine the range of attribute values; Clearly define the source component, source table, and source field information for the attribute; Define the business rules and business definition information for the attributes.

4. A database modeling apparatus, characterized in that, The device includes: The analysis module is used to analyze the business requirements and data requirements in the business system to obtain the subject domain of the business system and the data range corresponding to the subject domain; The first refinement module is used to refine the subject domain step by step to obtain at least one secondary data domain at a level below the subject domain. The clustering module is used to perform the following for each of the secondary data domains: clustering the data granularity of the secondary data domains to obtain the target entities of the secondary data domains; The second refinement module is used to refine the target entity to obtain the logical entity corresponding to the target entity; The definition module is used to define the attributes of the logical entity based on the data items within the data range; The clustering module is further configured to perform clustering based on the feature parameters of the data granularity to obtain target entities in the secondary data domain, wherein the feature parameters of the data granularity within the same target entity satisfy a preset similarity condition; the feature parameters include at least one of the following: The degree of refinement of the data granularity; The source of the data granularity; The application scenarios of the data granularity; The granularity set generated after clustering the data granularity includes single granularity and combined granularity. Single granularity represents a uniquely identified data granularity, and combined granularity is a combination of related single granularities. The second refinement module includes: a data granularity definition unit, used to define the data granularity within the data range and describe the business meaning contained in the data granularity according to the business items of the business system; The first attaching unit is used to attach the data granularity to the secondary data field corresponding to the target entity to obtain the logical entity corresponding to the target entity. The definition module includes: a data granularity selection unit, used to select a target data granularity from the data granularities included in the data range; the target data granularity is any data granularity included in the data range; a data granularity processing unit, used to process data items belonging to the target data granularity so that different processed data items have different names and meanings; a second attachment unit, used to attach the processed data items to the logical entity, the process of attaching data items to the logical entity is to attach the processed data items to the same set of data granularities, and realize the attachment of data items to the logical entity according to the one-to-one relationship between the set of data granularities and the logical entity at the logical model level; and an attribute definition unit, used to define the attributes of the logical entity with attached data items, to describe the detailed information contained in the attributes of the logical entity, so that the logical entity can be used to guide the modeling of the physical model.

5. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the database modeling method as described in any one of claims 1-3.

6. A computer storage medium, characterized in that, The computer storage medium stores computer program instructions, which, when executed by a processor, implement the database modeling method as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Method of standardizing power enterprise data resources

    CN108280562A

  • Pyramid model designing method to service in data warehouse building model

    CN1897026A