Data Model Generation Method, Apparatus, Electronic Device, and Readable Storage Medium

By labeling the full-link data set and building a blood relationship table, a standard data model is generated, which solves the problem of inaccurate positioning in the generation of traditional data models and improves the accuracy of data link positioning.

CN114840720BActive Publication Date: 2025-06-17CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210589944.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-26
Publication Date
2025-06-17
Estimated Expiration
2042-05-26

AI Technical Summary

Technical Problem

The generation of traditional data models has problems such as inconsistent index calibers, duplicate data construction and chimney-style development, resulting in inaccurate data link positioning.

Method used

By obtaining the data label set, marking the full-link data set, generating a tagged link data set, and then building a full-link blood relationship table, generating a directed graph of the original data, and link iterative optimization of the original data model to obtain a standard data model.

Benefits of technology

It improves the accuracy of data model construction and the accuracy of data link positioning, and can refine the data of different links in a more granular manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114840720B_ABST
    Figure CN114840720B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence technology, and discloses a data model generation method, including: obtaining a data label set, marking a full-link data set according to the data label set to obtain a marked link data set, generating a full-link lineage table according to the marks in the marked link data set, generating an original data directed graph based on the full-link lineage table, using the original data directed graph as an original data model, and performing link iteration optimization on the original data model to obtain a standard data model. In addition, the present invention also relates to blockchain technology, and the full-link data set can be obtained from nodes of a blockchain. The present invention also proposes a device for generating a data model, an electronic device, and a computer-readable storage medium. The present invention can generate a data model capable of accurately positioning a data link.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device, and computer-readable storage medium for generating a data model. Background Art

[0002] With the development of business, the amount of data has shown an explosive growth, and data lineage query and data link positioning have become increasingly important. Under the existing technology, data models are widely used to locate problems generated by data links, such as star models, snowflake models, wide table models, etc. However, the generation of traditional data models will cause the following problems: inconsistent metric calibers, duplicate data construction, chimney-style development, etc., resulting in inaccurate data link positioning. Therefore, there is an urgent need for a data model that can efficiently locate data links. Summary of the Invention

[0003] The present invention provides a method, device, electronic device, and readable storage medium for generating a data model, and its main purpose is to generate a data model that can accurately locate data links.

[0004] To achieve the above object, a method for generating a data model provided by the present invention includes:

[0005] Obtain a data tag set, and mark the full-link data set according to the data tag set to obtain a marked link data set;

[0006] Generate a full-link lineage table according to the marks in the marked link data set;

[0007] Generate an original data directed graph based on the full-link lineage table, and use the original data directed graph as the original data model;

[0008] Perform link iteration optimization on the original data model to obtain a standard data model.

[0009] Optionally, the step of marking the full-link data set according to the data tag set to obtain a marked link data set includes:

[0010] Mark the data of different links in the full-link data set according to the link tags in the data tag set to obtain a marked data set for multiple links;

[0011] Perform data application marking and database table marking on the data in the marked data set for multiple links, and summarize all the marked data of the multiple links to obtain the marked link data set.

[0012] Optionally, the step of generating a full-link lineage table according to the marks in the marked link data set includes:

[0013] Perform a top-down traversal operation on the tags in the marked link data set according to a preset traversal statement;

[0014] Summarize the data corresponding to all traversed tags and arrange them in top-down order to obtain the full-link blood relationship table.

[0015] Optionally, generating a directed graph of the original data based on the full-link blood relationship table includes:

[0016] Extract all table-type data in the full-link blood relationship table and use each table-type data as a vertex to obtain a vertex set;

[0017] Connect the vertices in the vertex set in reverse order from bottom to top to obtain the directed graph of the original data.

[0018] Optionally, iteratively optimizing the original data model to obtain a standard data model includes:

[0019] Determine the link of a preset theme in the original data model as the target link;

[0020] Determine the library table with the most usage times in the target link as the secondary end point;

[0021] Iteratively optimize the original data model according to the starting point, end point and the secondary end point in the original data model to obtain the standard data model.

[0022] Optionally, iteratively optimizing the original data model according to the starting point, end point and the secondary end point in the original data model to obtain the standard data model includes:

[0023] Summarize the data of all vertices from the starting point to the secondary end point to obtain a first local data set, and summarize the data of all vertices from the secondary end point to the end point to obtain a second local data set;

[0024] Generate a first local blood relationship table according to the first local data set, and generate a second local blood relationship table according to the second local data set;

[0025] Generate a first local directed graph and a second local directed graph according to the first local blood relationship table and the second local blood relationship table, and respectively determine the secondary end points in the first local directed graph and the second local directed graph;

[0026] Iteratively optimize the first local directed graph and the second local directed graph according to the secondary end points in the first local directed graph and the second local directed graph until the level of the target link converges to obtain the standard data model.

[0027] Optionally, before marking the full-link data set according to the data tag set, the method further includes:

[0028] Using a bilateral test elimination method to eliminate outliers from the data in the full-link data set, obtaining a data set with outliers removed;

[0029] Using a preset missing value detection function to detect missing values in the data set with outliers removed, and receiving filled data to fill the missing values, obtaining the full-link data set after data cleaning.

[0030] To solve the above problems, the present invention also provides a data model generation device, the device includes:

[0031] A data marking module, configured to obtain a data tag set, and mark the full-link data set according to the data tag set, obtaining a marked link data set;

[0032] A lineage table generation module, configured to generate a full-link lineage table according to the marks in the marked link data set;

[0033] An original model generation module, configured to generate an original data directed graph based on the full-link lineage table, and use the original data directed graph as an original data model;

[0034] A standard model generation module, configured to perform link iteration optimization on the original data model to obtain a standard data model.

[0035] To solve the above problems, the present invention also provides an electronic device, the electronic device includes:

[0036] A memory, storing at least one computer program; and

[0037] A processor, executing the computer program stored in the memory to implement the above-mentioned data model generation method.

[0038] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned data model generation method.

[0039] By marking the full-link data set, the present invention obtains a marked link data set, and generates a full-link lineage table according to the marks in the marked link data set, which can refine the data of different links at a finer granularity and improve the accuracy of data model construction. And based on the full-link lineage table, a directed graph of the original data is generated, and the original data model is iteratively optimized for the link to obtain a standard data model. Since the standard data model is constructed based on the full-link lineage table, the accuracy of data link positioning can be improved. Therefore, the data model generation method, device, electronic device and computer-readable storage medium provided by the present invention can generate a data model capable of accurate data link positioning. Description of the Drawings

[0040] Figure 1 It is a schematic flowchart of the data model generation method provided by an embodiment of the present invention;

[0041] Figure 2 It is a functional module diagram of the data model generation device provided by an embodiment of the present invention;

[0042] Figure 3 It is a schematic structural diagram of an electronic device for implementing the data model generation method provided by an embodiment of the present invention.

[0043] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0044] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0045] The embodiments of the present application provide a data model generation method. The execution subject of the data model generation method includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiments of the present application. In other words, the data model generation method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0046] Refer to Figure 1As shown in the figure, it is a schematic flowchart of the data model generation method provided by an embodiment of the present invention. In this embodiment, the data model generation method includes:

[0047] S1. Obtain a data tag set, and mark the full-link data set according to the data tag set to obtain a marked link data set.

[0048] In the embodiment of the present invention, the data tag set includes data application tags and database table tags. Among them, the data application tags include tags such as link, data domain, theme (primary / secondary / tertiary theme), and importance level. The database table tags include tags such as data domain, theme (primary / secondary / tertiary theme), business process, granularity, table level, and table type. The full-link data set includes all data from the application layer (APP) to the operation storage data layer (ODS layer) of different links (paths of business activities, such as query business links).

[0049] Specifically, before marking the full-link data set according to the data tag set, the method further includes:

[0050] Using the bilateral test elimination method to eliminate outliers from the data in the full-link data set to obtain a data set with outliers removed;

[0051] Using a preset missing value detection function to detect missing values in the data set with outliers removed, and receiving filling data to fill the missing values to obtain the full-link data set after data cleaning.

[0052] In the embodiment of the present invention, the bilateral test elimination method includes:

[0053]

[0054] Among them, represents the average value of the data in the full-link data set, S represents the standard deviation of the data in the full-link data set, and Y i represents any data in the full-link data set. G represents the test value. When G is greater than a preset test threshold, it is determined that Y i is outlier data.

[0055] The missing value detection function can be the missmap function missing function. If no data missing value is detected, no processing is performed. If a data missing value is detected, in the embodiment of the present invention, an alarm is sent to the user side, and data filling is performed according to the received filling data.

[0056] Specifically, marking the full-link data set according to the data tag set to obtain a marked link data set includes:

[0057] Mark the data of different links in the full-link data set according to the link tags in the data tag set, and obtain a marked data set for multiple links;

[0058] Perform data application marking and database table marking on the data in the marked data sets of the multiple links, and summarize all the marked data of the multiple links to obtain the marked link data set.

[0059] In an embodiment of the present invention, for example, for a marked data set marked with a data query service (a certain link), add a data application mark according to the application (APP) passed by the data query service, and add a database table mark according to the database table used.

[0060] In an embodiment of the present invention, by marking the full-link data set to obtain a marked link data set, the data of different links can be refined with a finer granularity, improving the accuracy of data model construction.

[0061] S2. Generate a full-link lineage table according to the marks in the marked link data set.

[0062] In an embodiment of the present invention, the full-link lineage table includes all data lineage relationships from the data application layer (APP layer) to the operational data store layer (ODS layer), including application name, application APP layer library table, data domain, subject, importance level, link level in the full link, library table, data domain, subject, business process, granularity, table level, dimension, index, table type, CPU consumption, library table size, change times, completion time, etc.

[0063] Specifically, generating the full-link lineage table according to the marks in the marked link data set includes:

[0064] Perform a top-down traversal operation on the marks in the marked link data set according to a preset traversal statement;

[0065] Summarize the data corresponding to all the traversed marks and arrange them in a top-down order to obtain the full-link lineage table.

[0066] In an optional embodiment of the present invention, the preset traversal statement can be an SQL statement, etc. Through top-down traversal, that is, starting from the data application layer (APP) and going from top to bottom, a full-link lineage table to the operational data store layer (ODS layer) is generated.

[0067] In an embodiment of the present invention, through marking and top-down traversal, the relationships of all data in different links can be extracted, thereby improving the accuracy of data model generation.

[0068] S3. Generate a directed graph of the original data based on the full - link lineage relationship table, and use the directed graph of the original data as the original data model.

[0069] In the embodiment of the present invention, in the directed graph G0=(V, E) of the original data, it includes a vertex set V and an edge set E. Among them, the starting point in the vertex set is the ODS - layer table, the ending point is the APP - layer table, and the remaining vertices are arbitrary tables. The distance between any two tables in the vertex set is 1, and the direction in the edge set is from the ODS layer to the APP layer.

[0070] Specifically, generating the directed graph of the original data based on the full - link lineage relationship table includes:

[0071] Extract all table - type data in the full - link lineage relationship table, and use each table - type data as a vertex to obtain a vertex set;

[0072] Connect the vertices in the vertex set in reverse order from bottom to top to obtain the directed graph of the original data.

[0073] In the embodiment of the present invention, by extracting all library tables in the full - link lineage relationship table within the same theme (a certain business activity link) and converting them in reverse (i.e., in the order from bottom to top) into a weighted directed graph G0=(V, E) from the starting point to the ending point, the original data model can be constructed quickly and accurately.

[0074] S4. Perform link iteration optimization on the original data model to obtain a standard data model.

[0075] In the embodiment of the present invention, the link iteration optimization refers to starting from different links and respectively iteratively finding the local optimal links of each link until the full - link optimal is reached.

[0076] Specifically, performing link iteration optimization on the original data model to obtain a standard data model includes:

[0077] Determine the link of the preset theme in the original data model as the target link;

[0078] Determine the library table with the most usage times in the target link as the secondary ending point;

[0079] Perform iterative optimization on the original data model according to the starting point, ending point, and the secondary ending point in the original data model to obtain the standard data model.

[0080] Specifically, performing iterative optimization on the original data model according to the starting point, ending point, and the secondary ending point in the original data model to obtain the standard data model includes:

[0081] Summarize the data of all vertices from the starting point to the secondary end point to obtain a first local data set, and summarize the data of all vertices from the secondary end point to the end point to obtain a second local data set;

[0082] Generate a first local blood relationship table according to the first local data set, and generate a second local blood relationship table according to the second local data set;

[0083] Generate a first local directed graph and a second local directed graph according to the first local blood relationship table and the second local blood relationship table, and respectively determine the secondary end points in the first local directed graph and the second local directed graph;

[0084] Perform link iteration optimization on the first local directed graph and the second local directed graph respectively according to the secondary end points in the first local directed graph and the second local directed graph until the level of the target link converges to obtain the standard data model.

[0085] In an optional embodiment of the present invention, the full-link library table within the theme is divided into two parts: starting point >> secondary end point >> end point. The local optimal links of each part are found respectively, and then the full-link optimal is obtained. The termination condition is that the levels of all links cannot be reduced any more. For example, for the data S from the starting point to the secondary end point, it is optimized and merged according to the steps of S2 to generate a new table, and iteratively optimized. Finally, the shortest path from the starting point to the secondary end point is found, so that the link reaches the local optimal. The data processing steps from the secondary end point to the end point are similar to the data processing steps of brushing data from the starting point to the secondary end point, and will not be elaborated here.

[0086] In the embodiment of the present invention, starting from the application layer (APP) to the operation storage data layer (ODS layer), all application data is covered to avoid omission, and the data model is optimized based on the reverse full-link blood relationship directed graph. With the help of the concept of graph theory, from local optimal to global optimal, the positioning accuracy of the data model is improved.

[0087] The present invention marks the full-link data set to obtain a marked link data set, and generates a full-link blood relationship table according to the marks in the marked link data set, which can refine the data of different links with finer granularity and improve the accuracy of data model construction. And based on the full-link blood relationship table, an original data directed graph is generated, and link iteration optimization is performed on the original data model to obtain a standard data model. Since the standard data model is constructed based on the full-link blood relationship table, the accuracy of data link positioning can be improved. Therefore, the data model generation method proposed by the present invention can generate a data model capable of accurate data link positioning.

[0088] As Figure 2 shown, it is a functional module diagram of a data model generation device provided by an embodiment of the present invention.

[0089] The data model generation device 100 according to the present invention can be installed in an electronic device. According to the functions implemented, the data model generation device 100 may include a data tagging module 101, a lineage table generation module 102, a raw model generation module 103, and a standard model generation module 104. The modules in the present invention may also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0090] In this embodiment, the functions of each module / unit are as follows:

[0091] The data tagging module 101 is used to obtain a data tag set, and tag the full-link data set according to the data tag set to obtain a tagged link data set;

[0092] The lineage table generation module 102 is used to generate a full-link lineage table according to the tags in the tagged link data set;

[0093] The raw model generation module 103 is used to generate a raw data directed graph based on the full-link lineage table, and use the raw data directed graph as a raw data model;

[0094] The standard model generation module 104 is used to perform link iteration optimization on the raw data model to obtain a standard data model.

[0095] Specifically, the specific implementation manners of each module of the data model generation device 100 are as follows:

[0096] Step 1: Obtain a data tag set, and tag the full-link data set according to the data tag set to obtain a tagged link data set.

[0097] In the embodiment of the present invention, the data tag set includes data application tags and database table tags. Among them, the data application tags include tags such as link, data domain, subject (primary / secondary / tertiary subject), importance level, etc., and the database table tags include tags such as data domain, subject (primary / secondary / tertiary subject), business process, granularity, table level, table type, etc. The full-link data set includes all data from the application layer (APP) to the operational storage data layer (ODS layer) of different links (paths of business activities, such as query business links).

[0098] Specifically, before tagging the full-link data set according to the data tag set, the method further includes:

[0099] Outliers are removed from the data in the full-link data set using the bilateral test removal method to obtain a data set with outliers removed.

[0100] The missing value detection function is used to detect missing values in the data in the data set with outliers removed, and the received filling data is used to fill the missing values, so as to obtain the full-link data set after data cleaning.

[0101] In the embodiment of the present invention, the bilateral test removal method includes:

[0102]

[0103] Among them, represents the average value of the data in the full-link data set, S represents the standard deviation of the data in the full-link data set, and Y i represents any data in the full-link data set. G represents the test value. When G is greater than the preset test threshold, it is determined that Y i is abnormal data.

[0104] The missing value detection function can be the missmap function missing function. If no data missing value is detected, no processing is performed. If a data missing value is detected, in the embodiment of the present invention, an alarm is sent to the user side, and data filling is performed according to the received filling data.

[0105] Specifically, the marking of the full-link data set according to the data label set to obtain the marked link data set includes:

[0106] Mark the data of different links in the full-link data set according to the link labels in the data label set to obtain the marked data sets of multiple links;

[0107] Perform data application marking and database table marking on the data in the marked data sets of multiple links, and summarize all the marked data of multiple links to obtain the marked link data set.

[0108] In the embodiment of the present invention, for example, for the marked data set marked with the data query service (a certain link), data application marking is added according to the application (APP) passed by the data query service, and database table marking is added according to the used database table.

[0109] In the embodiment of the present invention, by marking the full-link data set to obtain the marked link data set, the data of different links can be refined with a finer granularity, improving the accuracy of data model construction.

[0110] Step 2: Generate a full-link lineage table according to the marks in the marked link data set.

[0111] In the embodiments of the present invention, the full-linkage lineage table contains all data lineage relationships from the data application layer (APP layer) to the operational data store layer (ODS layer), including application name, APP layer library tables of the application, data domain, subject, importance level, link levels in the full link, library tables, data domain, subject, business process, granularity, table levels, dimensions, metrics, table types, CPU consumption, library table size, number of changes, completion time, etc.

[0112] Specifically, generating the full-linkage lineage table according to the tags in the marked link data set includes:

[0113] Performing a top-down traversal operation on the tags in the marked link data set according to a preset traversal statement;

[0114] Summarizing the data corresponding to all traversed tags and arranging them in a top-down order to obtain the full-linkage lineage table.

[0115] In an alternative embodiment of the present invention, the preset traversal statement can be an SQL statement, etc. Through top-down traversal, that is, starting from the data application layer (APP) and going from top to bottom, a full-linkage lineage table from the operational data store layer (ODS layer) is generated.

[0116] In the embodiments of the present invention, through tagging and top-down traversal, the relationships of all data in different links can be extracted, thereby improving the accuracy of data model generation.

[0117] Step 3: Generating an original data directed graph based on the full-linkage lineage table, and using the original data directed graph as the original data model.

[0118] In the embodiments of the present invention, in the original data directed graph G0=(V, E), it includes a vertex set V and an edge set E. Among them, the starting point in the vertex set is the ODS layer table, the end point is the APP layer table, and the remaining vertices are any tables. The distance between any two tables in the vertex set is 1, and the direction in the edge set is from the ODS layer to the APP layer.

[0119] Specifically, generating the original data directed graph based on the full-linkage lineage table includes:

[0120] Extracting all table-type data in the full-linkage lineage table and using each table-type data as a vertex to obtain a vertex set;

[0121] Connecting the vertices in the vertex set in a reverse order from bottom to top to obtain the original data directed graph.

[0122] In an embodiment of the present invention, by extracting all library tables in the full-link blood relationship table within the same theme (a certain business activity link) and converting them in reverse (i.e., in a bottom-up order) into a weighted directed graph G0=(V, E) from the starting point to the ending point, the original data model can be constructed quickly and accurately.

[0123] Step 4: Perform link iteration optimization on the original data model to obtain a standard data model.

[0124] In an embodiment of the present invention, the link iteration optimization refers to starting from different links and iteratively finding the local optimal links for each link until the full-link optimum is reached.

[0125] Specifically, the performing link iteration optimization on the original data model to obtain a standard data model includes:

[0126] Determine the link of the preset theme in the original data model as the target link;

[0127] Determine the library table with the most usage times in the target link as the sub-ending point;

[0128] Perform iteration optimization on the original data model according to the starting point, ending point, and the sub-ending point in the original data model to obtain the standard data model.

[0129] In detail, the performing iteration optimization on the original data model according to the starting point, ending point, and the sub-ending point in the original data model to obtain the standard data model includes:

[0130] Summarize the data of all vertices from the starting point to the sub-ending point to obtain a first local data set, and summarize the data of all vertices from the sub-ending point to the ending point to obtain a second local data set;

[0131] Generate a first local blood relationship table according to the first local data set, and generate a second local blood relationship table according to the second local data set;

[0132] Generate a first local directed graph and a second local directed graph according to the first local blood relationship table and the second local blood relationship table, and respectively determine the sub-ending points in the first local directed graph and the second local directed graph;

[0133] Perform link iteration optimization on the first local directed graph and the second local directed graph respectively according to the sub-ending points in the first local directed graph and the second local directed graph until the level of the target link converges to obtain the standard data model.

[0134] In an optional embodiment of the present invention, the full-link library tables within the subject are divided into two parts: starting point >> secondary end point >> end point. The local optimal links of each part are found respectively, and then the optimal full link is obtained. The termination condition is that the levels of all links cannot be reduced any further. For example, for the data S from the starting point to the secondary end point, it is optimized and merged according to the steps of S2 to generate a new table, and iteratively optimized. Finally, the shortest path from the starting point to the secondary end point is found, making the link reach the local optimum. The data processing steps from the secondary end point to the end point are similar to the data processing steps of brushing data from the starting point to the secondary end point, and will not be elaborated here.

[0135] In the embodiment of the present invention, starting from the application layer (APP) to the operational data store layer (ODS layer), all application data is covered to avoid omission, and the data model is optimized based on the reverse full-link lineage relationship directed graph. With the help of the concept of graph theory, from local optimum to global optimum, the positioning accuracy of the data model is improved.

[0136] The present invention marks the full-link data set to obtain a marked link data set, and generates a full-link lineage table according to the marks in the marked link data set, which can refine the data of different links with finer granularity and improve the accuracy of data model construction. And based on the full-link lineage table, an original data directed graph is generated, and the original data model is iteratively optimized for links to obtain a standard data model. Since the standard data model is constructed based on the full-link lineage table, the accuracy of data link positioning can be improved. Therefore, the data model generation device proposed by the present invention can generate a data model capable of accurate data link positioning.

[0137] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the data model generation method provided by an embodiment of the present invention.

[0138] The electronic device may include a processor 10, a memory 11, a communication interface 12, and a bus 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a data model generation program.

[0139] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 11 can be an internal storage unit of the electronic device in some embodiments, such as the mobile hard disk of the electronic device. The memory 11 can also be an external storage device of the electronic device in some other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device. Further, the memory 11 can also include both the internal storage unit and the external storage device of the electronic device. The memory 11 can be used not only to store application software installed in the electronic device and various types of data, such as the code of the data model generation program, etc., but also to temporarily store the data that has been output or will be output.

[0140] The processor 10 can be composed of integrated circuits in some embodiments. For example, it can be composed of a single packaged integrated circuit, or can also be composed of multiple integrated circuits with the same or different functions packaged together, including the combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and circuits, and by running or executing the programs or modules (such as the data model generation program, etc.) stored in the memory 11, and calling the data stored in the memory 11, to execute various functions of the electronic device and process data.

[0141] The communication interface 12 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is usually used to establish a communication connection between this electronic device and other electronic devices. The user interface can be a display (Display), an input unit (such as a keyboard (Keyboard)). Optionally, the user interface can also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.

[0142] The bus 13 may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus 13 may be divided into an address bus, a data bus, a control bus, etc. The bus 13 is configured to implement connection communication between the memory 11 and at least one processor 10, etc.

[0143] Figure 3 Only an electronic device with components is shown. Those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the electronic device, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0144] For example, although not shown, the electronic device may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0145] Furthermore, the electronic device may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device and other electronic devices.

[0146] Optionally, the electronic device may further include a user interface. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device and to display a visual user interface.

[0147] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0148] The data model generation program stored in the memory 11 in the electronic device is a combination of multiple instructions, and when running in the processor 10, it can achieve:

[0149] Obtain a data tag set, mark the full-link data set according to the data tag set, and obtain a marked link data set;

[0150] Generate a full-link lineage table according to the marks in the marked link data set;

[0151] Generate an original data directed graph based on the full-link lineage table, and use the original data directed graph as the original data model;

[0152] Perform link iteration optimization on the original data model to obtain a standard data model.

[0153] Specifically, for the specific implementation method of the above instructions by the processor 10, reference can be made to the description of the relevant steps in the corresponding embodiments of the attached drawings, which will not be elaborated here.

[0154] Furthermore, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0155] The present invention also provides a computer-readable storage medium, and the readable storage medium stores a computer program, and when the computer program is executed by the processor of the electronic device, it can achieve:

[0156] Obtain a data tag set, mark the full-link data set according to the data tag set, and obtain a marked link data set;

[0157] Generate a full-link lineage table according to the marks in the marked link data set;

[0158] Generate an original data directed graph based on the full-link lineage table, and use the original data directed graph as the original data model;

[0159] Perform link iteration optimization on the original data model to obtain a standard data model.

[0160] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0161] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0162] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0163] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0164] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.

[0165] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, technologies, and application systems.

[0166] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0167] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.

[0168] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. The terms such as second are used to denote names and do not denote any particular order.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for generating a data model, characterized in that, The method includes: Obtain a data tag set, and mark the full-link data set according to the data tag set to obtain a marked link data set; Generate a full-link lineage table based on the marks in the marked link data set; Generate an original data directed graph based on the full-link lineage table, and use the original data directed graph as the original data model; Determine the link of a preset theme in the original data model as the target link; Determine the library table with the most usage times in the target link as the secondary end point; Iteratively optimize the original data model according to the starting point, end point, and the secondary end point in the original data model to obtain a standard data model; Among them, the iteratively optimizing the original data model according to the starting point, end point, and the secondary end point in the original data model to obtain a standard data model includes: aggregating the data of all vertices from the starting point to the secondary end point to obtain a first local data set, and aggregating the data of all vertices from the secondary end point to the end point to obtain a second local data set; generating a first local lineage table according to the first local data set, and generating a second local lineage table according to the second local data set; generating a first local directed graph and a second local directed graph according to the first local lineage table and the second local lineage table, and respectively determining the secondary end points in the first local directed graph and the second local directed graph; respectively performing link iterative optimization on the first local directed graph and the second local directed graph according to the secondary end points in the first local directed graph and the second local directed graph until the level of the target link converges to obtain a standard data model.

2. The method for generating a data model according to claim 1, characterized in that, The marking the full-link data set according to the data tag set to obtain a marked link data set includes: Mark the data of different links in the full-link data set according to the link tags in the data tag set to obtain a marked data set of multiple links; Perform data application marking and database table marking on the data in the marked data set of multiple links, and aggregate all the marked data of the multiple links to obtain the marked link data set.

3. The method for generating a data model according to claim 2, characterized in that, The generating a full-link lineage table based on the marks in the marked link data set includes: Perform a top-down traversal operation on the marks in the marked link data set according to a preset traversal statement; Aggregate the data corresponding to all the traversed marks and arrange them in a top-down order to obtain the full-link lineage table.

4. The method for generating a data model according to claim 1, characterized in that, The generating an original data directed graph based on the full-link lineage table includes: Extract all table-type data in the full-link lineage table, and use each table-type data as a vertex to obtain a vertex set; Connect the vertices in the vertex set in reverse order from bottom to top to obtain the original data directed graph.

5. The method for generating a data model according to claim 1, characterized in that, Before the marking the full-link data set according to the data tag set, the method further includes: Use a bilateral test elimination method to eliminate outliers in the full-link data set to obtain an outlier-eliminated data set; Use a preset missing value detection function to detect missing values in the data in the anomaly-removed data set, and receive filled data to fill in the missing values, so as to obtain the full-link data set after data cleaning.

6. A data model generation device, characterized in that, The device includes: A data marking module, configured to obtain a data label set, and mark the full-link data set according to the data label set to obtain a marked link data set; A lineage table generation module, configured to generate a full-link lineage table according to the marks in the marked link data set; An original model generation module, configured to generate an original data directed graph based on the full-link lineage table, and use the original data directed graph as the original data model; A standard model generation module, configured to determine the link of a preset theme in the original data model as the target link; determine the library table with the most usage times in the target link as the secondary end point; perform iterative optimization on the original data model according to the start point, end point, and the secondary end point in the original data model to obtain a standard data model; Among them, the iterative optimization of the original data model according to the start point, end point, and the secondary end point in the original data model to obtain a standard data model includes: summarizing the data of all vertices from the start point to the secondary end point to obtain a first local data set, and summarizing the data of all vertices from the secondary end point to the end point to obtain a second local data set; generating a first local lineage table according to the first local data set, and generating a second local lineage table according to the second local data set; generating a first local directed graph and a second local directed graph according to the first local lineage table and the second local lineage table, and respectively determining the secondary end points in the first local directed graph and the second local directed graph; performing link iterative optimization on the first local directed graph and the second local directed graph respectively according to the secondary end points in the first local directed graph and the second local directed graph until the level of the target link converges to obtain a standard data model.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can implement the data model generation method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by the processor, implements the data model generation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and device for generating query data model system

    CN111723074A

  • Data blood relationship analysis method and device, equipment and storage medium

    CN113486008A