Data table standardization method, device, equipment and computer storage medium
By automatically identifying business time fields and table categories, and generating standard table names and standard data items of standardized tables, the problem that data table standardization in the existing technology is solved, and efficient automatic standardization is achieved.
Patent Information
- Application Number
- CN202210320120.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-03-29
AI Technical Summary
In the prior art, standardization of field names and table names of data tables requires manual operation, which is time-consuming and labor-intensive and inefficient.
By automatically identifying the business time fields and table categories based on the original table information and data element benchmarking results of the source data table to be standardized, the standard table name and standard data item of the standardized table are generated.
Automatic standardization of field names and table names is realized, which avoids the time-consuming and labor-consuming problem of manual standardization and improves the efficiency of data table standardization.
Smart Images

Figure CN114648010B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to the field of data standardization technology, and provides a data table standardization method, device, equipment and computer storage medium. Background Art
[0002] With the popularization and development of Internet technology, data is growing rapidly and the types of data are becoming more and more diverse. The development of big data technology and artificial intelligence technology has provided basic conditions and application scenarios for the use of massive data. Since each business system is relatively independent and there may be problems such as inconsistent input standards, the data expression methods in each business system are messy and different, which brings difficulties to subsequent research and use. Therefore, in order to more conveniently put massive data into the research process and mine the value of data, data standardization is essential.
[0003] However, the current standardization process is usually adjusted manually, especially the naming of field names and table names of standardized tables is time-consuming and labor-intensive. Therefore, it is very necessary to be able to automatically standardize field names and table names. Summary of the invention
[0004] The embodiments of the present application provide a data table standardization method, apparatus, device and computer storage medium for implementing the standardization of field names and table names.
[0005] In one aspect, a data table standardization method is provided, the method comprising:
[0006] Based on the original table information of the source data table to be standardized and the data element benchmarking result of the source data table, determining the business time field included in the source data table;
[0007] Based on the original table information, table information is identified to determine the table category corresponding to the source data table; wherein the table category includes subject domain category, business category and partition mode category;
[0008] Based on the table category, generate a standard table name of a standardized table corresponding to the source data table;
[0009] Generate each standard data item of the standardized table based on the data element matching result, the original table information and the business time field;
[0010] The standardized table is obtained based on the standard table name and the various standard data items.
[0011] In one aspect, a data table standardization device is provided, the device comprising:
[0012] A business field identification unit, configured to determine a business time field included in a source data table based on original table information of the source data table to be standardized and a data element alignment result of the source data table;
[0013] A table information identification unit, used to identify the table information based on the original table information, and determine the table category corresponding to the source data table; wherein the table category includes a subject domain category, a business category, and a partitioning method category;
[0014] A table naming unit, used for generating a standard table name of a standardized table corresponding to the source data table based on the table category;
[0015] A data item naming unit, used to generate each standard data item of the standardized table based on the data element benchmarking result, the original table information and the business time field;
[0016] The standard table generating unit is used to obtain the standardized table based on the standard table name and the various standard data items.
[0017] Optionally, the device further comprises an automated benchmarking unit, configured to:
[0018] Extracting information from the source data table to obtain the original table information; wherein the original table information includes the table name and field information of the source data table;
[0019] A matching process is performed on each of the obtained field information to determine a matching result of a data element corresponding to each of the field information. The data element matching result includes a data element and a qualifier corresponding to each of the field information.
[0020] Optionally, the business field identification unit is specifically used to:
[0021] Based on the Chinese field information in the original table information and the data element matching result, determining the time field included in the source data table;
[0022] Based on the set non-business time field set, the non-business time fields in the time fields included in the source data table are filtered out;
[0023] The time fields remaining after filtering are determined as business time fields.
[0024] Optionally, the business field identification unit is further used to:
[0025] For each of the determined service time fields, if there is a service time field that does not correspond to all representation types, the missing representation types are completed;
[0026] For each of the non-business time fields, if there is a non-business time field including other representation types except the specified representation type, the other representation types are deleted.
[0027] Optionally, the table information identification unit is specifically used to:
[0028] Perform subject domain identification based on the table name and the field information to determine the subject domain category to which the source data table belongs;
[0029] Based on the table name and the field information, a partition mode is identified to determine the partition mode category to which the source data table belongs; wherein the partition mode category includes an incremental partition mode category and a full partition mode category;
[0030] Based on the table name, the business category to which the source data table belongs is extracted.
[0031] Optionally, the table information identification unit is specifically used to:
[0032] According to the order of priority of each candidate subject domain in the candidate subject domain set from high to low, the table name and the field information are matched with the keyword associated with each candidate subject domain in turn;
[0033] If the matching degree between the table name and the field information and the currently matching candidate subject domain is greater than the set matching degree threshold and meets the set requirements of the currently matching candidate subject domain, the currently matching candidate subject domain is determined as the subject domain category to which the source data table belongs.
[0034] Optionally, the table information identification unit is specifically used to:
[0035] Performing text preprocessing on the table name and the field information to obtain multiple candidate words;
[0036] Performing word vectorization on the multiple candidate words respectively to obtain word vectors corresponding to the multiple candidate words;
[0037] Based on the word vectors corresponding to the multiple candidate words, at least one keyword is determined from the multiple candidate words, and a table vector of the source data table is determined based on the at least one keyword;
[0038] Determine at least one candidate data table from the candidate data tables based on the similarity between the table vector of the source data table and the table vectors corresponding to the candidate data tables;
[0039] Based on the subject domain category corresponding to each of the at least one candidate data table, the subject domain category to which the source data table belongs is determined.
[0040] Optionally, the table information identification unit is specifically used to:
[0041] Extracting the initial business system name and the initial business name from the table name;
[0042] Standardizing the initial business system name to obtain a corresponding standard business system name;
[0043] The initial service name is standardized to obtain a corresponding standard service name.
[0044] Optionally, the data item naming unit is specifically used for:
[0045] For each field information, the following operations are performed respectively to generate the standard data items of each field information in the standardization table:
[0046] For a piece of field information, if the data element matching result corresponding to the piece of field information is a name, determining that the standard data item corresponding to the piece of field information is a corresponding source data item in the source data table;
[0047] If the data element matching result corresponding to the field information is not a name, determining whether there is a corresponding qualifier for the field information;
[0048] If there is a qualifier, then based on the result of matching the corresponding qualifier with the data element, determine the standard data item corresponding to the one field information;
[0049] If there is no qualifier, the standard data item corresponding to the one field information is determined based on the corresponding data element matching result.
[0050] Optionally, the data item naming unit is further used to:
[0051] Determine whether the one field information is a business time field;
[0052] If the one field information is a business time field, based on the representation type of the one field information, adding a type identifier of a corresponding representation type to the standard data item corresponding to the one field information;
[0053] If the one field information is a non-business time field, it is determined whether there is duplication in each standard data item, and if there is duplication, a distinguishing mark is added to the repeated standard data item.
[0054] On the one hand, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the above methods when executing the computer program.
[0055] In one aspect, a computer storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the steps of any of the above methods are implemented.
[0056] In one aspect, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps of any of the above methods.
[0057] In an embodiment of the present application, on the one hand, table information is identified based on the original table information to determine the table category corresponding to the source data table, such as the subject domain category, business category and partition method category, and then based on the table category, the standard table name of the standardized table corresponding to the source data table is automatically generated; on the other hand, the business time field contained in the source data table can be determined through the original table information of the source data table to be standardized and the data element benchmarking result of the source data table, and then the various standard data items of the standardized table are automatically generated based on the data element benchmarking result, the original table information and the business time field, thereby combining the automatic identification of the source data table to achieve automatic standardization of field names and table names, avoiding the time-consuming and labor-intensive naming of field names and table names of standardized tables caused by manual standardization, and improving the efficiency of data table standardization. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0059] Figure 1 A schematic diagram of an application scenario provided for an embodiment of the present application;
[0060] Figure 2 A schematic diagram of a process for standardizing a data table provided in an embodiment of the present application;
[0061] Figure 3 A flowchart of a data table standardization process provided in an embodiment of the present application;
[0062] Figure 4 A schematic diagram of a process for identifying a subject domain based on a priority ranking method provided in an embodiment of the present application;
[0063] Figure 5A schematic diagram of a process for implementing subject domain classification of a source data table using a classification model provided in an embodiment of the present application;
[0064] Figure 6 A schematic diagram of a process for generating standard data items provided in an embodiment of the present application;
[0065] Figure 7 A schematic diagram of a structure of a data table standardization device provided in an embodiment of the present application;
[0066] Figure 8 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0067] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other arbitrarily. In addition, although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that here.
[0068] It can be understood that in the specific implementation of the present application, related data such as data tables to be standardized are involved. When the above embodiments of the present application are applied to specific products or technologies, if user data tables are involved, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0069] To facilitate understanding of the technical solutions provided in the embodiments of the present application, some key terms used in the embodiments of the present application are explained here:
[0070] Business time field and non-business time field: Business time field refers to the field related to the actual business, such as transfer time, etc. Non-business time field refers to other fields except business time field. Non-business time field can be the time related to the operation of the data table, such as storage time, update time, modification time, deletion time, creation time, partition time, collection time, entry time, import time and addition time, etc. In actual applications, due to different businesses, the business time field may be different, so the non-business time field exclusion method can be used to determine whether a field is a business time field.
[0071] Subject domain: Subject domain usually refers to a collection of closely related data themes. These data themes can be divided into different subject domains according to the business focus. Each subject domain can have multiple themes. For example, the subject domain can be a relationship subject domain, a track subject domain, a person subject domain, an address subject domain, an item subject domain, an event subject domain, and an organization subject domain. Taking the relationship subject domain as an example, it mainly involves a collection of themes related to relationships. For example, these themes can include keywords that can represent a relationship, such as relationship, association, contact, address book, friend information, group information, case transfer, person case, property rights, branch, and marriage.
[0072] Partitioning method: Generally speaking, partitioning methods can include incremental partitioning or full partitioning. Correspondingly, the types of data tables include incremental tables and full tables. Incremental tables are when the amount of data is too large. Data tables can be stored or distributed incrementally. Specifically, when storing or distributing data, only incremental data is involved, and the data of the entire data table is not involved. In the case of a full table, regardless of whether the data has not changed, it must be stored or distributed again. Each time the data is stored or distributed, all the data is distributed. For example, it can be divided into incremental tables and full tables according to the data stored every day and whether it is partitioned by day. Then the full table will store all the data of each day, and the incremental table will store the data that increases every day compared to the previous day. Usually, the incremental table and the full table have different table suffixes. For example, the full table can be identified with the suffix _df, and the incremental table can be indicated with the suffix _di.
[0073] In practical applications, the partitioning method to be used can be selected according to the type of data. For example, for registration data, since the integrity of the data needs to be ensured, the full partitioning method can be used, while for trajectory or perception data, the focus is more on the current data, so the incremental partitioning method can be used.
[0074] Representation type: refers to different ways of expressing the time field. Generally speaking, the representation type can include time type, character type and integer type. These three representation types are different ways of expressing time, but the time or date they express is the same.
[0075] The following is a brief introduction to the design concept of the embodiments of the present application.
[0076] At present, in order to more conveniently put massive data into the research process and explore the value of data, data standardization is essential.
[0077] However, the current standardization process is usually adjusted manually, especially the naming of field names and table names of standardized tables is time-consuming and labor-intensive, and the efficiency is extremely low. Therefore, it is very necessary to be able to automatically standardize field names and table names.
[0078] In view of this, an embodiment of the present application provides a method based on data table standardization. In this method, on the one hand, table information is identified based on the original table information to determine the table category corresponding to the source data table, such as the subject domain category, business category and partition method category, and then based on the table category, the standard table name of the standardized table corresponding to the source data table is automatically generated. On the other hand, the business time field contained in the source data table can be determined through the original table information of the source data table to be standardized and the data element benchmarking result of the source data table, and then the various standard data items of the standardized table are automatically generated based on the data element benchmarking result, the original table information and the business time field. Combined with the automatic identification of the source data table, the automatic standardization of field names and table names is realized, avoiding the time-consuming and labor-intensive naming of the field names and table names of the standardized table caused by manual standardization, and improving the efficiency of data table standardization. Therefore, in the massive data scenario, the standardization process can be completed quickly, and massive data can be more conveniently put into the research process to mine data value.
[0079] In addition, the embodiments of the present application realize automatic generation of standardized table structures through algorithm integration, and the algorithms involved include business time field recognition algorithm, data subject domain recognition algorithm, partition mode recognition algorithm, business system and business name extraction algorithm, standardized table naming algorithm and data item naming algorithm. By automatically generating standardized table structures, it is possible to realize automatic modeling of standard data and generate standardized table names and data item names.
[0080] The following briefly introduces the application scenarios to which the technical solutions of the embodiments of the present application can be applied. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application and are not limited. In the specific implementation process, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.
[0081] The solution provided in the embodiment of the present application can be applied to data standardization scenarios of most business systems, such as public security business data standardization, administrative business data standardization, and office business data standardization. Figure 1 As shown, it is a schematic diagram of an application scenario provided by an embodiment of the present application. In this scenario, it can include a terminal device 101, a data table standardization device 102 and a database 103.
[0082] The terminal device 101 may be, for example, a mobile phone, a tablet computer (PAD), a laptop computer, a desktop computer, a smart TV, a smart vehicle-mounted device, and a smart wearable device. The terminal device 101 may be installed with a search application, and the application involved in the embodiment of the present application may be a software client, or a client such as a web page or a small program, and the server is a background server corresponding to the software or web page, small program, etc., without limiting the specific type of the client.
[0083] The data table standardization device 102 can execute the steps of the data table standardization method provided in the embodiment of the present application to realize the data table standardization function. For example, it can be a terminal device with a certain computing power, or an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, i.e., content delivery networks (CDNs), as well as big data and artificial intelligence platforms, but is not limited to these.
[0084] The data table standardization device 102 may include one or more processors 1021, a memory 1022, and an I / O interface 1023 for interacting with a terminal device, etc. The memory 1022 of the data table standardization device 102 may also store program instructions of the data table standardization method provided in the embodiment of the present application, and when these program instructions are executed by the processor 1021, they can be used to implement the steps of the data table standardization method provided in the embodiment of the present application, so as to implement the data table standardization process.
[0085] The database 103 may be a database of any structure, and may be used to store source data tables to be standardized and standardized tables obtained by standardization.
[0086] Specifically, the user terminal device 101 can pre-specify the storage location of the source data table that needs to be standardized in the database 103, and then the data table standardization device 102 can obtain the source data table from the database 103, and then standardize the source data table, and store the obtained standardized table in the database 103.
[0087] In one embodiment, when a user wants to search for data, he or she may input a search keyword through a search application in the terminal device 101 , and the database 103 may perform a corresponding search in the standardized table based on the search keyword.
[0088] In one implementation, the data of each standardized table obtained may also be used as training text for training a specific business model for use in actual business scenarios.
[0089] In one implementation, data statistics can also be performed based on the obtained standardized tables. Since similar data items are unified as the same data item after standardization, similar data in statistics will be counted in the same item, thereby improving the accuracy of statistical data.
[0090] The terminal device 101, the data table standardization device 102 and the database 103 may be directly or indirectly connected to each other through one or more networks. The network may be a wired network or a wireless network, for example, the wireless network may be a mobile cellular network, or a wireless fidelity (Wireless-Fidelity, WIFI) network, and of course, other possible networks, which are not limited in the embodiment of the present invention.
[0091] It should be noted that in the embodiment of the present application, the number of terminal devices 101 can be one or more, and similarly, the number of data table standardization devices 102 can be one or more, that is, there is no limitation on the number of terminal devices 101 or data table standardization devices 102.
[0092] In a possible application scenario, the relevant data (such as data tables, etc.) involved in the embodiments of the present application can be stored using cloud storage technology. Cloud storage is a new concept that extends and develops from the concept of cloud computing. A distributed cloud storage system refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (or storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.
[0093] In a possible application scenario, in order to reduce the communication delay of retrieval, the database 103 can deploy corresponding servers in various regions, or for load balancing, different servers can serve terminal devices 101 in different regions respectively. For example, the terminal device 101 is located at location a and establishes a communication connection with the server at service location a; the terminal device 101 is located at location b and establishes a communication connection with the server at service location b. Multiple servers form a data sharing system to achieve data sharing through blockchain.
[0094] Each server in the data sharing system has a node identifier corresponding to the server. Each server in the data sharing system can store the node identifiers of other servers in the data sharing system, so that the generated blocks can be broadcast to other servers in the data sharing system according to the node identifiers of other servers. A node identifier list can be maintained in each server, and the server name and node identifier are stored in the node identifier list accordingly. The node identifier can be the Internet Protocol (IP) address of the interconnection protocol between networks and any other information that can be used to identify the node.
[0095] Of course, the method provided in the embodiment of the present application is not limited to Figure 1The application scenario shown or Figure 2 The architecture can also be used in other possible application scenarios, which are not limited by the embodiments of the present application. Figure 1 The functions that can be implemented by each device in the application scenario shown will be described in the subsequent method embodiments, and will not be described in detail here.
[0096] The method flow provided in each embodiment of the present application can be Figure 1 The data table standardization device 102 or the terminal device 101 may be used to execute the process, or the data table standardization device 102 and the terminal device 101 may be used to execute the process together. Here, the process is mainly described by taking the data table standardization device 102 as an example.
[0097] See also Figure 2 , which is a flow chart of the data table standardization method provided in an embodiment of the present application.
[0098] Step 201: Based on the original table information of the source data table to be standardized and the data element matching result of the source data table, determine the business time field contained in the source data table.
[0099] Step 202: Identify table information based on the original table information to determine the table category corresponding to the source data table; wherein the table category includes subject domain category, business category, and partitioning method category.
[0100] Step 203: Based on the table category, generate a standard table name of the standardized table corresponding to the source data table.
[0101] Step 204: Generate various standard data items of the standardized table based on the data element matching results, original table information and business time field.
[0102] Step 205: Obtain a standardized table based on the standard table name and each standard data item.
[0103] In an embodiment of the present application, on the one hand, table information is identified based on the original table information to determine the table category corresponding to the source data table, such as the subject domain category, business category and partition method category, and then based on the table category, the standard table name of the standardized table corresponding to the source data table is automatically generated. On the other hand, the business time field contained in the source data table can be determined through the original table information of the source data table to be standardized and the data element benchmarking result of the source data table, and then the various standard data items of the standardized table are automatically generated based on the data element benchmarking result, the original table information and the business time field. Combined with the automatic identification of the source data table, automatic standardization of field names and table names is achieved, avoiding the time-consuming and labor-intensive naming of field names and table names of standardized tables caused by manual standardization, and improving the efficiency of data table standardization. Therefore, in massive data scenarios, the standardization process can be completed quickly, and massive data can be more conveniently put into the research process to mine data value.
[0104] See also Figure 3 As shown, it is a flowchart of the data table standardization process provided in an embodiment of the present application.
[0105] Step 301: Obtain a source data table to be standardized.
[0106] In an embodiment of the present application, the storage location of the source data table that needs to be standardized can be pre-specified. For example, a list of source data tables and the storage location of each source data table in the list can be given, and then before executing the standardization process, the corresponding source data table can be obtained based on the specified storage location.
[0107] In one implementation, data table standardization can be integrated into a standardized system platform. When data table standardization is required, the source table data to be standardized can be connected to the system platform, and table queries can be performed by calling the interface provided by the system. For example, the table name of the source table can be entered to perform source table queries, thereby obtaining the source data table to be standardized.
[0108] Step 302: Automated benchmarking is performed on the source data table through an automated benchmarking function to obtain data element benchmarking results of the source data table.
[0109] In the embodiments of the present application, the table names and each data item in the standardized table should be able to correspond to the standard data, while the data in the source data table may not be able to correspond to the standard data. For example, a different expression from the standard is used for data with the same meaning. For example, the "customer number" specified in the standard expresses the meaning of the customer's identification, while other expressions may be used in the source data table, such as "customer unified number", "customer number", "customer ID", etc. Although the meanings expressed are the same, this is not conducive to subsequent data mining and may bring certain recognition obstacles to computer processing. Therefore, it is necessary to standardize the data and use a unified expression. Therefore, before generating the standardized table, the source data table needs to be benchmarked to map all the fields included in the source data table into standard expressions.
[0110] Specifically, information extraction can be performed on the source data table to obtain original table information. For example, the original table information may include the table name and field information of the source data table. In addition, it may also include table attribute information of the source data table, such as the partitioning method of the table, or it may also include the representation type corresponding to each field information in the source data table. For example, for the time field, it may include representation types such as time type, character type, and integer type. Of course, other possible information can also be extracted, and the embodiments of the present application are not limited to this.
[0111] For each field information obtained, benchmarking is performed separately to determine the data element benchmarking results corresponding to each field information. The data element benchmarking results are the results of tool operations on the source table. The result file is a mapping relationship. The data element benchmarking results include the data elements and qualifiers corresponding to each field information. Among them, taking the above-mentioned integration of data table standardization into a standardized system platform as an example, the benchmarking function can also be integrated into an automated function of the standardized system platform. When benchmarking is required, the function call interface of the standardized system platform can be called to execute it.
[0112] In an embodiment of the present application, after extracting the original table information and obtaining the corresponding data element benchmarking results, the original table information and the data element benchmarking results can be used to automatically generate a standardized table structure, which is specifically described as follows.
[0113] Step 303: Based on the original table information of the source data table to be standardized and the data element matching result of the source data table, determine the business time field contained in the source data table.
[0114] In specific implementation, the business time field recognition algorithm can be integrated into the standardized system platform, so that when the business time field needs to be recognized, the business time field recognition algorithm can be called to implement the recognition process.
[0115] Specifically, the time field identifier can be set in advance, so that the time field contained in the source data table can be determined according to the Chinese field information in the original table information, the data element matching result and the set time field identifier. Among them, the characteristics of the data element include the above-mentioned time field identifier, or called the expression word, so that when the expression word of the data element is "time", "date", "date time", etc., it indicates that the corresponding field is a time field, and then the entire source data table can be queried based on this, so as to filter out all the time fields in the source data table.
[0116] In the specific implementation process, based on the actual business, the time field can be divided into business time field and non-business time field. The business time field refers to the field related to the actual business, such as transfer time, etc. The non-business time field refers to other fields except the business time field. The non-business time field can be, for example, the time related to the operation of the data table. Due to different businesses, the business time field may be different. Therefore, the method of directly setting rules to determine the business time field may not be accurate when the business changes. However, the type of non-business time field is usually relatively fixed. Therefore, the detection rules of the non-business time field are pre-set to detect the non-business time field from the found time field, and then the non-business time field exclusion method can be used to determine whether a field is a business time field.
[0117] For example, the non-business time field set is set to include time fields such as storage time, update time, modification time, deletion time, creation time, partition time, collection time, entry time, import time and addition time. Then, when the time field obtained by the above filtering contains a word in the set non-business time field set, it will be filtered out from the time field contained in the source data table, and the remaining time field after filtering will be determined as the business time field.
[0118] In an embodiment of the present application, a determined business time field is identified, for example, corresponding attributes of the business time field may be added or a label may be added to indicate that the field is a business time field. At the same time, the business time field may also be standardized to facilitate subsequent data item naming operations.
[0119] Specifically, considering that the business time field usually expresses business-related time and is relatively more important, for each determined business time field, if there is a business time field that does not correspond to all representation types, the missing representation type can be supplemented. In other words, the business time field is automatically processed, and when there is a missing representation type, the missing representation type is supplemented. For example, when the representation type includes three types: time type, character type, and integer type, and the business time field only has one representation type, the other two types are supplemented to facilitate subsequent retrieval in any manner.
[0120] For each non-business time field, one representation type can be specified to be retained, so if there is a non-business time field including other representation types except the specified representation type, the other representation types except the specified representation type will be deleted. For example, when the non-business time field is specified to be represented by the time type, if the non-business time field has multiple representation types, only the time type will be retained and the other representation types will be deleted.
[0121] In the embodiment of the present application, the process of performing table information identification based on the original table information to obtain the table category corresponding to the source data table may specifically include the following steps 304 to 306, which are introduced one by one below.
[0122] Step 304: Perform subject domain identification based on the table name and field information in the original table information to determine the subject domain category to which the source data table belongs.
[0123] In an embodiment of the present application, the subject domain identification algorithm can be integrated into a standardized system platform, so that when the subject domain needs to be identified, the subject domain identification algorithm can be called to implement the identification process, that is, the subject involved in the source data table can be reflected based on the table name and the content involved in the field information, and therefore the subject domain to which the source data table belongs can be identified based on the table name and field information.
[0124] In one implementation, the table name and field information may be matched with the keywords associated with each candidate subject domain in the candidate subject domain set in descending order of priority. If the degree of match between the table name and field information and the currently matching candidate subject domain is greater than a set matching degree threshold and meets the set requirements of the currently matching candidate subject domain, the currently matching candidate subject domain is determined to be the subject domain category to which the source data table belongs.
[0125] See Table 1 below for some examples of subject domain categories. In Table 1, the priorities from high to low are relationship subject domain, trajectory subject domain, person subject domain, address subject domain, object subject domain, event subject domain, and organization subject domain. Among them, the trajectory subject domain involves the time of trajectory occurrence, so the source data table must contain the business time field, and the address subject domain involves the specific location, so the source data table must contain the latitude and longitude fields.
[0126] Priority Subject Domain Category Other requirements 1 Relationship Subject Domain 2 Track subject domain Field requirements in the table: including business time field 3 People Subject Domain 4 Address Subject Field Field requirements in the table: Contains latitude and longitude fields 5 Item Subject Field 6 Event subject field 7 Organization Subject Domain
[0127] Table 1
[0128] Combined with the priority ranking in Table 1 above, see Figure 4 The figure shows a flowchart of identifying subject domains based on priority sorting.
[0129] S401: Determine whether the source data table belongs to the relational subject domain.
[0130] Specifically, the relationship subject domain mainly involves a table that represents a mutual relationship. For example, a data table belonging to the system subject domain may contain keywords that can represent mutual relationships, such as relationship, association, contact, address book, friend information, group information, case transfer, human case, property rights, branch, and marriage.
[0131] Then, for the source data table, the name and field information in the source data table can be matched with the keywords involved in the relational subject domain, and whether the source data table belongs to the relational subject domain can be determined based on the matching degree.
[0132] S402: If the judgment result of S401 is no, it is determined whether the source data table belongs to the track subject domain.
[0133] When the matching degree between the name and field information in the source data table and the relational subject domain is greater than the set matching degree threshold, the source data table is considered to belong to the relational subject domain.
[0134] If the matching degree between the name and field information in the source data table and the relationship subject domain is not greater than the set matching degree threshold, it is considered that the source data table does not belong to the relationship subject domain, and then it is sorted according to the priority to continue to determine whether the source data table belongs to the trajectory subject domain.
[0135] Specifically, the trajectory subject domain mainly involves representing a table that reflects the historical trajectory. For example, the data table belonging to the trajectory subject domain may include keywords that can represent the historical trajectory, such as trajectory, record, accommodation, registration, payment, consumption, rental, snapshot, ticket booking, entry, intercom, and express delivery.
[0136] Similarly, for the source data table, the name and field information in the source data table can be matched with the keywords involved in the trajectory subject domain, and it can be determined whether the name and field information in the source data table meets the requirements of the trajectory subject domain, and then whether the source data table belongs to the trajectory subject domain can be determined based on the matching degree and whether the requirements are met.
[0137] S403: If the judgment result of S402 is no, it is judged whether the source data table belongs to the human subject domain.
[0138] When the matching degree between the name and field information in the source data table and the trajectory subject domain is greater than the set matching degree threshold, and meets the requirements of the trajectory subject domain, that is, the source data table contains the business time field, then the source data table is considered to belong to the trajectory subject domain.
[0139] If the matching degree between the name and field information in the source data table and the trajectory subject domain is not greater than the set matching degree threshold, or meets the requirements of the trajectory subject domain, that is, the source data table does not contain the business time field, then it is considered that the source data table does not belong to the trajectory subject domain, and then it is sorted according to the priority to continue to determine whether the source data table belongs to the person subject domain.
[0140] Specifically, the people subject domain mainly involves representing a table that reflects information related to people. For example, the data table belonging to the people subject domain may contain population, people, personnel, basic personnel information, personnel information, social security information, disabled persons' federation information, provident fund information, or personnel title attribute words and other keywords that can represent personnel information. For example, personnel title attribute words include appraiser, employee, customer, user, teacher, civil servant, driver, car owner, Internet user, student, member, shareholder, list and tour guide, etc.
[0141] Similarly, for the source data table, the name and field information in the source data table can be matched with the keywords involved in the human subject domain, and whether the source data table belongs to the human subject domain can be determined based on the degree of match.
[0142] S404: If the determination result of S403 is no, determine whether the source data table belongs to the address subject domain.
[0143] When the matching degree between the name and field information in the source data table and the person subject domain is greater than the set matching degree threshold, the source data table is considered to belong to the person subject domain.
[0144] If the match between the name and field information in the source data table and the person subject domain is not greater than the set match threshold, it is considered that the source data table does not belong to the person subject domain, and then it is sorted according to priority to continue to determine whether the source data table belongs to the address subject domain.
[0145] Specifically, the address subject domain mainly involves representing a table that reflects address information. For example, a data table belonging to the address subject domain may include keywords that can represent address information, such as longitude and latitude, points, sites, places, addresses, hostels, hotels, Internet cafes, and hospitals.
[0146] Similarly, for the source data table, the name and field information in the source data table can be matched with the keywords involved in the address subject domain, and it can be determined whether the name and field information in the source data table meets the requirements of the address subject domain, that is, whether the source data table contains longitude and latitude fields, and then whether the source data table belongs to the address subject domain can be determined based on the matching degree and whether the requirements are met.
[0147] S405: If the determination result of S404 is no, determine whether the source data table belongs to the item subject domain.
[0148] When the match between the name and field information in the source data table and the address subject domain is greater than the set match threshold, and the source data table contains longitude and latitude fields, the source data table is considered to belong to the address subject domain.
[0149] If the match between the name and field information in the source data table and the address subject domain is not greater than the set match threshold, or if the source data table does not contain longitude and latitude fields, it is considered that the source data table does not belong to the address subject domain. Then, according to the priority sorting, it is further determined whether the source data table belongs to the item subject domain.
[0150] Specifically, the item subject domain mainly involves representing a table that reflects information related to items. For example, the data table belonging to the item subject domain may include electric vehicles, motor vehicles, items, property, decisions, certificates, hot spots, terminals, equipment, documents, card information, base stations, databases, hardware, home appliances, vehicles, channels, and other keywords that can represent items.
[0151] Similarly, for the source data table, the name and field information in the source data table can be matched with the keywords involved in the item subject domain, and whether the source data table belongs to the item subject domain can be determined based on the degree of match.
[0152] S406: If the judgment result of S405 is no, it is determined whether the source data table belongs to the event subject domain.
[0153] When the matching degree between the name and field information in the source data table and the item subject domain is greater than the set matching degree threshold, the source data table is considered to belong to the item subject domain.
[0154] If the matching degree between the name and field information in the source data table and the item subject domain is not greater than the set matching degree threshold, it is considered that the source data table does not belong to the item subject domain, and then it is sorted according to priority to continue to determine whether the source data table belongs to the event subject domain.
[0155] Specifically, the event subject domain mainly involves representing a table that reflects the events that have occurred. For example, a data table belonging to the event subject domain may include changes, cases, police information, penalties, judgments, violations, permits, causes of action, measures, rules, statistical information, and other keywords that can represent the occurrence of events.
[0156] Similarly, for the source data table, the name and field information in the source data table can be matched with the keywords involved in the event subject domain, and whether the source data table belongs to the event subject domain can be determined based on the degree of match.
[0157] S407: If the judgment result of S406 is no, it is determined whether the source data table belongs to the organization subject domain.
[0158] When the matching degree between the name and field information in the source data table and the event subject domain is greater than the set matching degree threshold, the source data table is considered to belong to the event subject domain.
[0159] If the match between the name and field information in the source data table and the event subject domain is not greater than the set match threshold, it is considered that the source data table does not belong to the event subject domain, and then it is sorted according to priority to continue to determine whether the source data table belongs to the organization subject domain.
[0160] Specifically, the organizational subject domain mainly involves representing a table that reflects the events that have occurred. For example, the data table belonging to the organizational subject domain may include keywords such as changes, cases, police information, penalties, judgments, violations, permits, causes of action, measures, rules, statistical information, etc. that can represent the occurrence of events.
[0161] Similarly, for the source data table, the name and field information in the source data table can be matched with the keywords involved in the organization subject domain, and whether the source data table belongs to the organization subject domain can be determined based on the matching degree. When the matching degree between the name and field information in the source data table and the organization subject domain is greater than the set matching degree threshold, it is considered that the source data table belongs to the organization subject domain. If the matching degree between the name and field information in the source data table and the organization subject domain is not greater than the set matching degree threshold, it is considered that the source data table does not belong to the organization subject domain, and the process ends.
[0162] In another implementation, a classification model may also be used to implement subject domain classification of the source data table.
[0163] For details, see Figure 5As shown, it is a schematic diagram of the process of using the classification model to implement the subject domain classification of the source data table. Here, the K-nearest neighbor (KNN) model is specifically introduced as an example. In practical applications, other possible classification models can also be used to implement the classification of the subject domain, and the embodiment of the present application does not limit this.
[0164] S501: Perform text preprocessing on the table name and field information to obtain multiple candidate words.
[0165] Specifically, the text preprocessing process is the process of extracting keywords to represent the text. For Chinese text preprocessing, it mainly includes two stages: text segmentation and stop word removal. After text segmentation and stop word removal, keywords are formed for subsequent processing.
[0166] S502: Perform word vectorization on multiple candidate words respectively to obtain word vectors corresponding to the multiple candidate words.
[0167] Specifically, the purpose of word vectorization is to convert the keywords after text preprocessing into vector format. The accuracy of the vector determines the quality of subsequent subject domain classification.
[0168] In one implementation, a Bag Of Words (BOW) model or a Vector Space Model may be used to implement word vectorization. Of course, other possible vectorization models may also be used, and this embodiment of the present application does not limit this.
[0169] S503: Based on the word vectors corresponding to the multiple candidate words, determine at least one keyword from the multiple candidate words, and determine the table vector of the source data table based on the at least one keyword.
[0170] Since the source data table may include numerous words, some of which may be irrelevant to classification, it is possible to filter out keywords as the basis for subsequent subject domain classification, and it is also possible to reduce the number of words, thereby improving classification efficiency.
[0171] In one embodiment, feature extraction of the text representation method of the vector space model corresponds to the selection of feature items and the calculation of feature weights. The basic idea of feature selection is to sort by word frequency, select the feature items with the highest scores, and filter out the rest.
[0172] Specifically, the feature value may be calculated using the following formula, namely, term frequency-inverse document frequency (TF-IDF). The larger the TF-IDF value is, the greater the probability that the word is a keyword.
[0173]
[0174]
[0175] TF-IDF=TF*IDF
[0176] The IDF denominator is added with 1 to avoid calculation errors caused by the denominator being 0.
[0177] Then, the vector of the source data table is calculated based on the filtered keywords. Generally speaking, the sample vector should be able to achieve a center distance as small as possible with similar samples and a center distance as large as possible with heterogeneous samples.
[0178] S504: Determine at least one candidate data table from each candidate data table based on the similarity between the table vector of the source data table and the table vectors corresponding to each candidate data table.
[0179] In an embodiment of the present application, the cosine value of the vector angle can be used to measure the similarity, so as to select at least one candidate data table that is closest to the source data table from the candidate data table, for example, K candidate data tables are selected, and the candidate data tables are data tables with determined subject domain categories, or data tables with determined probabilities of each subject domain category to which they belong.
[0180] S505: Determine the subject domain category to which the source data table belongs based on the subject domain category corresponding to each of the at least one candidate data table.
[0181] Specifically, the weight of at least one candidate data table in each subject domain category is calculated in turn, and the largest weight is selected, or the weighted sum is performed to assist in determining the subject domain category to which the source data table belongs. Alternatively, a voting mechanism may be used, where when the candidate data table belongs to subject domain A, subject domain A increases by one vote, and the subject domain category with the highest number of votes in at least one candidate data table is selected as the subject domain category to which the source data table belongs.
[0182] Step 305: Identify the partition mode based on the table name and field information in the original table information to determine the partition mode category to which the source data table belongs.
[0183] In the embodiment of the present application, the partition mode recognition algorithm can be integrated into the standardized system platform, so that when the partition mode needs to be recognized, the partition mode recognition algorithm can be called to implement the recognition process.
[0184] Specifically, the partitioning method category may include an incremental partitioning category and a full partitioning category, which may be identified based on the source table name and Chinese fields to identify whether the source data table is registration data or trajectory and perception data. If it is registration data, the partitioning method category is a full partitioning category, and the standardized table belongs to the full table. The suffix may be a representation suffix of the full table, such as _di. If it is trajectory and perception data, the partitioning method category is an incremental partitioning category, and the standardized table belongs to the incremental table. The suffix may be a representation suffix of the incremental table, such as _df.
[0185] Step 306: Based on the table name in the original table information, extract the business category to which the source data table belongs.
[0186] In an embodiment of the present application, the business system and business name extraction algorithm can be integrated into a standardized system platform, so that when the business system and business name need to be extracted, the business system and business name extraction algorithm can be called to implement the identification process.
[0187] The service category may include a service system name and a service name.
[0188] Specifically, the initial business system name and the initial business name contained in the table name can be extracted, and the name may not correspond to the standard, so the initial business system name can be standardized to obtain the corresponding standard business system name, and the initial business name can be standardized to obtain the corresponding standard business name. In other words, in actual application, the ideal extraction result should include the business system and business name, but if the input source table name does not completely contain the extracted content, partial extraction can be performed, and the business name can be optimized during the extraction process. For example, the keyword "xx table" in the table name can be removed, and the table name "xx information" can be completed. For example, the population information table can be optimized to population information, and the resident population table can be optimized to resident population information.
[0189] Step 307: Based on the identified subject domain category, partition method category, and business category, a standard table name of a standardized table corresponding to the source data table is generated.
[0190] Specifically, based on the subject domain category, partition method category and business category obtained in the above processes, a standard table name can be generated in the format of "dwd_subject domain_business system_business name_partition suffix". Of course, in actual application, it is not limited to this format. Other systems can be set according to user needs, and the embodiments of the present application do not limit this.
[0191] Based on the above process, the automated construction indicated by the standard can be realized. The above process can be standardized for various business systems, has high applicability and a wide range of applications.
[0192] Step 308: Generate various standard data items of the standardized table based on the data element matching results, original table information and business time field.
[0193] In the embodiment of the present application, standard data items can be automatically generated based on the data element matching results and the qualifier extraction results. Figure 6 As shown in FIG. 1 , a schematic diagram of a process of generating corresponding standard data items is shown in FIG. 1 , which takes field information A as an example. The generation of standard data items includes two stages, namely, Figure 6 The basic item basic naming stage and the business field naming and whole table verification stage shown in the figure include steps S601 to S605 , and the business field naming and whole table verification stage include steps S606 to S608 .
[0194] S601: For field information A, determine whether the data element matching result is "name".
[0195] S602: If S601 is yes, that is, if the data element matching result corresponding to the field information A is a name, it is determined that the standard data item corresponding to the field information A is the corresponding source data item in the source data table, that is, the standard data item is a "source data item".
[0196] S603: If S601 is no, that is, if the data element matching result corresponding to field information A is not a name, determine whether there is a corresponding qualifier in field information A. For example, "mother" in "mother's ID number" is a qualifier, which is used to limit the attributes of subsequent words.
[0197] S604: If S603 is yes, that is, if there is a qualifier, then based on the matching result of the corresponding qualifier and the data element, determine the standard data item corresponding to the field information A, for example, the standard data item is "qualifier_data element".
[0198] S605: If S603 is no, that is, if there is no qualifier, then based on the corresponding data element matching result, determine the standard data item corresponding to the field information A, for example, the standard data item is "data element".
[0199] S606: Determine whether the field information A is a business time field;
[0200] S607: If S606 is yes, that is, if the field information A is a business time field, based on the representation type of the field information A, a type identifier of a corresponding representation type is added to the standard data item corresponding to the field information A.
[0201] For example, add a "_time type" suffix to the date and time type, add a "_character type" suffix to the character type, and add a "_integer type" suffix to the integer type to distinguish the data types.
[0202] S608: If S606 is no, that is, if the field information A is a non-business time field, then the whole table is checked. The whole table check means that after the data items are named, if there are fields with the same data items, they need to be distinguished. For example, after the data items are named, if there are other data items with the same name as the standard data item corresponding to the field information A, then a distinguishing mark is added to the standard data item name corresponding to the field information A and the repeated data item, such as adding a digital number to distinguish them.
[0203] After the standard table name and the standard data items are obtained based on the above process, a standardized table can be obtained based on the standard table name and the standard data items.
[0204] Step 309: Obtain a standardized table based on the standard table name and each standard data item.
[0205] In the embodiment of the present application, the field names and table names can be automatically standardized through the above process. After standardization, a physical model of data modeling can be automatically generated on the system, and the standard table can be uploaded to the standard database for subsequent work needs.
[0206] In summary, in the embodiment of the present application, by inputting the table name of the source table, the system can automatically query this table, automatically benchmark it, and after benchmarking, the standardized naming of the data item field name and the table name can be automatically realized. Among them, the automatic generation of the standard table structure is realized by algorithm integration, and the algorithms involved include the business time field recognition algorithm, the data subject domain recognition algorithm, the partition method recognition algorithm, the business system and business name extraction algorithm, the standard table naming algorithm and the data item naming algorithm, and the automatic generation of the standard table structure can realize the automatic modeling of the standard data, generate the standard table name and the data item name, and no manual intervention is required in the process of standardizing the data table. The target data item can be automatically generated according to the data element benchmarking result and the qualifier extraction result. The degree of automation is high, which can solve the large amount of time and energy consumed by the prior art for data table standardization, and can greatly improve the time cost and manpower cost. And it can be modified according to the actual business needs, that is, it can be applied to each business system, and it can also be updated according to the needs of the needs, so as to realize the automatic generation of table names and field names.
[0207] See also Figure 7 Based on the same inventive concept, the embodiment of the present application further provides a data table standardization device 70, which includes:
[0208] The business field identification unit 701 is used to determine the business time field included in the source data table based on the original table information of the source data table to be standardized and the data element matching result of the source data table;
[0209] The table information identification unit 702 is used to identify the table information based on the original table information and determine the table category corresponding to the source data table; wherein the table category includes the subject domain category, the business category and the partitioning method category;
[0210] A table naming unit 703 is used to generate a standard table name of a standardized table corresponding to a source data table based on the table category;
[0211] A data item naming unit 704 is used to generate each standard data item of the standardized table based on the data element benchmarking result, the original table information and the business time field;
[0212] The standard table generating unit 705 is used to obtain a standardized table based on the standard table name and each standard data item.
[0213] Optionally, the device further includes an automated benchmarking unit 706, which is used to:
[0214] Extract information from the source data table to obtain original table information; wherein the original table information includes the table name and field information of the source data table;
[0215] The obtained fields of information are respectively subjected to a benchmarking process to determine the data element benchmarking results corresponding to the respective fields of information. The data element benchmarking results include the data elements and qualifiers corresponding to the respective fields of information.
[0216] Optionally, the business field identification unit 701 is specifically configured to:
[0217] Based on the Chinese field information in the original table information and the data element matching results, determine the time field included in the source data table;
[0218] Based on the set non-business time field set, filter out the non-business time fields in the time fields contained in the source data table;
[0219] The time fields remaining after filtering are determined as business time fields.
[0220] Optionally, the business field identification unit 701 is further configured to:
[0221] For each determined business time field, if there is a business time field that does not correspond to all representation types, the missing representation type is completed;
[0222] For each non-business time field, if there is a non-business time field including other representation types except the specified representation type, the other representation types are deleted.
[0223] Optionally, the table information identification unit 702 is specifically used to:
[0224] Identify subject domains based on table names and field information to determine the subject domain category to which the source data table belongs;
[0225] Identify the partitioning mode based on the table name and field information, and determine the partitioning mode category to which the source data table belongs; the partitioning mode categories include incremental partitioning category and full partitioning category;
[0226] Based on the table name, extract the business category to which the source data table belongs.
[0227] Optionally, the table information identification unit 702 is specifically used to:
[0228] According to the order of priority of each candidate subject domain in the candidate subject domain set from high to low, the table name and field information are matched with the keywords associated with each candidate subject domain in turn;
[0229] If the matching degree between the table name and field information and the currently matching candidate subject domain is greater than the set matching degree threshold and meets the set requirements of the currently matching candidate subject domain, the currently matching candidate subject domain is determined to be the subject domain category to which the source data table belongs.
[0230] Optionally, the table information identification unit 702 is specifically used to:
[0231] Perform text preprocessing on table names and field information to obtain multiple candidate words;
[0232] Perform word vectorization on multiple candidate words respectively to obtain word vectors corresponding to the multiple candidate words;
[0233] Based on the word vectors corresponding to the multiple candidate words, at least one keyword is determined from the multiple candidate words, and a table vector of the source data table is determined based on the at least one keyword;
[0234] Determine at least one candidate data table from each candidate data table based on the similarity between the table vector of the source data table and the table vectors corresponding to each candidate data table;
[0235] Based on the subject domain category corresponding to each of the at least one candidate data table, the subject domain category to which the source data table belongs is determined.
[0236] Optionally, the table information identification unit 702 is specifically used to:
[0237] Extract the initial business system name and the initial business name from the table name;
[0238] Standardizing the initial business system name to obtain a corresponding standard business system name;
[0239] The initial service name is standardized to obtain a corresponding standard service name.
[0240] Optionally, the data item naming unit 704 is specifically used to:
[0241] For each field information, perform the following operations respectively to generate the standard data items of each field information in the standardization table:
[0242] For a field information, if the data element matching result corresponding to the field information is a name, it is determined that the standard data item corresponding to the field information is a corresponding source data item in the source data table;
[0243] If the data element matching result corresponding to a field information is not a name, determine whether a field information has a corresponding qualifier;
[0244] If there is a qualifier, a standard data item corresponding to the field information is determined based on the corresponding qualifier and the data element matching result;
[0245] If there is no qualifier, a standard data item corresponding to the field information is determined based on the corresponding data element matching result.
[0246] Optionally, the data item naming unit 704 is specifically used to:
[0247] Determine whether a field information is a business time field;
[0248] If a field information is a business time field, based on the representation type of the field information, a type identifier of the corresponding representation type is added to the standard data item corresponding to the field information;
[0249] If the one field information is a non-business time field, it is determined whether there is duplication in each standard data item, and if there is duplication, a distinguishing mark is added to the repeated standard data item.
[0250] Through the above device, the standard table structure can be automatically generated through algorithm integration, and the standard data can be automatically modeled to generate standard table names and data item names. No manual intervention is required in the process of data table standardization. The target data items can be automatically generated based on the data element benchmarking results and qualifier extraction results. The high degree of automation can solve the large amount of time and energy consumed by the existing technology for data table standardization, and can greatly improve time and labor costs. It can also be modified according to actual business needs, which can be applied to various business systems, and can also be updated according to targeted needs to achieve automatic generation of table names and field names.
[0251] The device can be used to execute the methods shown in the various embodiments of the present application. Therefore, for the functions that can be implemented by the various functional modules of the device, reference can be made to the description of the aforementioned embodiments and no further details will be given.
[0252] See also Figure 8 Based on the same technical concept, the present application embodiment also provides a computer device 80, which can be Figure 1 As shown in the terminal device or server, the computer device 80 may include a memory 801 and a processor 802 .
[0253] The memory 801 is used to store computer programs executed by the processor 802. The memory 801 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the computer device, etc. The processor 802 may be a central processing unit (CPU), or a digital processing unit, etc. The specific connection medium between the memory 801 and the processor 802 is not limited in the embodiments of the present application. The embodiments of the present application are Figure 8 In the embodiment, the memory 801 and the processor 802 are connected via a bus 803. The bus 803 is connected to the processor 802 via a bus 803. Figure 8 The connection between other components is shown by bold lines, and is not intended to be limiting. The bus 803 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0254] The memory 801 may be a volatile memory, such as a random-access memory (RAM); the memory 801 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), or the memory 801 may be any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 801 may be a combination of the above memories.
[0255] The processor 802 is used to execute the method executed by the device in each embodiment of the present application when calling the computer program stored in the memory 801.
[0256] In some possible implementations, various aspects of the method provided in the present application may also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of the method according to the various exemplary implementations of the present application described above in this specification. For example, the computer device can execute the method executed by the device in each embodiment of the present application.
[0257] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0258] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0259] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A data table standardization method, characterized in that: The method comprises: Based on the original table information of the source data table to be standardized and the data element benchmarking result of the source data table, determine the business time field contained in the source data table; wherein the original table information includes the table name and field information of the source data table; Based on the original table information, table information is identified to determine the table category corresponding to the source data table; wherein the table category includes subject domain category, business category and partition mode category; Based on the table category, generate a standard table name of a standardized table corresponding to the source data table; For each field information included in the original table information, the following operations are performed respectively to generate respective standard data items for each field information: For a piece of field information, if the data element matching result corresponding to the piece of field information is a name, determining that the standard data item corresponding to the piece of field information is a corresponding source data item in the source data table; If the data element matching result corresponding to the field information is not a name, determining whether there is a corresponding qualifier for the field information; If there is a qualifier, then based on the result of matching the corresponding qualifier with the data element, determine the standard data item corresponding to the one field information; If there is no qualifier, determining the standard data item corresponding to the one field information based on the corresponding data element matching result; Determine whether the one field information is a business time field; If the one field information is a business time field, based on the representation type of the one field information, adding a type identifier of a corresponding representation type to the standard data item corresponding to the one field information; If the one field information is a non-business time field, determining whether there is duplication in each standard data item, and if there is duplication, adding a distinguishing mark to the repeated standard data item; The standardized table is obtained based on the standard table name and the standard data items of each field information.
2. The method according to claim 1, characterized in that Before determining the business time field included in the source data table based on the original table information of the source data table to be standardized and the data element matching result of the source data table, the method further includes: Extracting information from the source data table to obtain the original table information; A matching process is performed on each of the obtained field information to determine a matching result of a data element corresponding to each of the field information. The data element matching result includes a data element and a qualifier corresponding to each of the field information.
3. The method according to claim 2, characterized in that Based on original table information of a source data table to be standardized and a data element benchmarking result of the source data table, determining a business time field included in the source data table includes: Based on the Chinese field information in the original table information and the data element matching result, determining the time field included in the source data table; Based on the set non-business time field set, the non-business time fields in the time fields included in the source data table are filtered out; The time fields remaining after filtering are determined as business time fields.
4. The method according to claim 3, characterized in that After determining the business time field included in the source data table based on the original table information of the source data table to be standardized and the data element matching result of the source data table, the method further includes: For each of the determined service time fields, if there is a service time field that does not correspond to all representation types, the missing representation types are completed; For each of the non-business time fields, if there is a non-business time field including other representation types except the specified representation type, the other representation types are deleted.
5. The method according to claim 2, characterized in that Performing table information recognition based on the original table information to determine the table category corresponding to the source data table includes: Perform subject domain identification based on the table name and the field information to determine the subject domain category to which the source data table belongs; Based on the table name and the field information, a partition mode is identified to determine the partition mode category to which the source data table belongs; wherein the partition mode category includes an incremental partition mode category and a full partition mode category; Based on the table name, the business category to which the source data table belongs is extracted.
6. The method according to claim 5, characterized in that Performing subject domain identification based on the table name and the field information to determine the subject domain category to which the source data table belongs includes: According to the order of priority of each candidate subject domain in the candidate subject domain set from high to low, the table name and the field information are matched with the keyword associated with each candidate subject domain in turn; If the matching degree between the table name and the field information and the currently matching candidate subject domain is greater than the set matching degree threshold and meets the set requirements of the currently matching candidate subject domain, the currently matching candidate subject domain is determined as the subject domain category to which the source data table belongs.
7. The method according to claim 5, characterized in that Performing subject domain identification based on the table name and the field information to determine the subject domain category to which the source data table belongs includes: Performing text preprocessing on the table name and the field information to obtain multiple candidate words; Performing word vectorization on the multiple candidate words respectively to obtain word vectors corresponding to the multiple candidate words; Based on the word vectors corresponding to the multiple candidate words, at least one keyword is determined from the multiple candidate words, and a table vector of the source data table is determined based on the at least one keyword; Determine at least one candidate data table from the candidate data tables based on the similarity between the table vector of the source data table and the table vectors corresponding to the candidate data tables; Based on the subject domain category corresponding to each of the at least one candidate data table, the subject domain category to which the source data table belongs is determined.
8. The method according to claim 5, characterized in that Based on the table name, extracting the business category to which the source data table belongs, including: Extracting the initial business system name and the initial business name from the table name; Standardizing the initial business system name to obtain a corresponding standard business system name; The initial service name is standardized to obtain a corresponding standard service name.
9. A data table standardization device, characterized in that: The device comprises: A business field identification unit, configured to determine the business time field contained in the source data table based on the original table information of the source data table to be standardized and the data element matching result of the source data table; wherein the original table information includes the table name and field information of the source data table; A table information identification unit, used to identify the table information based on the original table information, and determine the table category corresponding to the source data table; wherein the table category includes a subject domain category, a business category, and a partitioning method category; A table naming unit, used for generating a standard table name of a standardized table corresponding to the source data table based on the table category; The data item naming unit is used to perform the following operations for each field information included in the original table information, and generate a standard data item for each field information: for a field information, if the data element matching result corresponding to the field information is a name, determine that the standard data item corresponding to the field information is the corresponding source data item in the source data table; if the data element matching result corresponding to the field information is not a name, determine whether there is a corresponding qualifier for the field information; if there is a qualifier, determine the standard data item corresponding to the field information based on the corresponding qualifier and data element matching result; if there is no qualifier, determine the standard data item corresponding to the field information based on the corresponding data element matching result; determine whether the field information is a business time field; if the field information is a business time field, add a type identifier of the corresponding representation type to the standard data item corresponding to the field information based on the representation type of the field information; if the field information is a non-business time field, determine whether there is duplication in each standard data item, and if there is duplication, add a distinguishing identifier to the repeated standard data item; The standard table generating unit is used to obtain the standardized table based on the standard table name and the standard data items of each field information.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
11. A computer storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer program product comprising computer program instructions, characterized in that When the computer program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Method and device for storing data and equipment
CN108664499A
Data benchmarking method and device and storage device
CN110795482A