Data table classification method and device, medium and electronic equipment

By obtaining the operation statements, key foreign key information and content characteristics of the data table, using the data table classification model to generate feature vectors and assign weights, the problem of non-standard named data table classification in the database is solved, and a more accurate and adaptable data table classification is achieved.

CN120336528AActive Publication Date: 2025-07-18TIANJIN TIANHE DIGITAL IND TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510813984.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

It is difficult for the prior art to efficiently and accurately classify non-standard named data tables in the database, resulting in inefficient data management and significantly increasing the difficulty of data retrieval.

Method used

By obtaining the operation statements, key foreign key information and content characteristics of the data table in the target time window, the data table classification model is used to generate the first, second and third feature vectors, and different weights are assigned according to the weight allocation layer, and finally the classification results are obtained at the classification layer.

Benefits of technology

It realizes a more accurate and objective classification of data tables, has stronger migration and adaptability, and is suitable for data tables with different internal data characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336528A_ABST
    Figure CN120336528A_ABST
Patent Text Reader

Abstract

The invention provides a data table classification method and device, a medium and electronic equipment, and relates to the field of data table classification, and the method comprises the steps: obtaining each operation statement of a to-be-classified data table in a target time window, so as to obtain a first feature vector YT; obtaining each piece of key foreign key information of the to-be-classified data table to obtain a second feature vector ET; obtaining content feature information of the to-be-classified data table to obtain a third feature vector; inputting the third feature vector into a weight distribution layer of a data table classification model to obtain a first weight corresponding to the YT and a second weight corresponding to the ET; obtaining a target feature vector according to the YT, the first weight, the ET and the second weight; and inputting the target feature vector into a classification layer of a data table classification model to obtain a classification result corresponding to the to-be-classified data table. The method has higher mobility and adaptability, and the obtained classification result is more accurate and objective.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In the big data era, the number of data tables stored in the database is huge and the types are complex. Efficient and accurate classification of data tables has become a key link to improve data management and retrieval efficiency. At present, the industry generally classifies data tables in the database based on rules such as standard file names and storage paths. When facing data tables with standardized naming, such methods can achieve fast and accurate classification, significantly improving data management efficiency.

[0003] However, with the increasing diversification of data sources and the diversification of data generation methods, there are a large number of data tables with non-standard named table names and field names widely existing in the database. Due to the lack of unified and standardized naming rules, these data tables are difficult to be effectively classified directly using the existing classification methods based on standard naming or paths, resulting in low data management efficiency and a significant increase in the difficulty of data retrieval, which greatly affects the overall efficiency of data processing and analysis. There is an urgent need to study new data table classification technologies to solve this problem. Summary of the Invention

[0004] In view of the above technical problems, the present application provides a data table classification method, device, medium and electronic device, which at least partially solve the problems existing in the prior art.

[0005] In the first aspect of the present application, a data table classification method is provided. The method includes: Obtain each operation statement of the data table to be classified within the target time window, and input it into the text encoding layer of the data table classification model to obtain the first feature vector YT; where YT represents the operation characteristics of the data table to be classified in chronological order within the target time window; the end time of the target time window is the current time; Obtain each key foreign key information of the data table to be classified, and according to the graph feature extraction layer of the data table classification model, obtain the second feature vector ET; where ET represents the data reference characteristics of the data table to be classified; the key foreign key information is the foreign key information between the data table to be classified and the key data table; the key data table and the data table to be classified are in the same database; and the key data table can determine the data table category according to the preset classification rules; the data table to be classified cannot determine the data table category according to the preset classification rules; Obtain the content feature information of the data table to be classified, and according to the graph feature extraction layer of the data table classification model, obtain the third feature vector; where the third feature vector represents the data content characteristics of the data table to be classified; Input the third feature vector into the weight assignment layer of the data table classification model to obtain the first weight corresponding to YT and the second weight corresponding to ET; According to YT, the first weight, ET and the second weight, obtain the target feature vector; Input the target feature vector into the classification layer of the data table classification model to obtain the classification result corresponding to the data table to be classified.

[0006] In a second aspect of the present application, a data table classification device is provided, and the device includes: A first vector acquisition unit, configured to acquire each operation statement of the data table to be classified within a target time window and input it into the text encoding layer of the data table classification model to obtain a first feature vector YT; where YT represents the operation characteristics of the data table to be classified in chronological order within the target time window; the end time of the target time window is the current time; A second vector acquisition unit, configured to acquire each key foreign key information of the data table to be classified and obtain a second feature vector ET according to the graph feature extraction layer of the data table classification model; where ET represents the data reference characteristics of the data table to be classified; the key foreign key information is the foreign key information between the data table to be classified and the key data table; the key data table and the data table to be classified are in the same database; and the key data table can determine the data table category according to a preset classification rule; the data table to be classified cannot determine the data table category according to the preset classification rule; A third vector acquisition unit, configured to acquire the content feature information of the data table to be classified and obtain a third feature vector according to the graph feature extraction layer of the data table classification model; where the third feature vector represents the data content characteristics of the data table to be classified; A weight acquisition unit, configured to input the third feature vector into the weight assignment layer of the data table classification model to obtain a first weight corresponding to YT and a second weight corresponding to ET; A target vector acquisition unit, configured to obtain a target feature vector according to YT, the first weight, ET, and the second weight; A classification unit, configured to input the target feature vector into the classification layer of the data table classification model to obtain the classification result corresponding to the data table to be classified.

[0007] In a third aspect of the present application, a non-transitory computer-readable storage medium is provided, and at least one instruction or at least one program segment is stored in the storage medium, and at least one instruction or at least one program segment is loaded and executed by a processor to implement the foregoing data table classification method.

[0008] In a fourth aspect of the present application, an electronic device is provided, including a processor and the foregoing non-transitory computer-readable storage medium.

[0009] The present application has at least the following beneficial effects: The data table classification method provided by this application first obtains each operation statement of the data table to be classified within the target time window and inputs it into the text encoding layer of the data table classification model to obtain the first feature vector. Here, obtain the operation characteristics of the data table to be classified in chronological order within the target time window. Secondly, obtain each key foreign key information of the data table to be classified and, according to the graph feature extraction layer of the data table classification model, obtain the second feature vector. Here, obtain the reference relationship characteristics between the data table to be classified and other key data tables with clear classifications. After that, obtain the content feature information of the data table to be classified and, according to the graph feature extraction layer of the data table classification model, obtain the third feature vector. Here, the third feature vector characterizes the features of the internal data of the data table to be classified. Then, obtain the first weight corresponding to the first feature vector and the second weight corresponding to the second feature vector according to the third feature vector. Here, since the first feature vector and the second feature vector are the influences of different external factors on the data table to be classified, and the third feature vector analyzes the internal data features of the data table to be classified. If the third feature vector characterizes the data as more chaotic and less standard, then at this time, the first feature vector is more important for determining the final category, and the corresponding first weight is larger. On the contrary, if the third feature vector characterizes the data as relatively standard, then the possibility of classifying this data as non-raw data is greater. At this time, the second feature vector may be more important for determining the final category, and the corresponding second weight is larger. Thus, after assigning weights to the first feature vector and the second feature vector, obtain the final target feature vector and obtain the final classification result. This application not only considers the connection between the data table to be classified and external factors, including data reference relationship characteristics and operation characteristics of the data table to be classified, but also assigns corresponding weights to the data reference relationship characteristics and operation characteristics of the data table to be classified based on the characteristics of the internal data of the data table to be classified, so that data tables with different internal data characteristics construct different target feature vectors; it has stronger migration and adaptability, and the obtained classification result is more accurate and objective. Description of the Drawings

[0010] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0011] Figure 1 Flowchart of the data table classification method provided by the embodiment of this application; Figure 2 Structural block diagram of the data table classification device provided by the embodiment of this application. Detailed Embodiments

[0012] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0013] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0014] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present application, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. Additionally, this apparatus and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.

[0015] Please refer to Figure 1 As shown, an embodiment of the present application provides a method for classifying data tables, and the method includes: Step S100, obtaining each operation statement of the data table to be classified within a target time window, and inputting it into the text encoding layer of the data table classification model to obtain a first feature vector YT; where YT represents the operation characteristics of the data table to be classified in chronological order within the target time window; the end time of the target time window is the current time.

[0016] Specifically, the target time window can be of any time length. Preferably, the target time window can be from the start of the incoming storage of the data table to be classified to the current time. The longer the target time window, the richer the operation features that can be obtained. Here, the operation features can be extracted from SQL statements, where the operations can be querying, processing, calculating, etc. Since the data in different data tables may be different, the operations performed on different data tables may also be different. Sort the operation features within the target time window in chronological order to obtain the first feature vector with temporal characteristics.

[0017] The text encoding layer of the data table classification model is used to encode the obtained several operation statements and convert the operation statements into the final first feature vector with temporal characteristics.

[0018] Step S200: Obtain each key foreign key information of the data table to be classified, and according to the graph feature extraction layer of the data table classification model, obtain the second feature vector ET; where ET represents the data reference feature of the data table to be classified; the key foreign key information is the foreign key information between the data table to be classified and the key data table; the key data table and the data table to be classified are in the same database; and the key data table can determine the data table category according to the preset classification rules; the data table to be classified cannot determine the data table category according to the preset classification rules.

[0019] Specifically, the foreign key is the matching relationship between a certain field value in a data table and a certain field value in another data table. In this embodiment, each key foreign key information of the data table to be classified represents the information that the data table to be classified references other key data tables or the information that the data table to be classified is referenced by other key data tables. Here, the key data table is a data table that has determined the data table category according to the preset classification rules and is in the same database as the data table to be classified. It should be noted that the preset classification rules can be a method of classifying standard-named data tables based on standard information such as standard data table names and / or storage paths. That is, the preset classification rules cannot classify this type of non-standard-named data tables to be classified.

[0020] Step S300: Obtain the content feature information of the data table to be classified, and according to the graph feature extraction layer of the data table classification model, obtain the third feature vector; where the third feature vector represents the data content feature of the data table to be classified.

[0021] Specifically, the content feature information can include feature information such as the unique value ratio of fields, the fluctuation of the corresponding field values within each field, the total number of fields, and the correlation matrix between fields. Extract the features of the correlation matrix according to the graph encoding layer and jointly generate the third feature vector with other feature information.

[0022] Step S400: Input the third feature vector into the weight assignment layer of the data table classification model to obtain the first weight corresponding to YT and the second weight corresponding to ET.

[0023] Specifically, the weight assignment layer of the data table classification model is used to obtain the first weight and the second weight according to the input third feature vector. In one embodiment, if the third feature vector indicates that the data is more chaotic and less standard, then at this time, the greater the possibility that this data is classified as raw data, and the first feature vector is relatively more important for determining the final category, so the corresponding first weight is greater; conversely, if the third feature vector indicates that the data is relatively standard, then the greater the possibility that this data is classified as non-raw data. At this time, the second feature vector is relatively more important for determining the final category, so the corresponding second weight is greater.

[0024] Step S500: Obtain the target feature vector according to YT, the first weight, ET, and the second weight.

[0025] Step S600: Input the target feature vector into the classification layer of the data table classification model to obtain the classification result corresponding to the data table to be classified.

[0026] Thus, after assigning weights to the first feature vector and the second feature vector, the final target feature vector is obtained, and the final classification result is acquired.

[0027] Here, the classification result can be any one of raw data, data after cleaning, data after aggregation and summary, and application metric data.

[0028] Among them, the original data is the ODS data, which includes the fields directly synchronized by the business system (such as user_id_raw, order_date_source); its corresponding typical SQL operations include full - volume insertion, simple cleaning, and simple data addition, etc., which have characteristics such as low complexity and no cross - table association; the data after cleaning is the DWD data (such as user_id_cleaned, order_date_utc); its corresponding typical SQL operations include field type conversion, data desensitization, and lightweight deduplication, etc.; it has characteristics such as medium complexity, multi - table association, clear and standardized steps, etc.; the data after aggregation and summary is the DWS data (such as daily_sales_total, monthly_active_users); its corresponding typical SQL operations include aggregation, window functions, cross - fact - table association to calculate metrics, etc.; it has characteristics such as pre - calculation of key metrics, time - dimension aggregation, and serving multiple business lines, etc.; the application - metric data is the ADS data (such as retention_rate_7d, conversion_funnel), and its corresponding typical SQL operations include complex business formulas, data binning, and directly exporting the results to reports or API interfaces, etc.; it has characteristics such as business - term naming, very few underlying associations, and final metric output, etc.

[0029] It can be understood that there is a progressive relationship among the above - mentioned data - table categories, which are obtained through step - by - step processing. That is, the data after cleaning can be obtained by processing the original data, the data after aggregation and summary can be obtained by further processing, and the application - metric data can be obtained by further processing. However, the above four types of data can exist in the same database simultaneously.

[0030] This application not only considers the relationship between the data table to be classified and external factors, including the data - reference relationship characteristics and the operation characteristics of the data table to be classified, but also assigns corresponding weights to the data - reference relationship characteristics and the operation characteristics of the data table to be classified based on the characteristics of the internal data of the data table to be classified, so that data tables with different internal data characteristics construct different target feature vectors; it has stronger migration and adaptability, and the obtained classification results are more accurate and objective.

[0031] In an exemplary embodiment of this application, step S200 includes: Step S210, obtaining each operation statement of the data table to be classified within the target time window to obtain an operation - statement list; among them, each operation statement has a corresponding execution time.

[0032] Step S220, inputting the operation - statement list into the text - encoding layer of the data - table classification model to extract the key operation characteristics in each operation statement; among them, the key operation characteristics include: operation - keyword characteristics, field - name characteristics, and conditional - expression characteristics.

[0033] Specifically, the operation keyword features include SELECT, INSERT, UPDATE, DELETE, etc. These keywords are used as features. A keyword dictionary can be established, with each keyword corresponding to a unique index. The field name features include table names and field names, etc. Similarly, a dictionary is established for them and a unique index is assigned. The conditional expression features include the conditions in the WHERE clause, etc. The type of condition (such as equal to, greater than, less than, etc.) and the columns involved can be used as features.

[0034] Step S230, construct the first feature vector YT = (YT1, YT2,..., YT i ,..., YT n ); i = 1, 2,..., n; where n is the number of operation statements; YT i is the list of feature values corresponding to the operation statement ranked at the i-th position in chronological order of execution time; YT i = (YT i,1 , YT i,2 ,..., YT i,j ,..., YT i,m , G i ); j = 1, 2,..., m; m is the total number of key operation features; YT i,j is the feature value of the j-th key operation feature of the operation statement ranked at the i-th position in chronological order of execution time; G i is the normalized time corresponding to the operation statement ranked at the i-th position in chronological order of execution time; G i conforms to the following characteristics: G i = (T i - T min ) / (T min - T max ); where T i is the execution time of the operation statement ranked at the i-th position in chronological order of execution time; T min is the earliest execution time among the execution times corresponding to all operation statements in the operation statement list; T max is the latest execution time among the execution times corresponding to all operation statements in the operation statement list.

[0035] Specifically, a feature bit is reserved for the normalized time to represent the execution order of the SQL statement, that is, G i .

[0036] For the rest, assign a feature bit to each keyword. If an operation statement contains a certain keyword, the value of the corresponding feature bit is set to 1; otherwise, it is set to 0. Similarly, assign feature bits to each table name and field name according to all table names and field names. If an operation statement involves a certain table name or column name, the corresponding feature bit is set to 1; otherwise, it is set to 0. The same applies to conditional expressions. If an operation statement involves a certain conditional expression, the corresponding feature bit is set to 1; otherwise, it is set to 0.

[0037] Sort the feature vectors of all SQL statements according to the normalized time values to ensure that they are arranged in the order of execution time. Concatenate the sorted feature vectors in sequence to form a long feature vector, which represents the timing information of all operation statements. The timing encoding can clearly present the execution order of each operation statement. By analyzing these features arranged in time order, it is possible to clearly understand the entire data processing flow, know which operations are executed first and which are executed later, which helps to understand how the data is gradually processed and transformed. It can also help identify the dependency relationships between operation statements. If the execution result of one statement is used as the input of another statement, then this sequence and dependency relationship will be reflected in the timing features.

[0038] In an exemplary embodiment of the present application, step S200 includes: Step S210, obtain each data table in the database where the data table to be classified is located except the data table to be classified, to obtain a data table identifier list SB = (SB1, SB2,..., SB x ,..., SB y ); x = 1, 2,..., y; where y is the number of data tables in the database where the data table to be classified is located except the data table to be classified.

[0039] Step S220, classify each data table in SB according to a preset classification rule to obtain a classified data table identifier list SBY = (SBY1, SBY2,..., SBY a ,..., SBY b ); a = 1, 2,..., b; where b is the number of classified data tables; SBY a is the a-th classified data table; b ≤ y.

[0040] Specifically, the preset classification rule can be a method for classifying data tables with standard names based on standard information such as standard data table names and / or storage paths. That is, the preset classification rule cannot classify data tables with non-standard names in the category of data tables to be classified. That is, the classified data tables are also data tables in the database where the data table to be classified is located that can determine the data table type using the preset classification rule.

[0041] Step S230: Traverse SBY according to each key foreign key information of the data table to be classified, so as to obtain the key data table identifier list GSBY = (GSBY1, GSBY2,..., GSBY c ,..., GSBY d ); c = 1, 2,..., d; where d is the number of key data tables; GSBY c is the c-th key data table; d ≤ b; the key data table is a classified data table that has at least one reference relationship with the data table to be classified.

[0042] Specifically, obtain d key data tables in SBY that have a reference or be-referenced relationship with the data table to be classified.

[0043] Step S240: According to each key foreign key information of the data table to be classified and GSBY, obtain a directed graph; where the directed graph includes nodes and directed edges; the nodes include the data table node to be classified and the key data table nodes; the directed edges represent the reference relationships between the data table node to be classified and each key data table node; the directed edges are one-way arrows; the arrows point to the referencing party; each of the nodes also includes the type information of the corresponding data table to be classified or key data table.

[0044] Specifically, the directed graph takes the data table node to be classified as the center and radiates outwards. Each directed edge represents the reference or be-referenced relationship between each connected key data table node and the data table to be classified. Further, the directed edge is a one-way arrow; the arrow points to the referencing party; as an example: the arrow points to the data table to be classified and away from key data table A. The reference relationship between the data table to be classified and A is that a certain field value in the data table to be classified references a certain field value in key data table A.

[0045] Step S250: Input the directed graph into the graph feature extraction layer of the data table classification model to obtain the second feature vector ET.

[0046] Specifically, take the directed graph as an image containing a large amount of information, and continue to extract graph features from it to obtain the second feature vector ET. As an example: the graph feature can be graph density, etc.

[0047] The second feature vector obtained in this embodiment makes full use of the classified data tables that have a foreign key relationship with the data table to be classified to construct a directed graph, and can obtain which category of data tables the data table to be classified has a reference relationship with, as well as the direction and closeness degree of the reference relationship. It has an important reference role in finally determining the category to which the data table to be classified belongs.

[0048] In an exemplary embodiment of the present application, the data table to be classified contains several fields; the fields are divided into numerical fields and non-numerical fields; each field has several field values. Step S300 includes: Step S310, obtaining the unique value ratio corresponding to each field in the data table to be classified, so as to obtain the average unique value ratio WP corresponding to the data table to be classified.

[0049] Specifically, the unique value ratio of a field refers to the non-repetition ratio of all field values corresponding to the field in the data table; as an example: for field A, the corresponding field values in a certain data table are 100, and these 100 field values are all different, then the unique value ratio of field A is 100%. And the average unique value ratio corresponding to the data table to be classified is the average of the unique value ratios of all fields included in the data table to be classified. The larger WP is, the higher the degree of data dispersion is, each value is relatively unique, and there are fewer duplicate data. On the contrary, if WP is smaller, it means that the degree of data dispersion is low and there are more duplicate data.

[0050] Step S320, obtaining the numerical fluctuation ratio ZD = SC / f of the data table to be classified; where SC is the number of fluctuating numerical fields in the data table to be classified; f is the total number of numerical fields in the data table to be classified; the fluctuating numerical field is a field whose corresponding standard deviation is greater than the preset standard deviation threshold.

[0051] Specifically, the fluctuating numerical field can also be a field whose corresponding variance is greater than the preset variance threshold. The larger ZD is, the greater the difference between the individual data values in the data table to be classified is, the more widely distributed they are, and there are more fields that may have some extreme values. On the contrary, the smaller ZD is, the greater the difference between the individual data values in the data table to be classified is, the more widely distributed they are, and there are fewer fields that may have some extreme values. And the later in the data processing, after processing such as standardization, the standard deviation of the field is smaller. Therefore, in one embodiment, the larger ZD is, the greater the possibility that the data table is in the initial stage, that is, the original data table. On the contrary, the possibility is smaller.

[0052] Step S330, obtaining the correlation matrix XG corresponding to the data table to be classified; where XG meets the following conditions: ; where e = 1, 2,..., f; XG e,f is the correlation coefficient between the e-th numerical field included in the data table to be classified and the f-th numerical field included in the data table to be classified.

[0053] Here, XG e,f = (cov(XG e , XG f )) / (σ XGe ×σXGf ); where, cov(XG e , XG f ) is the covariance between the e-th numerical field and the f-th numerical field included in the data table to be classified; σ XGe is the standard deviation of the e-th numerical field included in the data table to be classified; σ XGf is the standard deviation of the f-th numerical field included in the data table to be classified.

[0054] In one embodiment, the larger XG is, the higher the correlation between fields is, that is, the greater the possibility that the data table to be classified is the original data. This is because in the original data, due to no duplicate removal, cleaning, standardization, aggregation and other processing, there may be highly correlated fields, etc.; conversely, if XG is smaller, it indicates that the correlation between fields is lower, that is, the possibility that the data table to be classified is the original data is smaller. And the possibility of being in the later stage of processing is greater.

[0055] Step S340, input XG into the graph feature extraction layer of the data table classification model to obtain the image feature TXG corresponding to XG.

[0056] Step S350, according to WP, ZD and TXG, obtain the third feature vector ST = (WP, ZD, z, TXG); where, z is the total number of fields in the data table to be classified.

[0057] Specifically, finally, the above features and the total number feature of fields in the data table to be classified are jointly used as the third feature vector. Here, the third feature vector can represent the data dispersion feature, field feature, and association feature between fields in the data table to be classified.

[0058] In an exemplary embodiment of the present application, the target feature vector MT meets the following conditions: MT = (α × YT, β × ET); where, α is the first weight; β is the second weight.

[0059] Specifically, the obtained first weight and second weight are used to adjust the proportion of the first feature vector and the second feature vector in the target feature vector, so that the obtained target feature vector can better reflect the characteristics of the data table to be classified. The target feature vectors corresponding to different characteristic data tables to be classified are different, so that the migration and adaptability of the finally obtained classification result are stronger and more accurate.

[0060] It should be noted that the above steps are all implemented by each module in the data table classification model.

[0061] Please refer to Figure 2As shown in the figure, an embodiment of the present application provides a data table classification device 100, and the device includes: A first vector acquisition unit 110, configured to acquire each operation statement of the data table to be classified within a target time window, and input it into the text encoding layer of the data table classification model to obtain a first feature vector YT; where YT represents the operation features of the data table to be classified in chronological order within the target time window; the end time of the target time window is the current time.

[0062] A second vector acquisition unit 120, configured to acquire each key foreign key information of the data table to be classified, and obtain a second feature vector ET according to the graph feature extraction layer of the data table classification model; where ET represents the data reference features of the data table to be classified; the key foreign key information is the foreign key information between the data table to be classified and the key data table; the key data table and the data table to be classified are within the same database; and the key data table can determine the data table category according to the preset classification rules; the data table to be classified cannot determine the data table category according to the preset classification rules.

[0063] A third vector acquisition unit 130, configured to acquire the content feature information of the data table to be classified, and obtain a third feature vector according to the graph feature extraction layer of the data table classification model; where the third feature vector represents the data content features of the data table to be classified.

[0064] A weight acquisition unit 140, configured to input the third feature vector into the weight assignment layer of the data table classification model to obtain a first weight corresponding to YT and a second weight corresponding to ET.

[0065] A target vector acquisition unit 150, configured to obtain a target feature vector according to YT, the first weight, ET, and the second weight.

[0066] A classification unit 160, configured to input the target feature vector into the classification layer of the data table classification model to obtain a classification result corresponding to the data table to be classified.

[0067] An embodiment of the present application further provides a computer program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps in the methods according to various exemplary embodiments of the present application described above in this specification.

[0068] In addition, although the steps of the methods in the present application are described in a specific order in the drawings, this does not require or imply that these steps must be executed in this specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0069] From the descriptions of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0070] In an exemplary embodiment of the present application, an electronic device capable of implementing the above method is further provided.

[0071] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.

[0072] The electronic device according to this embodiment of the present application. The electronic device is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present application.

[0073] The electronic device is presented in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: at least one of the above-mentioned processors, at least one of the above-mentioned memories, and a bus connecting different system components (including the memory and the processor).

[0074] Among them, the memory stores program codes, and the program codes can be executed by the processor, so that the processor executes the steps according to various exemplary embodiments of the present application described in the above "Exemplary Method" section of this specification.

[0075] The memory may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) and / or a cache memory, and may further include a read-only memory (ROM).

[0076] The memory may further include a program / utility having a set (at least one) of program modules, and such program modules include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.

[0077] The bus can represent one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of the various bus architectures.

[0078] The electronic device can also communicate with one or more external devices (such as a keyboard, a pointing device, a Bluetooth device, etc.), can also communicate with one or more devices that enable a user to interact with the electronic device, and / or can communicate with any device that enables the electronic device to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface. Moreover, the electronic device can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter. As shown in the figure, the network adapter communicates with other modules of the electronic device through the bus. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0079] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0080] In an exemplary embodiment of the present application, a computer-readable storage medium is also provided, on which a program product capable of implementing the above method of this specification is stored. In some possible implementation manners, various aspects of the present application can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present application described in the above "Exemplary Method" section of this specification.

[0081] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0082] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0083] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0084] The program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0085] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, and are not for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0086] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0087] The above is only the specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for classifying data tables, characterized in that, The method includes: Obtain each operation statement of the data table to be classified within the target time window, and input it into the text encoding layer of the data table classification model to obtain a first feature vector YT; where YT represents the operation characteristics of the data table to be classified in chronological order within the target time window; the end time of the target time window is the current time; Obtain each key foreign key information of the data table to be classified, and according to the graph feature extraction layer of the data table classification model, obtain a second feature vector ET; where ET represents the data reference characteristics of the data table to be classified; the key foreign key information is the foreign key information between the data table to be classified and the key data table; the key data table and the data table to be classified are within the same database; and the key data table can determine the data table category according to the preset classification rules; the data table to be classified cannot determine the data table category according to the preset classification rules; Obtain the content feature information of the data table to be classified, and according to the graph feature extraction layer of the data table classification model, obtain a third feature vector; where the third feature vector represents the data content characteristics of the data table to be classified; Input the third feature vector into the weight assignment layer of the data table classification model to obtain a first weight corresponding to YT and a second weight corresponding to ET; Obtain the target feature vector according to YT, the first weight, ET and the second weight; Input the target feature vector into the classification layer of the data table classification model to obtain the classification result corresponding to the data table to be classified.

2. The data table classification method according to claim 1, wherein Obtain each operation statement of the data table to be classified within the target time window, and input it into the text encoding layer of the data table classification model to obtain a first feature vector, including: Obtain each operation statement of the data table to be classified within the target time window to obtain an operation statement list; where each operation statement has a corresponding execution time; Input the operation statement list into the text encoding layer of the data table classification model, and extract the key operation features in each operation statement; where the key operation features include: operation keyword features, field name features and conditional expression features; Construct the first feature vector YT=(YT1, YT2,..., YT i ,..., YT n ); i = 1, 2,..., n; where n is the number of operation statements; YT i is the list of eigenvalue corresponding to the operation statement ranked at the i-th position in the chronological order of execution time; YT i =(YT i,1 , YT i,2 ,..., YT i,j ,..., YT i,m , G i ); j = 1, 2,..., m; m is the total number of key operation features; YT i,j is the eigenvalue of the j-th key operation feature of the operation statement ranked at the i-th position in the chronological order of execution time; G i is the normalized time corresponding to the operation statement ranked at the i-th position in the chronological order of execution time; G i conforms to the following characteristics: G i =(T i -T min ) / (T min -T max ); where T i is the execution time of the operation statement ranked at the i-th position in the chronological order of execution time; T min is the earliest execution time among the execution times corresponding to all operation statements in the operation statement list; T max is the latest execution time among the execution times corresponding to all operation statements in the operation statement list.

3. The data table classification method according to claim 1, characterized in that Obtain each key foreign key information of the data table to be classified, and according to the graph feature extraction layer of the data table classification model, obtain a second feature vector, including: Obtain each data table in the database where the data table to be classified is located, except for the data table to be classified, to obtain a data table identifier list SB = (SB1, SB2,..., SB x ,..., SB y ); x = 1, 2,..., y; where y is the number of data tables in the database where the data table to be classified is located, except for the data table to be classified; Classify each data table in SB according to the preset classification rules to obtain a list of classified data table identifiers SBY = (SBY1, SBY2,..., SBY a ,..., SBY b ); a = 1, 2,..., b; where b is the number of classified data tables; SBY a is the a-th classified data table; b ≤ y; Traverse SBY according to each key foreign key information of the data table to be classified, so as to obtain the key data table identification list GSBY = (GSBY1, GSBY2,..., GSBY c ,..., GSBY d ); c = 1, 2,..., d; where d is the number of key data tables; GSBY c is the c-th key data table; d ≤ b; the key data table is a classified data table that has at least one reference relationship with the data table to be classified; According to each key foreign key information of the data table to be classified and GSBY, obtain a directed graph; where the directed graph includes nodes and directed edges; the nodes include the data table node to be classified and the key data table node; the directed edge represents the reference relationship between the data table node to be classified and each key data table node; the directed edge is a one-way arrow; the arrow points to the referencing party; each node also includes the type information of the corresponding data table to be classified or key data table; Input the directed graph into the graph feature extraction layer of the data table classification model to obtain a second feature vector ET.

4. The data table classification method according to claim 1, characterized in that The data table to be classified contains several fields; the fields are divided into numerical fields and non-numerical fields; each field has several field values.

5. The data table classification method according to claim 4, wherein Obtain the content feature information of the data table to be classified, and according to the graph feature extraction layer of the data table classification model, obtain a third feature vector, including: Obtain the unique value ratio corresponding to each field in the data table to be classified to obtain the average unique value ratio WP corresponding to the data table to be classified; Obtain the numerical fluctuation ratio ZD = SC / f corresponding to the data table to be classified; where SC is the number of fluctuating numerical fields in the data table to be classified; f is the total number of numerical fields in the data table to be classified; the fluctuating numerical field is a field whose corresponding standard deviation is greater than the preset standard deviation threshold; Obtain the correlation matrix XG corresponding to the data table to be classified; wherein, XG meets the following conditions: ; where e = 1, 2, ..., f; XG e,f is the correlation coefficient between the e-th numerical field and the f-th numerical field included in the data table to be classified; Input XG into the graph feature extraction layer of the data table classification model to obtain the image feature TXG corresponding to XG; According to WP, ZD, and TXG, obtain the third feature vector ST = (WP, ZD, z, TXG); where z is the total number of fields in the data table to be classified.

6. The data table classification method according to claim 1, wherein The target feature vector MT meets the following conditions: MT = (α × YT, β × ET); Where α is the first weight; β is the second weight.

7. The data table classification method according to claim 1, wherein The classification result is any one of the original data, the data after cleaning, the data after aggregation and summarization, and the application metric data.

8. A data table classification device, characterized in that, The device includes: The first vector acquisition unit is used to acquire each operation statement of the data table to be classified within the target time window and input it into the text encoding layer of the data table classification model to obtain the first feature vector YT; where YT represents the operation characteristics of the data table to be classified in chronological order within the target time window; the end time of the target time window is the current time; The second vector acquisition unit is used to acquire each key foreign key information of the data table to be classified and obtain the second feature vector ET according to the graph feature extraction layer of the data table classification model; where ET represents the data reference characteristics of the data table to be classified; the key foreign key information is the foreign key information between the data table to be classified and the key data table; the key data table and the data table to be classified are in the same database; and the key data table can determine the data table category according to the preset classification rules; the data table to be classified cannot determine the data table category according to the preset classification rules; The third vector acquisition unit is used to acquire the content feature information of the data table to be classified and obtain the third feature vector according to the graph feature extraction layer of the data table classification model; where the third feature vector represents the data content characteristics of the data table to be classified; The weight acquisition unit is used to input the third feature vector into the weight assignment layer of the data table classification model to obtain the first weight corresponding to YT and the second weight corresponding to ET; The target vector acquisition unit is used to obtain the target feature vector according to YT, the first weight, ET, and the second weight; The classification unit is used to input the target feature vector into the classification layer of the data table classification model to obtain the classification result corresponding to the data table to be classified.

9. A non-transitory computer-readable storage medium, characterized in that, At least one instruction or at least one program is stored in the storage medium, and the at least one instruction or the at least one program is loaded and executed by a processor to implement the method described in any one of claims 1-7.

10. An electronic device, characterized in that, It includes a processor and the non-transitory computer-readable storage medium described in claim 9.

Citation Information

Patent Citations

  • Method and system for classifying data tables, terminal and storage medium

    CN109800422A

  • Table classification method and device, equipment, and storage medium

    CN112989050A

  • Data table classification method and device, equipment and storage medium

    CN115599975A

  • Database table classification treatment method and system based on multi-modal fusion

    CN119537489A

  • Information classification extraction method, apparatus, computer device and storage medium

    WO2021042503A1