A data processing method, device, and equipment

By automatically obtaining the attribute information of the data source and associating it with the data table, the cumbersome problem of users manually inputting meta information in cloud relational databases is solved, and the user experience and meta information construction efficiency is improved.

CN111723161BActive Publication Date: 2025-06-20ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910213125.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-03-20
Publication Date
2025-06-20
Estimated Expiration
2039-03-20

AI Technical Summary

Technical Problem

In a cloud relational database, users need to manually give the meta information of the data to facilitate the mapping of tables and data, resulting in large workloads, easy errors and poor user experience.

Method used

By obtaining the location information in the data processing request, the attribute information is automatically obtained from the data set of the data source, the data table is created, and the meta information of the data set is associated with the data table to achieve automatic mapping.

Benefits of technology

It reduces the workload of users to manually input meta information, improves the user experience, and greatly improves the efficiency of meta information construction, and improves the overall usage efficiency and experience of the data lake analysis system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111723161B_ABST
    Figure CN111723161B_ABST
Patent Text Reader

Abstract

The present application provides a data processing method, apparatus, and device. The method includes: obtaining a data processing request, where the data processing request includes location information of a data source; obtaining attribute information from a data set of the data source according to the location information; the data source includes a plurality of data sets, and the data set includes attribute information of the data set; creating a data table according to the attribute information, the data table corresponding to at least one data set, and associating meta-information corresponding to the at least one data set with the data table; and performing data processing by using the data table and the meta-information associated with the data table. Through the technical solution of the present application, meta-information and a data table can be automatically associated, thereby reducing the workload of users and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technologies, and in particular, to a data processing method, apparatus, and device. Background Art

[0002] Data Lake Analytics is used to provide serverless query analysis services for users, capable of analyzing and querying massive data in any dimension, and supporting functions such as high concurrency, low latency (millisecond-level response), real-time online analysis, and massive data query.

[0003] In a traditional relational database, if a user needs to use the database for query and analysis, the following operations are performed: creating a database; creating a Table (data table), where a Table refers to a set that associates and maintains all homogeneous records; importing data into the Table; and performing query and analysis based on the data in the Table. In a data lake analysis system, which provides a cloud relational database, different from the traditional relational database, if a user needs to use the database for query and analysis, the following operations are performed: creating a Table, mapping the Table to a partial data set of the currently affiliated data source; and performing query and analysis based on the Table.

[0004] In summary, it can be seen that in a traditional relational database, a Table is created first, and then data is imported into the Table; in a cloud relational database, a Table is created based on existing data, but there is no need to import data into the Table, and only the Table needs to be mapped to the data.

[0005] Obviously, in a cloud relational database, one of the core tasks is how to implement the mapping. In the traditional method, to implement the mapping, the following method can be adopted: the user specifies the mapping relationship between the Table and the data, that is, the user gives the meta-information of the data and binds the meta-information to the Table. However, when the user gives the meta-information, the user's workload is large and it is easy to make mistakes, resulting in a poor user experience. Summary of the Invention

[0006] This application provides a data processing method, which includes:

[0007] Obtaining a data processing request, where the data processing request includes the location information of the data source;

[0008] Obtaining attribute information from the data set of the data source according to the location information; where the data source includes multiple data sets, and the data set includes the attribute information of the data set;

[0009] Create a data table according to the attribute information, where the data table corresponds to at least one data set, and associate the meta information corresponding to the at least one data set with the data table;

[0010] Perform data processing using the data table and the meta information associated with the data table.

[0011] This application provides a data processing method, which is applied to a data lake analysis platform. The data lake analysis platform is used to provide serverless data processing services for users. The method includes:

[0012] Obtain a data processing request, where the data processing request includes the location information of the data source;

[0013] Obtain attribute information from the data sets of the data source according to the location information; where the data source includes multiple data sets, and the data set includes the attribute information of the data set;

[0014] Create a data table according to the attribute information, where the data table corresponds to at least one data set, and associate the meta information corresponding to the at least one data set with the data table;

[0015] Perform data processing using the data table and the meta information associated with the data table;

[0016] Wherein, the data source includes a cloud database provided by the data lake analysis platform.

[0017] This application provides a data processing method, and the method includes:

[0018] Obtain a data processing request, where the data processing request includes the location information of the data source;

[0019] Obtain attribute information from the data sets of the data source according to the location information;

[0020] Create a data table according to the attribute information, where the data table corresponds to at least one data set of the data source, and associate the meta information corresponding to the at least one data set with the data table;

[0021] Wherein, the association relationship between the data table and the meta information is used for data processing.

[0022] This application provides a data processing method, and the method includes:

[0023] Obtain a data processing request, where the data processing request includes the location information of the data source;

[0024] Obtain attribute information from the datasets of the data source according to the location information; wherein, the data source includes multiple datasets, and the dataset includes the attribute information of the dataset;

[0025] Cluster the multiple datasets according to the attribute information respectively corresponding to the multiple datasets to obtain a clustering set, wherein the clustering set includes at least one dataset;

[0026] Create a data table for the clustering set, and the data table corresponds to the at least one dataset;

[0027] Associate the meta-information corresponding to the at least one dataset with the data table;

[0028] Perform data processing using the data table and the meta-information associated with the data table.

[0029] This application provides a data processing method, and the method includes:

[0030] Obtain a data query request, and the data query request includes data table information;

[0031] Obtain the data table corresponding to the data table information and the meta-information associated with the data table; wherein, the data table is created according to the attribute information of the datasets in the data source, and the meta-information associated with the data table includes the meta-information corresponding to at least one dataset of the data source;

[0032] Process the query request using the data table and the meta-information associated with the data table.

[0033] This application provides a data processing device, and the device includes:

[0034] An obtaining module, configured to obtain a data processing request, and the data processing request includes the location information of the data source; obtain attribute information from the datasets of the data source according to the location information; wherein, the data source includes multiple datasets, and the dataset includes the attribute information of the dataset;

[0035] An association module, configured to create a data table according to the attribute information, the data table corresponds to at least one dataset, and associate the meta-information corresponding to the at least one dataset with the data table;

[0036] A processing module, configured to perform data processing using the data table and the meta-information associated with the data table.

[0037] This application provides a data processing device, including:

[0038] A processor and a machine-readable storage medium, where several computer instructions are stored on the machine-readable storage medium, and when the processor executes the computer instructions, the following processing is performed:

[0039] Obtain a data processing request, where the data processing request includes location information of a data source;

[0040] Obtain attribute information from the data set of the data source according to the location information; where the data source includes multiple data sets, and the data set includes the attribute information of the data set;

[0041] Create a data table according to the attribute information, where the data table corresponds to at least one data set, and associate the meta-information corresponding to the at least one data set with the data table;

[0042] Perform data processing using the data table and the meta-information associated with the data table.

[0043] Based on the above technical solution, in the embodiments of the present application, attribute information can be obtained from the data set of the data source, a data table can be created according to the attribute information, and the meta-information corresponding to the data set can be associated with the data table. That is to say, the meta-information and the data table can be automatically associated, without the user giving the meta-information and associating the meta-information with the data table, thereby reducing the workload of the user, improving the user experience, and greatly improving the construction efficiency of the meta-information, and enhancing the overall usage efficiency and experience of the data lake analysis system. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings in the embodiments of the present application.

[0045] Figure 1 It is a schematic flowchart of a data processing method in an embodiment of the present application;

[0046] Figure 2 It is a schematic structural diagram of a data lake analysis system in an embodiment of the present application;

[0047] Figure 3 It is a schematic diagram of obtaining data source information in an embodiment of the present application;

[0048] Figure 4 It is a schematic flowchart of a data processing method in an embodiment of the present application;

[0049] Figure 5It is a schematic structural diagram of a data processing device in an embodiment of the present application;

[0050] Figure 6 It is a schematic structural diagram of a data processing device in an embodiment of the present application. Specific embodiments

[0051] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and do not limit the present application. The singular forms "a", "the" and "said" used in the present application and the claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to any or all possible combinations of one or more of the associated listed items.

[0052] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, in addition, the word "if" used may be interpreted as "when" or "while" or "in response to determining".

[0053] In an embodiment of the present application, a data processing method is proposed, which can be applied to any device, such as any device of a data lake analysis system. Refer to Figure 1 As shown, it is a flowchart of the method, and the method may include:

[0054] Step 101, obtain a data processing request, and the data processing request includes location information of a data source.

[0055] Step 102, obtain attribute information from a data set of the data source according to the location information; wherein, the data source includes multiple data sets, and each data set includes attribute information of the data set.

[0056] Among them, the data set in this embodiment may be a sub-file of the data source or other types of data sets, and there is no limitation on this, as long as the data set includes multiple data of the data source.

[0057] In one example, obtaining attribute information from a data set of the data source according to the location information may include: determining whether the data table discovery function is enabled for the data processing request; if so, obtain attribute information from the data set of the data source according to the location information. If not, it is not necessary to obtain attribute information from the data set of the data source according to the location information, but a traditional process is used for processing.

[0058] Among them, determining whether the data processing request enables the data table discovery function may include: if the data processing request further includes automatic discovery indication information, it may be determined whether the data processing request enables the data table discovery function according to the automatic discovery indication information. For example, the automatic discovery indication information is used to indicate whether to enable the data table discovery function. If the automatic discovery indication information is used to indicate enabling the data table discovery function, it may be determined that the data processing request enables the data table discovery function according to the automatic discovery indication information. If the automatic discovery indication information is used to indicate not enabling the data table discovery function, it may be determined according to the automatic discovery indication information that there is no need to enable the data table discovery function for the data processing request.

[0059] Step 103: Create a data table according to the attribute information. The data table corresponds to at least one data set, and associate the meta-information corresponding to the at least one data set with the data table.

[0060] In one example, creating a data table according to the attribute information, where the data table corresponds to at least one data set, may include: clustering multiple data sets according to the attribute information respectively corresponding to the multiple data sets to obtain a clustering set, where each clustering set may include at least one data set. For each clustering set, create a data table for the clustering set, and the data table corresponds to the at least one data set.

[0061] Further, clustering multiple data sets according to the attribute information respectively corresponding to the multiple data sets to obtain a clustering set may include, but is not limited to: obtaining clustering indication information, where the clustering indication information may be used to indicate clustering sub-attributes; then, based on the attribute information respectively corresponding to the multiple data sets, the clustering sub-attributes respectively corresponding to the multiple data sets may be determined according to the clustering indication information, and the multiple data sets may be clustered according to the clustering sub-attributes respectively corresponding to the multiple data sets to obtain a clustering set.

[0062] Among them, clustering multiple data sets according to the clustering sub-attributes respectively corresponding to the multiple data sets to obtain a clustering set may include, but is not limited to: based on the clustering sub-attributes respectively corresponding to the multiple data sets, clustering the data sets with the same clustering sub-attributes into the same clustering set, and clustering the data sets with different clustering sub-attributes into different clustering sets. That is to say, the data sets with the same clustering sub-attributes may correspond to the same clustering set, and the data sets with different clustering sub-attributes may correspond to different clustering sets.

[0063] Among them, obtaining clustering indication information may include, but is not limited to: if the data processing request further includes clustering indication information, the clustering indication information may be obtained from the data processing request; or, obtaining pre-configured clustering indication information, such as obtaining pre-configured clustering indication information from this device.

[0064] In one example, according to the attribute information corresponding to multiple data sets, clustering the multiple data sets to obtain a clustering set may include: if the data processing request further includes filtering indication information, filtering the multiple data sets according to the filtering indication information to obtain a target data set; then, clustering the target data set based on the attribute information corresponding to the target data set to obtain a clustering set.

[0065] In one example, after creating a data table according to the attribute information, it may further include: if the data processing request further includes naming indication information, naming the data table according to the naming indication information.

[0066] In one example, associating the meta information corresponding to the at least one data set with the data table may include, but is not limited to: determining the meta information corresponding to the at least one data set according to the attribute information corresponding to the at least one data set, and associating the meta information with the data table.

[0067] Step 104, performing data processing by using the data table and the meta information associated with the data table.

[0068] In one example, the above execution order is only an example given for convenient description. In actual applications, the execution order between steps can also be changed, and no limitation is imposed on this execution order. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification, and the method may include more or fewer steps than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0069] Based on the above technical solution, in the embodiments of the present application, attribute information can be obtained from the data sets of the data source, a data table can be created according to the attribute information, and the meta information corresponding to the data set can be associated with the data table. That is to say, the meta information and the data table can be automatically associated without the user providing the meta information and associating the meta information with the data table, thereby reducing the workload of the user, improving the user experience, greatly improving the construction efficiency of the meta information, and enhancing the overall use efficiency and experience of the data lake analysis system.

[0070] Based on the same inventive concept as the above method, the embodiments of the present application also propose another data processing method, which can be applied to a data lake analysis platform (i.e., the cloud computing platform in the data lake analysis system). The data lake analysis platform is used to provide a serverless data processing service for users. The method includes:

[0071] A data processing request is obtained, and the data processing request may include location information of a data source, and the data source may include a cloud database provided by a data lake analysis platform. Then, attribute information is obtained from a data set of the data source according to the location information; wherein the data source may include multiple data sets, and each data set may include attribute information of the data set. A data table is created according to the attribute information, and the data table may correspond to at least one data set, and meta information corresponding to the at least one data set is associated with the data table. Then, data processing is performed using the data table and the meta information associated with the data table.

[0072] The above data sources may include cloud databases provided by the data lake analysis platform, and the cloud databases may be used to provide serverless query analysis services. The data lake analysis platform may be a storage-based cloud platform that focuses on data storage, or a computing-based cloud platform that focuses on data processing, or a comprehensive cloud computing platform that combines computing and data storage processing. There is no restriction on the data lake analysis platform.

[0073] The cloud database provided for the data lake analysis platform can be used to provide users with serverless query and analysis services. It can analyze and query massive amounts of data in any dimension, and supports high concurrency, low latency (millisecond-level response), real-time online analysis, massive data query and other functions.

[0074] Based on the same application concept as the above method, a data processing method is also proposed in an embodiment of the present application, which method may include: obtaining a data processing request, which may include location information of a data source; obtaining attribute information from a data set of the data source according to the location information; then, creating a data table according to the attribute information, which data table may correspond to at least one data set of the data source, and associating metadata corresponding to the at least one data set with the data table; wherein the association relationship between the data table and the metadata can be used for data processing, and there is no restriction on this data processing process.

[0075] Based on the same application concept as the above method, an embodiment of the present application further proposes a data processing method, which may include: obtaining a data processing request, where the data processing request may include location information of a data source. Obtaining attribute information from the data set of the data source according to the location information; wherein, the data source may include multiple data sets, and each data set may include attribute information of the data set. Then, clustering the multiple data sets according to the attribute information corresponding to each of the multiple data sets to obtain a clustering set, where the clustering set may include at least one data set. Creating a data table for the clustering set, where the data table may correspond to the at least one data set; associating the meta-information corresponding to the at least one data set with the data table. Performing data processing using the data table and the meta-information associated with the data table.

[0076] Based on the same application concept as the above method, an embodiment of the present application further proposes a data processing method, which may include: obtaining a data query request, where the data query request includes data table information.

[0077] Obtaining the data table corresponding to the data table information and the meta-information associated with the data table; wherein, the data table is created according to the attribute information of the data sets in the data source, and the meta-information associated with the data table may include the meta-information corresponding to at least one data set of the data source; specifically, for the creation process of the data table and the association process between the data table and the meta-information, reference may be made to the above embodiments and will not be elaborated herein.

[0078] Processing the query request using the data table and the meta-information associated with the data table, which will not be elaborated herein.

[0079] The above data processing method will be further described below in combination with specific application scenarios.

[0080] See Figure 2 As shown, it is a schematic structural diagram of a Data Lake Analytics system. The Data Lake Analytics system may include a client, a load balancing device, a front node (which may also be referred to as a front-end server), a compute node (which may also be referred to as a compute server), and a database. Of course, the Data Lake Analytics system may also include other servers, which are not limited herein.

[0081] In Figure 2 taking 3 front nodes as an example, in actual applications, the number of front nodes may also be other numbers, which are not limited herein. In Figure 2In this case, taking 4 computing nodes as an example, in actual applications, the number of computing nodes can also be other numbers, and there is no restriction on this. Since the processing flow of each front-end node is the same and the processing flow of each computing node is the same, for the convenience of description, in the subsequent embodiments, the processing flow of 1 front-end node and the processing flow of 1 computing node are taken as examples.

[0082] In Figure 2 this case, taking 5 databases as an example, in actual applications, the number of databases can also be other numbers, and there is no restriction on this. These databases are the data sources. In this embodiment, it can be a scenario for heterogeneous data sources, that is to say, these databases can be databases of the same type or different types. These databases can be relational databases or non-relational databases.

[0083] Furthermore, for each database, the type of this database can also include but is not limited to: OSS (Object Storage Service), TableStore, HBase (Hadoop Database), HDFS (Hadoop Distributed File System), MySQL (i.e., relational database), RDS (Relational Database Service), DRDS (Distribute Relational Database Service), RDBMS (Relational Database Management System), SQLServer (i.e., relational database), PostgreSQL (i.e., object-relational database), MongoDB (i.e., database based on distributed file storage), etc. Of course, the above are just several examples of database types, and there is no restriction on the type of this database.

[0084] Among them, the database is used to store various types of data, and there is no restriction on this data type. For example, it can be user data, commodity data, map data, video data, image data, audio data, etc.

[0085] Among them, the client can be an APP (Application) included in a terminal device (such as a PC (Personal Computer), a laptop, a mobile terminal, etc.), or it can be a browser included in the terminal device, and there is no restriction on this. The load balancing device is used to perform load balancing on the data requests of the client. For example, after receiving a data request, it loads the data request evenly to each front-end node.

[0086] In one example, multiple front-end nodes can be used to provide the same function, forming a resource pool of front-end nodes. For each front-end node in the resource pool, it is used to receive the data request sent by the client, perform SQL (Structured Query Language) parsing on the data request, generate multiple execution plans according to the parsing results, and process these execution plans. For example, the front-end node can send these execution plans to one or more computing nodes, and the computing nodes process the execution plans.

[0087] In one example, multiple computing nodes are used to provide the same function, forming a resource pool of computing nodes. For each computing node in the resource pool, if the computing node receives the execution plan sent by the front-end node, then the computing node can process the execution plan and return the processing result to the front-end node.

[0088] To sum up, the data lake analysis system adopts an architecture with separated storage and computing. The computing nodes read data from different data sources (Data Source), and these data sources are various types of databases.

[0089] In the data lake analysis system, what it provides is a cloud relational database. For example, Figure 2 each of the databases shown is a cloud relational database. If a user needs to use a database for processing, the data lake analysis system can create a database and create a data table in this database. Moreover, the data lake analysis system only needs to associate the meta-information of the data source (that is, Figure 2 the database) with this data table, and does not need to copy the data in the data source to this data table. In this way, for a processing request to access this data source (that is, a processing request to access this data table), the data lake analysis system can use the association relationship between the meta-information and the data table to link to the data source, query data from the data source, and process the data based on the data (such as querying and analyzing, etc.). After the data processing is completed, the currently created data table and database can be deleted.

[0090] In summary, when creating a data table, the data lake analysis system does not need to copy the data in the data source to the data table. Instead, it only needs to associate the metadata of the data source with the data table. That is to say, the data lake analysis system can obtain and store the metadata of the data source and associate the metadata with the data table. Further, as shown in Figure 2 During the execution of the front-end node and the computing node, the Meta (metadata) module can provide the metadata of the data table to the front-end node and the computing node.

[0091] In one example, after creating a database, before creating a data table, the data lake analysis system can also create a Schema in the database, and then create a data table in the Schema. The Schema is mapped to the data set of the data source. A Schema refers to managing and associating a group of tables or relationships.

[0092] In summary, when creating a data table, the data lake analysis system needs to obtain the metadata of the data source and associate the metadata with the data table. In the traditional method, the user needs to define the metadata of the data source. However, since the data lake analysis system supports many types of data sources, and the metadata of different types of data sources vary greatly. For example, for an OSS type of data source, different serialization tools and formats need to be defined, which are not required for an RDBMS type of data source. Therefore, having the user define the metadata of the data source has problems such as a large workload, prone to errors in metadata, and a poor user experience.

[0093] In contrast to the above method, in the embodiments of this application, the data lake analysis system (such as any device of the data lake analysis system, such as the front-end node, the computing node, the Meta module, etc.) can detect the metadata of the data source and associate the metadata with the data table, thereby mapping the metadata to the data table.

[0094] Specifically, for each type of data source, the data-related metadata is reflected in the data source and the original data. Therefore, the data lake analysis system can detect the metadata of the data source, thereby improving the construction efficiency of the metadata, enhancing the overall usage efficiency and experience of the data lake analysis system, optimizing the user's usage path, improving the construction efficiency of the metadata to a certain extent, and making users more adaptable to the arrival of the cloud era.

[0095] See Figure 3As shown in the figure, different types of data sources have different protocol methods. For these types of data sources, the data lake analysis system can obtain relevant meta-information and associate the meta-information with data tables. For example, for data sources such as RDBMS, RDS, SQLServer, and PostgreSQL, there is already the concept of data tables. Therefore, the data lake analysis system can directly obtain relevant information of the data tables through SQL statements and determine the meta-information of the data sources based on this information. Another example is that for data sources such as TableStore and HBase, there is also the concept of data tables, which is the concept of wide tables. Each row of data has primary key and non-primary key information. Therefore, the data lake analysis system can obtain the internal definition information of each data table through the RPC (Remote Procedure Call) interface and determine the meta-information of the data sources based on this information. Another example is that for data sources such as OSS, file, and NAS (Network Attached Storage), the data lake analysis system can obtain relevant file content through file interfaces or file-like interfaces, detect some content in the files according to different file types (such as csv, json, parquet, orc, etc.), analyze the field information in the files, and then construct the structure of each data table. In this way, the meta-information of the data sources can be analyzed.

[0096] In the above application scenario, for data sources such as OSS, file, and NAS, in order to obtain meta-information and associate the meta-information with data tables, the data processing method can be referred to Figure 4 as shown in the figure.

[0097] Step 401, the data lake analysis system obtains a data processing request, such as a data table creation request, etc.

[0098] Specifically, the client can send a data processing request to the data lake analysis system through a load balancing device. In this way, the data lake analysis system can obtain the data processing request.

[0099] Step 402, the data lake analysis system determines whether the data processing request enables the data table discovery function.

[0100] If yes, execute Step 403; if no, the traditional process can be used for processing.

[0101] In one example, the data processing request may include auto-discovery indication information for indicating whether to enable the data table discovery function. Based on this, if the auto-discovery indication information is used to indicate enabling the data table discovery function, it is determined that the data table discovery function is enabled for the data processing request, and step 403 is executed. If the auto-discovery indication information is used to indicate disabling the data table discovery function, it is determined that the data table discovery function is not enabled for the data processing request, and the traditional process is adopted for processing.

[0102] For example, the data processing request may include the auto-discovery indication information "discovertables = true", and this auto-discovery indication information is used to indicate enabling the data table discovery function. Alternatively, the data processing request may include the auto-discovery indication information "discovertables = false", and this auto-discovery indication information is used to indicate disabling the data table discovery function. Of course, the above is only an example and is not limited thereto.

[0103] In another example, if the data processing request does not include the auto-discovery indication information, it is default to enable the data table discovery function, or default to disable the data table discovery function, and this is not limited.

[0104] Step 403, the data lake analysis system obtains the attribute information from the data set of the data source.

[0105] In one example, the data processing request may include the location information of the data source, and the data source may include multiple data sets, and each data set includes the attribute information of the data set. Based on this, the data lake analysis system may obtain the attribute information from the data set of the data source according to the location information.

[0106] For example, the data source may include data set 1, data set 2, and data set 3, and the data processing request may include the location information "location = OSS: / / x.x.x.x:xxx / xxx". Based on this location information, the data lake analysis system determines that the type of the data source is OSS, and the location information of the data source is "x.x.x.x:xxx / xxx". Then, data set 1, data set 2, and data set 3 can be obtained from the location information, and the attribute information 1 of data set 1, the attribute information 2 of data set 2, and the attribute information 3 of data set 3 can be obtained.

[0107] The data lake analysis system can read part of the data in Dataset 1, and then obtain the attribute information 1 of Dataset 1. For example, by reading the first few lines of data in Dataset 1, the attribute information 1 of Dataset 1 can be included in these data. The attribute information 1 can include but is not limited to: file suffix (such as txt, jpg, png, etc.), file type (such as csv, json, parquet, orc, etc.), field information (i.e., column attributes, such as name, age, mobile phone number, ID card, etc.). There is no restriction on the content of this attribute information 1.

[0108] Similarly, referring to the acquisition method of the attribute information 1 of Dataset 1, the data lake analysis system can read part of the data in Dataset 2, and then obtain the attribute information 2 of Dataset 2. The data lake analysis system can read part of the data in Dataset 3, and then obtain the attribute information 3 of Dataset 3, which will not be elaborated here.

[0109] Step 404, the data lake analysis system filters multiple datasets to obtain a target dataset.

[0110] In one example, the data processing request can include filtering indication information, which is used to indicate how to filter multiple datasets of the data source. Therefore, the data lake analysis system can filter multiple datasets of the data source according to this filtering indication information to obtain a target dataset.

[0111] For example, the data processing request can include the filtering indication information "filefilters=json, csv". This filtering indication information "filefilters=json, csv" is used to filter out datasets of json type and csv type. That is to say, datasets of json type can be used as target datasets, and datasets of csv type can also be used as target datasets. However, datasets of other types are not used as target datasets.

[0112] Based on this, assuming that the file type in the attribute information 1 of Dataset 1 is csv, the file type in the attribute information 2 of Dataset 2 is csv, and the file type in the attribute information 3 of Dataset 3 is orc, the data lake analysis system can determine Dataset 1 and Dataset 2 as target datasets.

[0113] Step 405, the data lake analysis system clusters the target dataset according to the attribute information corresponding to the target dataset to obtain at least one clustering set, and each clustering set includes at least one target dataset.

[0114] Specifically, the data lake analysis system can obtain clustering indication information, which is used to indicate clustering sub-attributes. Based on the attribute information corresponding to the target data set, the clustering sub-attributes corresponding to the target data set are determined according to the clustering indication information. According to the clustering sub-attributes corresponding to the target data set, the target data set is clustered to obtain at least one clustering set. For example, the target data sets with the same clustering sub-attributes are clustered into the same clustering set, and the target data sets with different clustering sub-attributes are clustered into different clustering sets.

[0115] For example, the data processing request may include the clustering indication information "clusterAsTable = file type". This clustering indication information "clusterAsTable = file type" is used to indicate that the clustering sub-attribute is the file type. Based on this, the data lake analysis system can determine the file type corresponding to each target data set (the file type is one kind of the attribute information of the target data set), and can cluster the target data sets with the same file type into the same clustering set, and cluster the target data sets with different file types into different clustering sets.

[0116] Suppose the target data set includes data set 1 and data set 2. If the file type of data set 1 is the same as that of data set 2, then data set 1 and data set 2 are clustered into clustering set A. Suppose the target data set includes data set 1 and data set 2. If the file type of data set 1 is different from that of data set 2, then data set 1 is clustered into clustering set B1, and data set 2 is clustered into clustering set B2.

[0117] Of course, the file type is only an example of the clustering sub-attribute. The clustering sub-attribute can also be other sub-attributes, such as file suffix, field information, etc. No detailed restrictions are imposed on this clustering sub-attribute. In practical applications, since the user knows which data sets correspond to the same data table and knows the common characteristics of these data sets, this shared feature can be used as the clustering sub-attribute. For example, if the data sets with the same file type correspond to the same data table, then the clustering sub-attribute can be the file type. Therefore, the clustering indication information carried by the data processing request is used to indicate that the clustering sub-attribute is the file type. Another example is that if the data sets with the same file suffix correspond to the same data table, then the clustering sub-attribute can be the file suffix. Therefore, the clustering indication information carried by the data processing request is used to indicate that the clustering sub-attribute is the file suffix. Another example is that if the data sets with the same field information correspond to the same data table, then the clustering sub-attribute can be the field information. Therefore, the clustering indication information carried by the data processing request is used to indicate that the clustering sub-attribute is the field information.

[0118] In one example, if the clustering indication information is used to indicate that the clustering sub - attribute is the file suffix, the data lake analysis system can determine the file suffix corresponding to each target data set, cluster the target data sets with the same file suffix into the same clustering set, and cluster the target data sets with different file suffixes into different clustering sets. If the clustering indication information is used to indicate that the clustering sub - attribute is the field information, the data lake analysis system can determine the field information corresponding to each target data set, cluster the target data sets with the same field information into the same clustering set, and cluster the target data sets with different field information into different clustering sets.

[0119] Among them, data sets with the same field information can be defined as follows: if the similarity between the field information of data set 1 and the field information of data set 2 is greater than the threshold, it is determined that the field information of data set 1 is the same as the field information of data set 2; if the similarity between the field information of data set 1 and the field information of data set 2 is not greater than the threshold, it is determined that the field information of data set 1 is different from the field information of data set 2.

[0120] For example, the field information of data set 1 is name, age, mobile phone number, ID card, and the field information of data set 2 is name, age, mobile phone number, home address. Then the similarity between the field information of data set 1 and the field information of data set 2 is 75%. If the similarity of 75% is greater than the pre - configured threshold, it means the field information is the same; if the similarity of 75% is not greater than the pre - configured threshold, it means the field information is different.

[0121] Step 406, the data lake analysis system creates a data table for each clustering set. The data table corresponds to at least one target data set, that is, the target data sets corresponding to the clustering set to which the data table belongs.

[0122] For example, assuming that clustering set A corresponds to data set 1 and data set 2, the data lake analysis system can create a data table A for clustering set A, and data table A corresponds to data set 1 and data set 2.

[0123] Another example, assuming that clustering set B1 corresponds to data set 1 and clustering set B2 corresponds to data set 2, the data lake analysis system can create a data table B1 for clustering set B1, and data table B1 corresponds to data set 1, and create a data table B2 for clustering set B2, and data table B2 corresponds to data set 2.

[0124] Step 407, the data lake analysis system names the data table.

[0125] In one example, the data processing request may include naming indication information, which is used to indicate the naming method of the data table. Therefore, the data table can be named according to the naming indication information.

[0126] For example, the data processing request includes naming indication information "tableRenamePrefix=XXX_" and naming indication information "tableRenameSuffix=_YYY". The naming indication information "tableRenamePrefix=XXX_" represents the prefix of the data table, and the naming indication information "tableRenameSuffix=_YYY" represents the suffix of the data table. Based on this, when the data lake analysis system names the data table, the prefix is "XXX_", and the suffix is "_YYY". For the content between the prefix and the suffix, it can be the name of the data source, can also be specified in the data processing request, or can be generated by the data lake analysis system itself, and there is no restriction on this.

[0127] Step 408, the data lake analysis system associates the meta-information corresponding to the data set with the data table.

[0128] For example, if data table A corresponds to data set 1 and data set 2, the meta-information corresponding to data set 1 and the meta-information corresponding to data set 2 can be associated with data table A. Another example, if data table B1 corresponds to data set 1 and data table B2 corresponds to data set 2, the meta-information corresponding to data set 1 can be associated with data table B1, and the meta-information corresponding to data set 2 can be associated with data table B2.

[0129] In one example, the data lake analysis system can determine the meta-information corresponding to the data set according to the attribute information corresponding to the data set, and associate the meta-information corresponding to the data set with the data table.

[0130] In one example, the meta-information corresponding to the data set can include, but is not limited to, one or any combination of the following: the location information of the data source, the communication protocol, the data distribution method, the data read / write protocol, the storage format, the field information (i.e., column attributes), and the field type (such as the field storage type and the field processing type).

[0131] Based on the attribute information corresponding to the data set, such as the file suffix, the file type, the field information, etc., the data lake analysis system can determine the following meta-information: the storage format (i.e., the storage format corresponding to the file suffix, such as txt, jpg, png, etc.), the field type (i.e., the field type corresponding to the file type, such as json, csv, parquet, orc, etc.), and the field information (i.e., column attributes, such as name, age, mobile phone number, ID card, etc.).

[0132] In addition, the data lake analysis system can also obtain other meta-information, such as the location information of the data source (which can be carried in the data processing request), the communication protocol, the data distribution method, and the data read / write protocol (such as OSS, etc., which can be carried in the data processing request), and there is no restriction on the obtaining method.

[0133] Step 409, the data lake analysis system processes data by using the data table and the meta-information associated with the data table. Specifically, after associating the meta-information corresponding to the data set with the data table, that is, after the data table has been established, the data table and the meta-information associated with the data table can be used to process the data.

[0134] For example, after receiving a data query request, the data query request may include data table information, and it is possible to obtain the data table corresponding to the data table information (that is, the data table created according to the attribute information of the data set in the data source, and the specific creation process can be seen in the above embodiments), and the meta-information associated with the data table (that is, the meta-information corresponding to at least one data set of the data source, and the specific content can be seen in the above embodiments). Then, the query request can be processed by using the data table and the meta-information associated with the data table. For example, data query and data analysis and other processing can be performed on the query request, and this processing process is not limited.

[0135] In one example, the above execution order is only an example given for convenient description. In actual applications, the execution order between steps can also be changed, and this execution order is not limited. Moreover, in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification. The steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0136] Based on the above technical solutions, in the embodiments of the present application, the attribute information can be obtained from the data sets in the data source, a data table can be created according to the attribute information, and the meta-information corresponding to the data set can be associated with the data table. That is to say, the meta-information and the data table can be automatically associated without the user providing the meta-information and associating the meta-information with the data table, thereby reducing the workload of the user, improving the user experience, greatly improving the construction efficiency of the meta-information, and enhancing the overall use efficiency and experience of the data lake analysis system.

[0137] In the above embodiments, by modifying the DDL (Data Definition Language, which is used to create, modify, delete, and query the definition information of data tables) enhancement capabilities (i.e., adding the ability to automatically discover data tables), and in cooperation with the data source detection capabilities (detecting table lists, detecting table structures, detecting local data in files, etc.), when creating a data table, the attribute information of each data set can be actively discovered. Then, through mechanisms such as filtering, aggregation, and renaming, a data table is created, and the meta-information is associated with the data table, thereby actively obtaining the meta-information of the data table, which greatly improves the efficiency and experience of users when using a heterogeneous cloud database product such as a data lake analysis system. For example, if a database has 100 data tables, in the traditional method, to create a set of database table structures, 101 SQL statements need to be executed, while in this solution, the SQL statements can be reduced to 1, greatly improving the usage efficiency, enhancing the usage experience, and being simple and universal.

[0138] In practical applications, if the data set does not include the corresponding attribute information (such as data sets in certain data formats that do not carry field information themselves), the user can also provide the attribute information corresponding to the data set. In this way, the data lake analysis system can process using the above process, and the processing of this scenario will not be elaborated here.

[0139] Based on the same application concept as the above method, an embodiment of this application also provides a data processing device, as Figure 5 shown in the structure diagram of the data processing device. The data processing device includes:

[0140] An acquisition module 51, configured to acquire a data processing request, where the data processing request includes the location information of a data source; and acquire attribute information from the data sets of the data source according to the location information; where the data source includes multiple data sets, and the data set includes the attribute information of the data set;

[0141] An association module 52, configured to create a data table according to the attribute information, where the data table corresponds to at least one data set, and associate the meta-information corresponding to the at least one data set with the data table;

[0142] A processing module 53, configured to perform data processing by using the data table and the meta-information associated with the data table.

[0143] When the association module 52 creates a data table according to the attribute information, it is specifically configured to:

[0144] Cluster the multiple data sets according to the attribute information corresponding to the multiple data sets respectively to obtain a clustering set, where the clustering set includes at least one data set;

[0145] Create a data table for the clustering set, where the data table corresponds to the at least one data set.

[0146] When the association module 52 clusters the multiple data sets according to the attribute information respectively corresponding to the multiple data sets to obtain a clustering set, it is specifically used for:

[0147] Obtain clustering indication information, where the clustering indication information is used to indicate clustering sub-attributes;

[0148] Based on the attribute information respectively corresponding to the multiple data sets, determine the clustering sub-attributes respectively corresponding to the multiple data sets according to the clustering indication information, and cluster the multiple data sets according to the clustering sub-attributes respectively corresponding to the multiple data sets to obtain a clustering set.

[0149] Based on the same application concept as the above method, an embodiment of the present application further provides a data processing device, including: a processor and a machine-readable storage medium, where several computer instructions are stored on the machine-readable storage medium, and when the processor executes the computer instructions, the following processing is performed:

[0150] Obtain a data processing request, where the data processing request includes the location information of the data source;

[0151] Obtain attribute information from the data sets of the data source according to the location information; where the data source includes multiple data sets, and the data set includes the attribute information of the data set;

[0152] Create a data table according to the attribute information, where the data table corresponds to at least one data set, and associate the meta-information corresponding to the at least one data set with the data table;

[0153] Perform data processing using the data table and the meta-information associated with the data table.

[0154] An embodiment of the present application further provides a machine-readable storage medium, where several computer instructions are stored on the machine-readable storage medium; when the computer instructions are executed, the following processing is performed:

[0155] Obtain a data processing request, where the data processing request includes the location information of the data source;

[0156] Obtain attribute information from the data sets of the data source according to the location information; where the data source includes multiple data sets, and the data set includes the attribute information of the data set;

[0157] Create a data table according to the attribute information, where the data table corresponds to at least one data set, and associate the meta-information corresponding to the at least one data set with the data table;

[0158] Perform data processing by using the data table and the meta-information associated with the data table.

[0159] See Figure 6 As shown, it is a structural diagram of a data processing device (i.e., any device in the data lake analysis system, such as a front-end node, a computing node, etc.) proposed in an embodiment of the present application. The data processing device 60 may include: a processor 61, a network interface 62, a bus 63, and a memory 64. The memory 64 may be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the memory 64 may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as an optical disc, a DVD, etc.).

[0160] The systems, devices, modules, or units illustrated in the above embodiments may be specifically implemented by a computer chip or an entity, or by a product with certain functions. A typical implementation device is a computer, and the specific form of the computer may be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.

[0161] For convenience of description, when describing the above devices, they are described as various units according to functions. Of course, when implementing the present application, the functions of each unit may be implemented in the same or multiple software and / or hardware.

[0162] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0163] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more flows and / or blocks Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks

[0164] Moreover, these computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one or more flows Figure 1 or one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks

[0165] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 or one or more flows and / or blocks Figure 1 or steps for implementing the functions specified in one or more blocks

[0166] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application

Claims

1. A data processing method, characterized in that, The method includes: Obtaining a data processing request, where the data processing request includes location information of a data source; If the data processing request further includes an auto-discovery indication message, then determining whether to enable a data table discovery function for the data processing request according to the auto-discovery indication message; If so, obtaining attribute information from the data sets of the data source according to the location information; where the data source includes multiple data sets, and the data set includes attribute information of the data set; Creating a data table according to the attribute information, where the data table corresponds to at least one data set, and associating meta-information corresponding to the at least one data set with the data table; Performing data processing by using the data table and the meta-information associated with the data table.

2. The method according to claim 1, characterized in that, Creating a data table according to the attribute information, where the data table corresponds to at least one data set, includes: Clustering the multiple data sets according to the attribute information respectively corresponding to the multiple data sets to obtain a clustering set, where the clustering set includes at least one data set; Creating a data table for the clustering set, where the data table corresponds to the at least one data set.

3. The method according to claim 2, characterized in that, Clustering the multiple data sets according to the attribute information respectively corresponding to the multiple data sets to obtain a clustering set, includes: Obtaining clustering indication information, where the clustering indication information is used to indicate clustering sub-attributes; Based on the attribute information respectively corresponding to the multiple data sets, determining the clustering sub-attributes respectively corresponding to the multiple data sets according to the clustering indication information, and clustering the multiple data sets according to the clustering sub-attributes respectively corresponding to the multiple data sets to obtain a clustering set.

4. The method according to claim 3, characterized in that, Clustering the multiple data sets according to the clustering sub-attributes respectively corresponding to the multiple data sets to obtain a clustering set, includes: Based on the clustering sub-attributes respectively corresponding to the multiple data sets, clustering the data sets with the same clustering sub-attributes into the same clustering set, and clustering the data sets with different clustering sub-attributes into different clustering sets.

5. The method according to claim 3, characterized in that, Obtaining clustering indication information, includes: If the data processing request further includes clustering indication information, then obtaining the clustering indication information from the data processing request; or obtaining pre-configured clustering indication information.

6. The method according to claim 2, characterized in that, Clustering the multiple data sets according to the attribute information respectively corresponding to the multiple data sets to obtain a clustering set, includes: If the data processing request further includes filtering indication information, then filtering the multiple data sets according to the filtering indication information to obtain a target data set; clustering the target data set based on the attribute information corresponding to the target data set to obtain a clustering set.

7. The method according to claim 1, characterized in that, After creating the data table according to the attribute information, further includes: if the data processing request further includes naming indication information, then naming the data table according to the naming indication information.

8. The method according to claim 1, characterized in that, Associating the meta-information corresponding to the at least one data set with the data table, includes: Determining the meta-information corresponding to the at least one data set according to the attribute information corresponding to the at least one data set, and associating the meta-information with the data table.

9. A data processing method, characterized in that, Applied to a data lake analysis platform, the data lake analysis platform is used to provide serverless data processing services for users, and the method includes: Obtain a data processing request, where the data processing request includes location information of a data source; If the data processing request further includes an automatic discovery indication message, then determine whether to enable the data table discovery function for the data processing request according to the automatic discovery indication message; If so, obtain attribute information from the data sets of the data source according to the location information; wherein, the data source includes multiple data sets, and the data set includes the attribute information of the data set; Create a data table according to the attribute information, the data table corresponds to at least one data set, and associate the meta information corresponding to the at least one data set with the data table; Perform data processing using the data table and the meta information associated with the data table; Wherein, the data source includes a cloud database provided by the data lake analysis platform.

10. A data processing method, characterized in that, The method includes: Obtain a data processing request, where the data processing request includes location information of a data source; If the data processing request further includes an automatic discovery indication message, then determine whether to enable the data table discovery function for the data processing request according to the automatic discovery indication message; If so, obtain attribute information from the data sets of the data source according to the location information; Create a data table according to the attribute information, the data table corresponds to at least one data set of the data source, and associate the meta information corresponding to the at least one data set with the data table; Wherein, the association relationship between the data table and the meta information is used for data processing.

11. A data processing method, characterized in that, The method includes: Obtain a data query request, where the data query request includes data table information; Obtain the data table corresponding to the data table information and the meta information associated with the data table; wherein, the data table is created according to the attribute information of the data sets in the data source, and the meta information associated with the data table includes the meta information corresponding to at least one data set of the data source; wherein, the process of obtaining the attribute information includes: after obtaining a data processing request, the data processing request includes location information and an automatic discovery indication message of a data source, determine whether to enable the data table discovery function for the data processing request according to the automatic discovery indication message; if so, obtain the attribute information from the data sets of the data source according to the location information; Process the query request using the data table and the meta information associated with the data table.

12. A data processing device, characterized in that, The device includes: An obtaining module, configured to obtain a data processing request, where the data processing request includes location information of a data source; if the data processing request further includes an automatic discovery indication message, then determine whether to enable the data table discovery function for the data processing request according to the automatic discovery indication message; if so, obtain attribute information from the data sets of the data source according to the location information; wherein, the data source includes multiple data sets, and the data set includes the attribute information of the data set; An association module, configured to create a data table according to the attribute information, where the data table corresponds to at least one data set, and associate meta-information corresponding to the at least one data set with the data table; A processing module, configured to perform data processing by using the data table and the meta-information associated with the data table.

13. The device according to claim 12, characterized in that, When creating the data table according to the attribute information, the association module is specifically configured to: Cluster the multiple data sets according to the attribute information respectively corresponding to the multiple data sets to obtain a clustering set, where the clustering set includes at least one data set; Create a data table for the clustering set, and the data table corresponds to the at least one data set.

14. The device according to claim 13, characterized in that, When clustering the multiple data sets according to the attribute information respectively corresponding to the multiple data sets to obtain a clustering set, the association module is specifically configured to: Obtain clustering indication information, where the clustering indication information is used to indicate clustering sub-attributes; Based on the attribute information respectively corresponding to the multiple data sets, determine the clustering sub-attributes respectively corresponding to the multiple data sets according to the clustering indication information, and cluster the multiple data sets according to the clustering sub-attributes respectively corresponding to the multiple data sets to obtain a clustering set.

15. A data processing equipment, characterized in that, Including: A processor and a machine-readable storage medium, where a plurality of computer instructions are stored on the machine-readable storage medium, and when the processor executes the computer instructions, the following processing is performed: Obtain a data processing request, where the data processing request includes location information of a data source; If the data processing request further includes automatic discovery indication information, then determine whether to enable the data table discovery function for the data processing request according to the automatic discovery indication information; If so, obtain attribute information from the data sets of the data source according to the location information; where the data source includes multiple data sets, and the data set includes attribute information of the data set; Create a data table according to the attribute information, where the data table corresponds to at least one data set, and associate meta-information corresponding to the at least one data set with the data table; Perform data processing by using the data table and the meta-information associated with the data table.

Citation Information

Patent Citations

  • Data storage, query and loading methods and devices

    CN104462362A

  • A method and apparatus for processing data

    CN109409419A