Data processing method, device, computer equipment and storage medium

By obtaining non-function information in the query statement to generate executable statements, the problem of built-in functions is solved between different system components, and the efficiency and scalability of cross-system components of unified SQL queries are realized.

CN117632993BActive Publication Date: 2025-05-30SF TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210958682.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-09
Publication Date
2025-05-30
Estimated Expiration
2042-08-09

AI Technical Summary

Technical Problem

The built-in functions between different system components are not equal, resulting in high redundancy, difficulty in utilization, and poor expansion capabilities when queries in unified SQL.

Method used

By obtaining non-function information in the query statement, executable statements are generated to avoid the conversion of function parts and ensure the function equality between different system components.

Benefits of technology

It realizes unified SQL query across system components, avoids problems caused by function ambiguousness, and improves query efficiency and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117632993B_ABST
    Figure CN117632993B_ABST
Patent Text Reader

Abstract

The present application provides a data processing method, apparatus, computer device, and storage medium. The method includes: obtaining a query statement, where the query statement includes multiple fields; determining a query database, non-function information of each of the fields, and a recursive function stack according to the query statement; generating an executable statement for the query database according to the non-function information; obtaining a data set to be parsed, where the data set is data returned by the query database running the executable statement; and obtaining a query result according to the query statement, the recursive function stack, and the data set. By using this method, non-function information that does not include function information is used as an executable statement for the query database, avoiding problems caused by unequal built-in functions between different system components.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of unified SQL query, and specifically relates to a data processing method, device, computer device, and storage medium. Background Art

[0002] In traditional DBMS (Database Manage System), most database execution languages are designed according to the SQL (Structured Query Language) standard; similarly, in big data components such as Hive, Spark, and Flink, there is also its own set of SQL standard designs. The idea of unified SQL is to define a set of SQL language standards in various components or programs that execute SQL, and transform them into DSL (Domain Specific Language) languages that can be executed in specified components or programs through a certain method, so as to manage multiple components or programs in disguise and simplify business operations.

[0003] Generally, each DBMS has its own UDF implementation method. For example, the Mysql database can support adding a dynamic link library packaged by implementing an interface to the Mysql folder or creating a new function in the database through the create function method. Currently popular big data components such as Hive, Spark, and Flink all support custom functions to expand the SQL functions of the component. However, this method is not based on unified SQL, and the problem of unequal built-in functions between different system components needs to be solved. Without unified management of UDF, it is easy to cause high redundancy, difficulty in utilization, and poor extensibility during subsequent iteration. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a data processing method, device, computer device, and storage medium, which convert non-function information that does not contain function information into an executable statement for querying a database, so as to avoid problems caused by unequal built-in functions between different system components.

[0005] In a first aspect, this application provides a data processing method, including:

[0006] Obtain a query statement, where the query statement includes multiple fields;

[0007] Determine a query database, non-function information of each field, and a recursive function stack according to the query statement;

[0008] Generate an executable statement for the query database according to the non-function information;

[0009] Obtain the data set to be parsed, where the data set is the data returned by running the executable statement on the query database;

[0010] Obtain the query result according to the query statement, the recursive function stack, and the data set.

[0011] In some embodiments of the present application, the non-function information includes condition information and query information. Determining the non-function information and the recursive function stack of each field according to the query statement includes:

[0012] Determine the query part and the condition part according to the query statement;

[0013] Obtain the condition parameters according to the condition part, and set the condition parameters as the condition information;

[0014] If the query part contains a function, determine the recursive function stack according to the function part of the query part, and determine the query information according to the non-function part of the query part;

[0015] If the query part does not contain a function, determine the query part as the query information.

[0016] In some embodiments of the present application, determining the recursive function stack according to the function part of the query part includes:

[0017] If the function part only contains one layer of function, determine the function name and the input parameter group of the function as the recursive function stack, and the input parameter group is the query parameter of the corresponding field of the function;

[0018] If the function part contains multiple layers of functions, analyze the function call relationship of the function part and the function names and input parameter groups of each layer of functions;

[0019] Determine the recursive function stack according to the function call relationship, the function names and input parameter groups of each layer of functions. The input parameter group of the top function of the recursive function stack is the query parameter of the corresponding field of the recursive function stack.

[0020] In some embodiments of the present application, the non-function information further includes parameter information. After determining the query part as the non-function information if the query part does not contain a function, it includes:

[0021] Obtain the preset query parameter type and the query parameters to be queried in each field;

[0022] If the query parameter to be queried does not belong to the query parameter type, determine the query parameter to be queried as the parameter information.

[0023] In some embodiments of the present application, determining the query part and the condition part according to the query statement includes:

[0024] Generating an abstract syntax tree according to the query statement;

[0025] Traversing the abstract syntax tree to determine condition identification points;

[0026] Determining the query part and the condition part according to the condition identification points.

[0027] In some embodiments of the present application, generating an executable statement for the query database according to the non-function information includes:

[0028] Generating a valid query statement according to the non-function information of each field;

[0029] Obtaining the transformation syntax of the query database;

[0030] Obtaining the executable statement according to the valid query statement and the transformation syntax.

[0031] In some embodiments of the present application, obtaining a query result according to the query statement, the recursive function stack, and the data set includes:

[0032] If the target field in the query statement contains a function, determining the query parameter of the target field according to the recursive function stack corresponding to the target field and the data set;

[0033] Determining the field query result corresponding to the target field according to the recursive function stack corresponding to the target field and the query parameter;

[0034] If the target field in the query statement does not contain a function, determining the field query result of the target field according to the non-function information corresponding to the target field and the data set;

[0035] Determining the query result according to the field query results of the fields in the query statement.

[0036] In a second aspect, the present application provides a data processing device, including:

[0037] A statement acquisition module, configured to acquire a query statement, where the query statement includes multiple fields;

[0038] A statement analysis module, communicatively connected to the statement acquisition module, and configured to determine a query database, non-function information, and a recursive function stack of each of the fields according to the query statement;

[0039] An instruction generation module, communicatively connected to the statement analysis module, for generating an executable statement for querying the database according to the non-function information;

[0040] A data acquisition module, communicatively connected to the instruction generation module, for acquiring a data set to be parsed, where the data set is the data returned by the query database running the executable statement;

[0041] A data parsing module, communicatively connected to the statement acquisition module, the statement analysis module, and the data acquisition module, for obtaining a query result according to the query statement, the recursive function stack, and the data set.

[0042] In some embodiments of the present application, the non-function information includes condition information and query information. The statement analysis module is further configured to determine a query part and a condition part according to the query statement; obtain condition parameters according to the condition part, and set the condition parameters as the condition information; if the query part includes a function, determine the recursive function stack according to the function part of the query part, and determine the query information according to the non-function part of the query part; if the query part does not include a function, determine the query part as the query information.

[0043] In some embodiments of the present application, the statement analysis module is further configured to, if the function part only includes one layer of function, determine the function name and the input parameter group of the function as the recursive function stack, where the input parameter group is the query parameter of the corresponding field of the function; if the function part includes multiple layers of functions, analyze the function call relationship of the function part and the function name and input parameter group of each layer of function; determine the recursive function stack according to the function call relationship, the function name and input parameter group of each layer of function, and the input parameter group of the top function of the recursive function stack is the query parameter of the corresponding field of the recursive function stack.

[0044] In some embodiments of the present application, the non-function information further includes parameter information. The statement analysis module is further configured to obtain a preset query parameter type and the query parameters to be queried in each field; if the query parameter to be queried does not belong to the query parameter type, determine the query parameter to be queried as the parameter information.

[0045] In some embodiments of the present application, the statement analysis module is further configured to generate an abstract syntax tree according to the query statement; traverse the abstract syntax tree to determine condition identification points; and determine the query part and the condition part according to the condition identification points.

[0046] In some embodiments of the present application, the instruction generation module is further configured to generate an effective query statement according to the non-function information of each field; obtain the conversion syntax of the query database; and obtain the executable statement according to the effective query statement and the conversion syntax.

[0047] In some embodiments of the present application, the data parsing module is further configured to: if the target field in the query statement contains a function, determine the query parameter of the target field according to the recursive function stack corresponding to the target field and the data set; determine the field query result corresponding to the target field according to the recursive function stack corresponding to the target field and the query parameter; if the target field in the query statement does not contain a function, determine the field query result of the target field according to the non-function information corresponding to the target field and the data set; and determine the query result according to the field query results of each field in the query statement.

[0048] In a third aspect, the present application further provides a computer device, which includes:

[0049] One or more processors;

[0050] A memory; and one or more applications, where the one or more applications are stored in the memory and are configured to be executed by the processor to implement the data processing method described above.

[0051] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is loaded by a processor to execute the steps in the data processing method described above.

[0052] In a fifth aspect, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in the first aspect above.

[0053] For the data processing method, apparatus, computer device, and storage medium described above, the function part and the non-function part of each field in the query statement are separated, and only the non-function information that does not contain function information is converted into an executable statement for querying the database, so that there is no function part in the executable statement, avoiding problems caused by unequal built-in functions between different system components. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0055] Figure 1It is a schematic flowchart of the data processing method in an embodiment of the present application;

[0056] Figure 2 It is a schematic flowchart of the data processing method in another embodiment of the present application;

[0057] Figure 3 It is a schematic diagram of the triples corresponding to the SELECT query part in another embodiment of the present application;

[0058] Figure 4 It is a schematic flowchart of performing a stack push operation on a data set in combination with a triple set in another embodiment of the present application;

[0059] Figure 5 It is a schematic structural diagram of the data processing device in an embodiment of the present application;

[0060] Figure 6 It is a schematic structural diagram of the computer device in an embodiment of the present application. Detailed implementation manners

[0061] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0062] In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality of" means two or more unless otherwise specifically defined.

[0063] In the description of the present application, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present application is not necessarily construed as being more preferred or more advantageous than other embodiments. In order for any person skilled in the art to implement and use the present invention, the following description is given. In the following description, details are set forth for the purpose of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed in the present application.

[0064] Refer toFigure 1 , an embodiment of the present application provides a data processing method, and the method includes steps S101 to S105, specifically as follows:

[0065] S101, obtain a query statement, where the query statement includes multiple fields.

[0066] Specifically, the query statement is a statement of unified SQL. The query statement includes multiple fields, and each field is separated by symbols, etc. Each field is the result to be queried. For example, the query statement is SELECT name, udfA(row_key) AS reverse_row_key, udfB(udfC(name, id)), CONCAT(name, '%') AS np FROM user WHERE inc_day = udfA(${yyyyMMdd}). Then, the query statement is separated into multiple fields by the symbol ",", namely: name, udfA(row_key) AS reverse_row_key, udfB(udfC(name, id)), and CONCAT(name, '%') AS np FROM user WHERE inc_day = udfA(${yyyyMMdd}). Correspondingly, the results to be queried are name, udfA(row_key), udfB(udfC(name, id)), and name, '%'. Each field may include a function, where the function may be an embedded function supported by unified SQL or a custom function. Inputting information such as definition meta-information, function name, function input and output parameter information, life cycle, and permissions determines the corresponding custom function.

[0067] Among them, an abstract syntax tree is generated according to the query statement, and syntax verification is performed to verify the legality of the query statement SQL input by the user. Warnings and prompts are given for items that do not conform to the syntax. For the functions included therein, if they are not embedded functions supported by SQL nor configured custom functions, warnings and prompts are also given.

[0068] S102, determine the query database, non-function information of each field, and recursive function stack according to the query statement.

[0069] Specifically, unified SQL supports querying multiple databases, but the query statement of unified SQL needs to be translated into a language executable by the specified database. Therefore, the query database, that is, the object of querying data, is determined by the fields in the query statement.

[0070] In addition, non-function information and a recursive function stack for each field in the query statement are determined, that is, the function part and the non-function part of each field in the query statement are separated, and the non-function information that does not contain function information is used as an instruction for subsequent queries to avoid problems caused by unequal built-in functions between different system components.

[0071] In one embodiment, the non-function information includes condition information and query information. This step includes: S201, determining a query part and a condition part according to the query statement; S202, obtaining condition parameters according to the condition part, and setting the condition parameters as the condition information; S203, if the query part includes a function, determining the recursive function stack according to the function part of the query part, and determining the query information according to the non-function part of the query part; S204, if the query part does not include a function, determining the query part as the query information.

[0072] Specifically, querying data in a database is performed under certain conditions. The preprocessing of the condition part in the query statement needs to be performed before converting the query statement of unified SQL into a specific executable statement. Therefore, the query part and the condition part are determined according to the query statement, and the condition part is preprocessed, that is, condition parameters are obtained according to the condition part, and the condition parameters are used as the condition information in the non-function information. Among them, the condition information can be a static condition or a dynamic condition. In this embodiment, preprocessing the condition part separately can better accommodate dynamic conditions.

[0073] In addition, the query part includes a first type of field and a second type of field. The first type of field includes a function, such as the field udfB(udfC(name,id)) mentioned above, and the second type of field does not include a function, such as the field name mentioned above. For the second type of field that does not include a function, there is no conflict in the conversion between unified SQL and the language for querying the database. Therefore, the second type of field is retained as the query information for converting the executable statement for querying the database.

[0074] For the first type of fields containing functions, the first type of fields can only contain fields of functions, such as the field udfB(udfC(name,id)) in the above example, or can be fields containing both functions and non-functions, such as the field udfA(row_key)AS reverse_row_key in the above example, which contains the function udfA(row_key) and the non-function reverse_row_key. Among them, the non-function is the name of the function, and the name is used to define and distinguish the function, and it is not a necessary part of the field and can be set according to user needs. If the functions contained therein are custom functions or functions that are not equivalent to the query database, the function part may cause errors in data query. Therefore, the function part and the non-function part in the first type of fields are separated. The non-function part is retained as query information for converting into an executable statement for querying the database, and the function part is converted into a recursive function stack according to a preset function mapping relationship. The function mapping relationship is used to sort out and analyze the function call logic of each function part, so as to quickly obtain the query information in combination with the recursive function stack after the parameters therein are subsequently queried.

[0075] It should be noted that if the first type of field only contains the function part and does not contain the non-function part, then the first type of field only has the corresponding recursive function stack and no corresponding query information.

[0076] In one embodiment, S201, determining a query part and a condition part according to the query statement includes: S301, generating an abstract syntax tree according to the query statement; S302, traversing the abstract syntax tree to determine a condition identification point; S303, determining the query part and the condition part according to the condition identification point.

[0077] Specifically, the condition identification point is the identification information of the condition part and can be set as needed. For example, taking the first where keyword in the query statement as the condition identification point, then the part after the where keyword in the query statement is the condition part, and the part before the where keyword is the query part. In addition, traversing the abstract syntax tree is a way to determine the condition identification point, and other search methods that can implement finding the condition identification point are also included within the scope of this embodiment.

[0078] In one embodiment, the step of determining the recursive function stack according to the function part of the query part includes: S401, if the function part only contains one layer of function, determine the function name and the input parameter group of the function as the recursive function stack, and the input parameter group is the query parameter of the corresponding field of the function; S402, if the function part contains multiple layers of functions, analyze the function call relationship of the function part and the function name and input parameter group of each layer of functions; S403, determine the recursive function stack according to the function call relationship, the function name and input parameter group of each layer of functions, and the input parameter group of the top function of the recursive function stack is the query parameter of the corresponding field of the recursive function stack.

[0079] Specifically, for each function, it contains two elements: the function name and the input parameter group. The input parameter group is the input parameter required by the function corresponding to the function name. Therefore, a binary tuple (function name, input parameter group) containing two elements, the function name and the input parameter group, is set to identify each function.

[0080] If the function part in the first type of field only contains one layer of function, directly determine the function name and the input parameter group of the function as the recursive function stack. The input parameter group is the query parameter of the corresponding field of the function, and the query parameter is the data that can be directly queried by the database. For example, if the first type of field is udfC(name,id), the corresponding recursive function stack is (udfC, [name,id]), and the input parameter group, that is, the query parameter, is [name,id], and name and id can be directly obtained from the database.

[0081] If the function part contains multiple layers of functions, analyze the function call relationship of the function part and the function name and input parameter group of each layer of functions. The function call relationship is the reference relationship between each layer of functions, and a certain function is the input parameter group of another function. For example, if the first type of field is udfB(udfC(name,id)), the function udfC(name,id) is the input parameter group of another function named udfB. Determine the recursive function stack according to the function call relationship, the function name and input parameter group of each layer of functions. Among them, a certain layer of function in the recursive function stack directly calls the adjacent layer of function. For example, if the first type of field is udfA(udfB(udfC(name,id))), the obtained recursive function stack from the bottom to the top is (udfA, [udfB(udfC(name,id))]), (udfB, [udfC(name,id)]), (udfC, [name,id]) in turn. The input parameter group of the top function of the recursive function stack does not have function input parameters and is the query parameter of the corresponding field of the recursive function stack, and the query parameter is the data that can be directly queried by querying the database.

[0082] In one embodiment, the non-function information further includes parameter information. After step S204, if the query part does not contain a function and the query part is determined as the non-function information, the following steps are included: S501, obtain the preset query parameter type and the query parameters to be queried in each of the fields; S502, if the query parameter to be queried does not belong to the query parameter type, determine the query parameter to be queried as the parameter information.

[0083] Specifically, a preset query parameter type is set in the unified SQL. Generally, the database queries the parameters corresponding to the query parameter type. However, since the query statement may contain custom functions, and there may be parameters to be queried based on new requirements in the custom functions. If these parameters are not in the preset query parameter type, and the recursive function stack corresponding to the function containing these parameters will not be included in the executable statement. Therefore, it is necessary to obtain the query parameters to be queried in each field of the query statement. If the query parameter to be queried does not belong to the query parameter type, then determine the query parameter to be queried as the parameter information, which is used as non-function information to generate the executable statement for querying the database.

[0084] S103, generate the executable statement for querying the database according to the non-function information.

[0085] Specifically, use the non-function information in the query statement to generate the executable statement for querying the database, and strip the functions in the query statement to avoid problems caused by the inequality of built-in functions between different system components.

[0086] In one embodiment, this step includes: S601, generate an effective query statement according to the non-function information of each field; S602, obtain the conversion syntax of the query database; S603, obtain the executable statement according to the effective query statement and the conversion syntax.

[0087] Specifically, generate an effective query statement according to the non-function information of each field. The non-function information includes condition information, query information, and parameter information. The condition information is the condition parameter corresponding to the condition part, and the parameter information is the query parameter to be queried that is newly added and not in the preset query parameter type. The query information is the non-function part included in the field of the query part. Among them, for the field of the query part that only contains the function part, there is no corresponding query information.

[0088] For example, a query statement is "SELECT name, udfA(row_key) AS reverse_row_key, udfB(udfC(name, id)), CONCAT(name, '%') AS np FROM user WHERE inc_day = udfA(${yyyyMMdd})". The first "where" keyword is the condition identification point. The part before it is the unqueried part, and the part after it is the condition part. In "s_date = udfA(${yyyyMMdd})", the user-defined function udfA and the user-defined dynamic input parameter ${yyyyMMdd} are involved. The specific value is obtained according to the unified SQL dynamic parameter. For example, here the dynamic input parameter represents the current date and its format, 20220429. The user-defined function udfA automatically adapts to the defined date format of the specific database and specific table, and thus obtains the corresponding condition information for the condition part. The duplicate check and filtering are performed on the query parameters of each field in the query part, and the parameter whose query parameter id is not of the query parameter type is obtained. Therefore, the query parameter id is the parameter information. For each field in the query part, "name" and "CONCAT(name, '%') AS np FROM user" do not contain functions, so they are directly used as query information. "udfB(udfC(name, id))" only contains functions, so there is no corresponding query information. The non-function part contained in "udfA(row_key) AS reverse_row_key" is "row_key" as the query information. Therefore, the final effective query statement is "SELECT id, name, row_key, CONCAT(name, '%') as np FROM user WHERE s_date = [preprocessing execution result]".

[0089] In addition, since the syntax rules of different databases are different, the conversion syntax of the query database corresponding to this query statement is obtained, and the effective query statement is converted into an executable statement of the query database based on the conversion syntax and then sunk to the query database for running.

[0090] S104, obtain the dataset to be parsed, where the dataset is the data returned by the query database running the executable statement.

[0091] Specifically, after the query database runs the executable statement, the queried dataset is returned. Therefore, the dataset to be parsed is obtained, and the dataset contains the data corresponding to each query field in the executable statement. Among them, the dataset can be the encrypted data in the query database.

[0092] S105, obtain the query result according to the query statement, the recursive function stack, and the dataset.

[0093] Specifically, for each field in the query statement, the corresponding query result is obtained from the data set. For the fields containing functions, they are parsed in combination with the recursive function stack to quickly and clearly obtain the query result with a clear logic.

[0094] In one embodiment, this step includes: S701, if the target field in the query statement contains a function, determine the query parameter of the target field according to the recursive function stack corresponding to the target field and the data set; S702, determine the field query result corresponding to the target field according to the recursive function stack corresponding to the target field and the query parameter; S703, if the target field in the query statement does not contain a function, determine the field query result of the target field according to the non - function information corresponding to the target field and the data set; S704, determine the query result according to the field query results of each field in the query statement.

[0095] Specifically, if the target field in the query statement does not contain a function, the corresponding field query result can be directly obtained from the data set according to the non - function information of this target field. For example, for the target field "name", directly obtain the information corresponding to "name" in the data set as the field query result of this target field. If the target field contains a function, obtain the corresponding parameters from the data set according to the query parameters required by the function in the target field, and then push the parameters onto the recursive function stack to obtain the field query result of the target field. For example, for the target field "udfB(udfC(name,id))", first obtain the query parameters "(name,id)", use the query parameters as the input parameter group into the function udfC to get the first - layer function result, and then use the first - layer function result as the input parameter group into the function udfB to get the field query result of this target field. Finally, combine the field query results of all target fields to obtain the query result.

[0096] In this embodiment, when converting the unified SQL into an executable statement for a specific database, the function part is stripped, so it supports users to customize functions based on the unified SQL, makes up for the gap in UDF of the data service provided by the unified SQL, and at the same time realizes non - invasiveness to each component.

[0097] An embodiment of this application provides a data processing method, which is applied to a unified SQL processing module and a UDF function management module. The unified SQL processing module defines the normative syntax of unified SQL, manages the executable information of each data source (including supported data types, database types that support target translation, specific database built-in function information, etc.), and is responsible for translating unified SQL into SQL executable by a specified database. The UDF function management module is based on the unified SQL processing module and is responsible for managing user-defined functions (including meta-information, including function names, function input and output parameter information, life cycle, permissions). It includes a preprocessing module and a postprocessing module. The preprocessing module preprocesses the query statement of unified SQL, and the postprocessing module processes the data processing operation returned after SQL execution.

[0098] Among them, for the convenience of logical sorting, each field in the query statement is mapped to a triple. The triple structure is (database field name / stack push automaton alias, returned alias, stack push automaton). The stack push automaton (i.e., recursive function stack) is a calculation stack implemented to solve the possible recursive call problem in the user-defined function UDF. The stack push automaton is generated during the UDF operation. The structure is a binary tuple. The first element is the function name, and the second element is the input parameter array. The function input parameter of the top element of the stack must be the database queryable field name. It should be noted that both the triple and the stack push automaton are only implementation methods for parsing non-function information and recursive function stacks in the query statement, and should not be understood that this embodiment is limited to this. As Figure 2 shown, the data processing method includes the following steps:

[0099] S1. Define UDF: The user defines the UDF function function through the specification of the UDF function management module, submits it to the UDF function management module for management in the form of a jar package, and configures the permission information of the UDF function. After that, the user-defined function can be used in the query statement.

[0100] S2. Input the query statement of unified SQL, such as SELECT name, udfA(row_key) AS reverse_row_key, udfB(udfC(name, id)), CONCAT(name, '%') AS np FROM user WHERE inc_day = udfA(${yyyyMMdd}).

[0101] S3. According to the first where keyword, the part after the where keyword, that is, the where condition part, is defined as the preprocessing part, and the part before the where keyword, that is, the SELECT query part, is defined as the postprocessing part. Abstract syntax trees are generated respectively based on the syntax rules of unified SQL, and legality checks are also performed simultaneously. For example, in the query statement of the above steps, SELECT name, udfA(row_key) AS reverse_row_key, udfB(udfC(name, id)), CONCAT(name, '%') AS np FROM user

[0102] For the postprocessing part, it is allocated to the postprocessing module; s_date = udfA(${yyyyMMdd}) is allocated to the preprocessing module.

[0103] S4. The preprocessing module traverses the abstract syntax tree generated by the where condition part in a depth-first manner, generates a stack push automaton for the UDF part, performs stack push operations, and replaces the tree nodes generated by user-defined variables with the operation results. For example, in the where condition part of the above steps, for the user-defined function udfA in s_date = udfA(${yyyyMMdd}) and the user-defined dynamic input parameter ${yyyyMMdd}. The specific value is obtained according to the dynamic parameter management in the unified SQL module. For example, here the dynamic input parameter represents the current date and its format, 20220429; the user-defined function udfA automatically adapts to the defined date format of the specific database and specific table. The constructed stack push automaton is (udfA, [20220429]).

[0104] S5. The post-preprocessing module traverses the abstract syntax tree generated by the SELECT query part in a depth-first manner, strips the UDF in the tree, generates a stack push automaton, and constructs a triple mapping each query field. The triple contains non-function information and a recursive function stack. For example, based on the triple structure (database field name / stack push automaton alias, returned alias, stack push automaton), non-existent elements are identified with specific strings. The triple corresponding to the SELECT query part of the above steps is as follows Figure 3 shown.

[0105] S6. Construct a triple set according to the triples mapping each query field.

[0106] S7. The effective query statement of the unified SQL to be finally executed for translation is obtained by splicing the abstract syntax trees obtained from the preprocessing module and the postprocessing module. For example, the effective query statement finally spliced from the query statements of the above steps is: SELECT id, name, row_key, CONCAT(name, '%') as np FROM user WHERE s_date = [preprocessing execution result]. Among them, for the fields containing functions, the first element of the non-stack push automaton alias in the triple is taken for splicing, such as row_key in the effective query statement. For the fields without functions, the entire field is spliced, such as name and CONCAT(name, '%') as np FROM user in the effective query statement. For the fields containing functions but without an alias returned, there is no element for splicing, such as the field udfB(udfC(name, id)) in the query statement where no element is retained in the effective query statement. In addition, the query parameters outside the preset query parameter types are also spliced, such as the query parameter id in the query statement.

[0107] S8. The unified SQL module translates the effective query statement into an executable statement for querying the database and assigns it to the query database component for execution.

[0108] S9. The query database returns the dataset to be parsed, and stack push operations are performed in combination with the triple set to obtain the query result. As Figure 4 shown, for each row of data, if there is a stack push automaton, stack push operations are performed; if there is no stack push automaton but there is an alias, the query information is obtained according to the alias; if there is neither a stack push automaton nor an alias, the query information is obtained according to the field name. Among them, the stack push operations are processed starting from the top of the stack according to the last-in, first-out principle of the stack, and the obtained result is returned to the next layer as an input parameter. For example, for the stack push automaton 2 as Figure 3 shown, the input parameter array of the top element of the stack must be the database return field. Execute udfC(name), and the returned result is used as the input parameter of udfB. Reach the bottom of the stack to obtain the final result, and combine the triple to obtain the corresponding name.

[0109] The UDF management system based on unified SQL in this embodiment can, by customizing UDFs, use the set specific custom UDF as a decryption tool to decrypt the encrypted dataset data for the interface, and simply and quickly implement the encryption and decryption process during the application of unified SQL. Similarly, through custom functions, processes such as translation and transposition are defined in UDFs, improving the reusability of the system code.

[0110] As Figure 5As shown in the figure, an embodiment of the present application provides a data processing device 900, including:

[0111] A statement acquisition module 910, configured to acquire a query statement, where the query statement includes multiple fields;

[0112] A statement analysis module 920, communicatively connected to the statement acquisition module 910, configured to determine a query database, non-function information of each of the fields, and a recursive function stack according to the query statement;

[0113] An instruction generation module 930, communicatively connected to the statement analysis module 920, configured to generate an executable statement for the query database according to the non-function information;

[0114] A data acquisition module 940, communicatively connected to the instruction generation module 930, configured to acquire a data set to be parsed, where the data set is data returned by the query database running the executable statement;

[0115] A data parsing module 950, communicatively connected to the statement acquisition module 910, the statement analysis module 920, and the data acquisition module 940, configured to obtain a query result according to the query statement, the recursive function stack, and the data set.

[0116] In some embodiments of the present application, the non-function information includes condition information and query information. The statement analysis module 920 is further configured to determine a query part and a condition part according to the query statement; obtain condition parameters according to the condition part, and set the condition parameters as the condition information; if the query part includes a function, determine the recursive function stack according to the function part of the query part, and determine the query information according to the non-function part of the query part; if the query part does not include a function, determine the query part as the query information.

[0117] In some embodiments of the present application, the statement analysis module 920 is further configured to, if the function part only includes one layer of function, determine the function name and the input parameter group of the function as the recursive function stack, where the input parameter group is the query parameter of the corresponding field of the function; if the function part includes multiple layers of functions, analyze the function call relationship of the function part and the function names and input parameter groups of each layer of functions; determine the recursive function stack according to the function call relationship, the function names and input parameter groups of each layer of functions, and the input parameter group of the top function of the recursive function stack is the query parameter of the corresponding field of the recursive function stack.

[0118] In some embodiments of the present application, the non - function information further includes parameter information, and the statement analysis module 920 is further configured to obtain a preset query parameter type and the query parameters to be queried in each of the fields; if the query parameter to be queried does not belong to the query parameter type, it is determined that the query parameter to be queried is the parameter information.

[0119] In some embodiments of the present application, the statement analysis module 920 is further configured to generate an abstract syntax tree according to the query statement; traverse the abstract syntax tree to determine condition identification points; and determine the query part and the condition part according to the condition identification points.

[0120] In some embodiments of the present application, the instruction generation module 930 is further configured to generate a valid query statement according to the non - function information of each field; obtain the conversion syntax of the query database; and obtain the executable statement according to the valid query statement and the conversion syntax.

[0121] In some embodiments of the present application, the data parsing module 950 is further configured to, if the target field in the query statement contains a function, determine the query parameter of the target field according to the recursive function stack corresponding to the target field and the data set; determine the field query result corresponding to the target field according to the recursive function stack corresponding to the target field and the query parameter; if the target field in the query statement does not contain a function, determine the field query result of the target field according to the non - function information corresponding to the target field and the data set; and determine the query result according to the field query results of each field in the query statement.

[0122] In some embodiments of the present application, the data processing device 900 may be implemented in the form of a computer program, and the computer program can run on a computer device as Figure 6 shown. In the memory of the computer device, each program module constituting the data processing device 900 can be stored, for example, Figure 5 the statement acquisition module 910, the statement analysis module 920, the instruction generation module 930, the data acquisition module 940, and the data analysis module 950 as shown. The computer program constituted by each program module enables the processor to execute the steps in the data processing method of each embodiment of the present application described in this specification.

[0123] For example, Figure 6 the computer device shown can be connected through, for example, Figure 5The statement acquisition module 910 in the data processing device 900 shown executes step S101. The computer device can execute step S102 through the statement analysis module 920. The computer device can execute step S103 through the instruction generation module 930. The computer device can execute step S104 through the data acquisition module 940. The computer device can execute step S105 through the data analysis module 950. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external computer device through a network connection. When the computer program is executed by the processor, it implements a data processing method.

[0124] Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component arrangement.

[0125] In some embodiments of this application, a computer device is provided, including one or more processors; a memory; and one or more application programs, where the one or more application programs are stored in the memory and are configured to be executed by the processor to perform the steps of the above data processing method. The steps of the data processing method here can be the steps in the data processing methods of the above various embodiments.

[0126] In some embodiments of this application, a computer-readable storage medium is provided, storing a computer program, and the computer program is loaded by the processor so that the processor performs the steps of the above data processing method. The steps of the data processing method here can be the steps in the data processing methods of the above various embodiments.

[0127] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0128] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0129] The above has introduced in detail a data processing method, device, computer device, and storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A data processing method, characterized in that, comprising: obtaining a query statement, the query statement including multiple fields; determining a query database, non-function information of each of the fields, and a recursive function stack according to the query statement; generating an executable statement for the query database according to the non-function information; obtaining a data set to be parsed, the data set being data returned by the query database running the executable statement; obtaining a query result according to the query statement, the recursive function stack, and the data set; The determining the query database, non-function information of each of the fields, and the recursive function stack according to the query statement includes: mapping each field in the query statement to a triple, the triple structure being (database field name or stack push automaton alias, returned alias, stack push automaton); the stack push automaton is a computing stack generated during the operation of a custom function to solve the problem of recursive calls in the custom function, and the structure of the stack push automaton is a binary tuple, the first element in the binary tuple being the function name, the second element being the input parameter group, and the function input parameter of the top element of the stack being the database queryable field name; constructing a triple set based on the triple; The generating the executable statement for the query database according to the non-function information includes: for the fields in the query statement that include functions, taking the first element of the non-stack push automaton alias in the triple for splicing; for the fields in the query statement that do not include functions, splicing the entire field; splicing query parameters other than the preset query parameter types to obtain an effective query statement; translating the effective query statement to obtain an executable statement for the query database; The obtaining the data set to be parsed, the data set being data returned by the query database running the executable statement, includes: allocating the executable statement to a query database component for running; the query database returns a data set to be parsed; The obtaining the query result according to the query statement, the recursive function stack, and the data set includes: if there is a stack push automaton in the triple, performing a stack push operation; if there is no stack push automaton in the triple but there is a returned alias, obtaining the query result according to the returned alias; if there is no stack push automaton and no returned alias in the triple, obtaining the query result according to the database field name.

2. The data processing method according to claim 1, characterized in that, the non-function information includes condition information and query information, and the determining the non-function information of each of the fields and the recursive function stack according to the query statement includes: determining a query part and a condition part according to the query statement; obtaining condition parameters according to the condition part, and setting the condition parameters as the condition information; if the query part includes a function, determining the recursive function stack according to the function part of the query part, and determining the query information according to the non-function part of the query part; if the query part does not include a function, determining the query part as the query information.

3. The data processing method according to claim 2, characterized in that, Determining the recursive function stack according to the function part of the query part includes: If the function part only contains one layer of function, determine the function name and the input parameter group of the function as the recursive function stack, and the input parameter group is the query parameter of the corresponding field of the function; If the function part contains multiple layers of functions, analyze the function call relationship of the function part and the function names and input parameter groups of each layer of functions; Determine the recursive function stack according to the function call relationship, the function names and input parameter groups of each layer of functions, and the input parameter group of the top function of the recursive function stack is the query parameter of the corresponding field of the recursive function stack.

4. The data processing method according to claim 2, wherein, The non-function information further includes parameter information. After determining the query part as the non-function information if the query part does not include a function, it includes: Obtain the preset query parameter type and the query parameters to be queried in each field; If the query parameter to be queried does not belong to the query parameter type, determine the query parameter to be queried as the parameter information.

5. The data processing method according to claim 2, wherein, Determining the query part and the condition part according to the query statement includes: Generate an abstract syntax tree according to the query statement; Traverse the abstract syntax tree to determine the condition identification points; Determine the query part and the condition part according to the condition identification points.

6. The data processing method according to claim 3, wherein, Obtaining the query result according to the query statement, the recursive function stack and the data set includes: If the target field in the query statement contains a function, determine the query parameter of the target field according to the recursive function stack corresponding to the target field and the data set; Determine the field query result corresponding to the target field according to the recursive function stack corresponding to the target field and the query parameter; If the target field in the query statement does not contain a function, determine the field query result of the target field according to the non-function information corresponding to the target field and the data set; Determine the query result according to the field query results of each field in the query statement.

7. The data processing method according to claim 1, wherein, Generating the executable statement of the query database according to the non-function information includes: Generate an effective query statement according to the non-function information of each field; Obtain the conversion syntax of the query database; Obtain the executable statement according to the effective query statement and the conversion syntax.

8. A data processing device, wherein, The device is used to execute the data processing method according to any one of claims 1-7, and the device includes: A statement acquisition module, configured to acquire a query statement, where the query statement includes multiple fields; A statement analysis module, communicatively connected to the statement acquisition module, configured to determine a query database, non-function information and a recursive function stack of each field according to the query statement; An instruction generation module, communicatively connected to the statement analysis module, for generating an executable statement for querying the database according to the non-function information; A data acquisition module, communicatively connected to the instruction generation module, for acquiring a data set to be parsed, where the data set is data returned by the database query running the executable statement; A data parsing module, communicatively connected to the statement acquisition module, the statement analysis module, and the data acquisition module, for obtaining a query result according to the query statement, the recursive function stack, and the data set.

9. A computer device, characterized in that, the computer device includes: one or more processors; a memory; and one or more applications, where the one or more applications are stored in the memory and configured to be executed by the processor to implement the data processing method according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that, a computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in the data processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Query statement generation method and device, computer equipment and storage medium

    CN114356968A