Data processing method and system, electronic equipment and readable storage medium

By converting the query language into a data query request representation and generating a query plan representation, decoupling the query language and execution logic solves the problem of low reusability of the data management system components, and efficient and flexible data processing and hardware collaboration are achieved, reducing development and maintenance costs.

CN120256450AActive Publication Date: 2025-07-04ZTE CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510743278.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The components of existing data management systems are low, resulting in high development and maintenance costs. Users need to learn a variety of incompatible structured query languages. There is a lack of unified standards between systems, which reduces work efficiency and hardware collaboration efficiency.

Method used

By converting the query language into a data query request representation, and generating a query plan representation based on the data query request representation, decoupling the query language and execution logic, and using the interface intermediate representation layer, the planned intermediate representation layer and the storage adaptation layer to achieve a modular design, supporting unified access to different execution engines and storage engines.

Benefits of technology

Improves reusability and flexibility of query execution, reduces development and maintenance costs, provides a consistent user experience, and supports efficient utilization of multiple workloads and heterogeneous hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256450A_ABST
    Figure CN120256450A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and system, electronic equipment and a readable storage medium, and belongs to the technical field of data processing. The method comprises the following steps: converting an obtained query language into a data query request representation according to at least one query clause; generating a query plan representation according to the data query request representation and the at least one execution operator; and obtaining a data processing result by executing a data processing operation corresponding to the query plan representation, the data processing operation including a data write-in operation and a data query operation. Through the mode, the query language and the execution logic can be decoupled, so that the same type of query language can be executed by using different execution logic, or the same execution logic can process different types of query languages, and the reusability of query execution is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a data processing method, system, electronic device, and readable storage medium. Background Art

[0002] Existing data management systems usually include database systems such as MySQL and MongoDB, as well as big data management systems such as Snowflake and Spark. Such systems usually use their respective SQL dialects or private API interfaces for data insertion, deletion, modification, and query. Such data management systems usually include a query engine, an execution engine, and a storage engine. Among them, the query engine is used to translate the SQL dialect into an abstract syntax tree and generate an optimized query plan, the execution engine is used for the specific execution of the query plan, and the storage engine is responsible for data access and transaction management. These several modules are tightly coupled together to form their respective data management systems.

[0003] Although most data management systems have similar components logically, existing database management systems are developed and maintained as a whole. The specific implementations of these modules are highly scattered and have little reusability, resulting in problems such as fragmentation, duplicate development, and high maintenance costs in existing data management systems. Summary of the Invention

[0004] Embodiments of this application provide a data processing method, system, electronic device, and readable storage medium to at least solve the problem of low component reusability in related data management systems.

[0005] To solve the above technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide a data processing method, including: converting the obtained query language into a data query request representation according to at least one query clause; generating a query plan representation according to the data query request representation and at least one execution operator; and obtaining a data processing result by performing a data processing operation corresponding to the query plan representation, where the data processing operation includes a data write operation and a data query operation.

[0006] In a second aspect, embodiments of this application provide a data processing system, including: an interface intermediate representation layer module for converting the obtained query language into a data query request representation according to at least one query clause; a plan intermediate representation layer module for generating a query plan representation according to the data query request representation and at least one execution operator; and a result obtaining module for obtaining a data processing result by performing a data processing operation corresponding to the query plan representation, where the data processing operation includes a data write operation and a data query operation.

[0007] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instructions that can run on the processor. When the program or instructions are executed by the processor, the steps of the method described in the first aspect above are implemented.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a program or instructions are stored. When the program or instructions are executed by a processor, the steps of the method described in the first aspect above are implemented.

[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute the steps of the method described in the first aspect above.

[0010] In the embodiment of the present application, according to at least one query clause, the obtained query language is converted into a data query request representation; according to the data query request representation and at least one execution operator, a query plan representation is generated; by executing a data processing operation corresponding to the query plan representation, a data processing result is obtained, where the data processing operation includes a data writing operation and a data query operation. In this way, by converting the query language into a data query request representation and generating a query plan representation according to the data query request representation, the query language can be decoupled from the execution logic, so that the same type of query language can be executed using different execution logics, or the same execution logic can process different types of query languages, thereby improving the reusability of query execution.

[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0013] Figure 1 Shows a schematic architecture diagram of a related data management system; Figure 2 Shows a schematic flowchart of a data processing method provided by some embodiments of the present application; Figure 3 Shows a schematic architecture diagram of a data management system provided by some embodiments of the present application; Figure 4 Shows a schematic structural diagram of a data processing system provided by some embodiments of the present application; Figure 5 Shows a schematic architecture diagram of an assembled data management system provided by some embodiments of the present application; Figure 6 Shows a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0014] Here, exemplary embodiments will be described in detail, and examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0015] As Figure 1 shown, a related data management system generally includes an interface layer 110, a query engine 120, an execution engine 130, and a storage engine 140; among them, the interface layer 110 is used to obtain a query language input by a user; the query engine 120 is used to translate the query language into an abstract syntax tree and generate a query plan; the execution engine 130 is used for the specific execution of the query plan; the storage engine 140 is used for data access and transaction management during the execution process. These several modules are tightly coupled together to form a data management system. Although most data management systems have similar components logically, existing database management systems are developed and maintained as a whole. Various database management systems usually use their own query languages or private interfaces for data addition, deletion, modification, and query operations. For example, relational databases use the standard SQL query language for data operations; NoSQL databases use their own private query languages or command interfaces to flexibly adapt to various data models; graph databases use graph query languages for operations on graph data. The query language or interface of each database management system is optimized for its data model and architecture. Therefore, their data addition, deletion, modification, and query operations are usually not interoperable. This fragmented situation has at least the following problems: 1) Developers need to reinvent the wheel between different systems, increasing the development and maintenance costs; 2) There is a lack of a unified standard between different data management systems, resulting in users needing to learn and adapt to multiple incompatible structured query languages (SQL) and non-SQL dialects, thus increasing the cognitive burden and learning costs of users; 3) The functions and semantics between different systems are inconsistent. When users use multiple systems, they need to continuously adapt to different behaviors and characteristics, reducing work efficiency.

[0016] Since each system has its unique Application Programming Interface (API) and characteristics, users need to spend a lot of time learning and mastering the usage methods of different systems. At the same time, due to the lack of reusability between the various modules of different systems, the development of new systems needs to start from scratch, resulting in an overly long development cycle and making it difficult to quickly launch new features and improvements. To quickly launch a prototype, developers often sacrifice the stability and maintainability of the code, leading to the continuous accumulation of technical debt. Moreover, due to the high fragmentation of data management systems, it is difficult for hardware vendors to optimize for specific data processing requirements, resulting in low coordination efficiency between hardware and software.

[0017] In view of the problems existing in the above data processing process, an embodiment of the present application provides a data processing method. By converting a query language into a data query request representation and generating a query plan representation according to the data query request representation, this method can decouple the query language from the execution logic, enabling the same type of query language to be executed using different execution logics, or the same execution logic to handle different types of query languages, thereby improving the reusability of query execution.

[0018] Please refer to Figure 2 , Figure 2 , which shows a schematic flowchart of the data processing method provided by some embodiments of the present application. The execution subject of this method can be a terminal device or a server. Among them, the terminal device can be a device such as a personal computer, or a mobile terminal device such as a mobile phone or a tablet computer. The terminal device can be the terminal device used by the user. The server can be an independent server or a server cluster composed of multiple servers. Moreover, the server can be the background server of a certain service or the background server of a certain platform (such as a data management system, a data processing platform, a query system, etc.). In the embodiments of the present application, the case where the execution subject is a server is taken as an example for description. For the case of the terminal device, it can be processed according to the following relevant content and will not be elaborated here. As shown in the figure, the data processing method 200 can include the following steps: Step 201: Convert the obtained query language into a data query request representation according to at least one query clause.

[0019] Among them, for the query language in the above step 201, it includes Structured Query Language (SQL), Graph Query Language (GraphQL), and natural language; for the query clauses in the above step 201, it includes the select_fields clause (SELECT statement) for indicating the fields returned by the query, the from_tables clause (FROM statement) for indicating the tables from which the data is sourced, the join_conditions clause (JOIN statement) for indicating the table join conditions, the where_conditions clause (WHERE statement) for indicating the filtering conditions of the query, the group_by clause (GROUP BY statement) for indicating the grouping fields, the having_conditions clause (HAVING statement) for indicating the filtering fields after grouping, the order_by clause (ORDER BY statement) for indicating the sorting rules, the limit clause (LIMIT statement) for indicating the limit on the number of rows returned by the query, the offset clause (OFFSET statement) for indicating the query offset, the aggregations clause (such as SUM, COUNT, AVG) for indicating the aggregation operations, the distinct clause for indicating whether to remove duplicates, and the subqueries clause for indicating subqueries.

[0020] In a specific implementation, as Figure 3 shown, the interface layer 310 of the data management system may include various forms of external interfaces, including SQL, GraphQL, Pandas, REST, natural language, etc. interfaces. Through the interface layer 310, various types of query languages can be obtained. For example, if the natural language input by the user is: "Find 10 books on natural language processing, sorted in descending order by publication time", if the natural language interface is used for query, the natural language can be converted into an interface Intermediate Representation (IR), and a trained model is used for the conversion. The format of the conversion prompt (Prompt) is {Convert "Find 10 books on natural language processing, sorted in descending order by publication time" to the interface IR format}; if the SQL interface is used for query, the SQL statement is as follows: SQL statement: SELECT title, author, publish_date FROM books WHERE topic = 'natural language processing' ORDER BY publish_date DESC LIMIT 10; According to at least one preset query clause, such as select_fields, from_tables, where_conditions, etc., the obtained query languages such as natural language and SQL are transformed into a more abstract intermediate representation, which mainly includes the basic information of the query, such as select fields, where conditions, orderBy sorting, and limit restrictions, etc.

[0021] In this way, through at least one query clause, the query language input by the user can be converted into a data query request representation that can be understood by the machine, so as to facilitate the subsequent execution of the data processing operation corresponding to the query language.

[0022] Step 202: Generate a query plan representation according to the data query request representation and at least one execution operator.

[0023] Among them, the execution operators in the above step 202 include the operator Scan for indicating a scan operation, the operator Filter for indicating a filtering operation, the operator Projection for indicating a projection operation, the operator Join for indicating a join operation, the operator Aggregation for indicating an aggregation operation, the operator Sort for indicating a sorting operation, the operator Limit for indicating a limit operation, and the operator Union for indicating a union operation.

[0024] In a specific implementation, according to the data query request representation converted in the above step 201 and execution operators such as Scan, Filter, and Projection, a query plan representation is generated. The query plan representation contains specific operators, such as Scan, Sort, Limit, etc., and these operators describe the execution order of the query.

[0025] In this way, by further converting the data query request representation into a query plan representation, the assembled hierarchical design in the query plan representation can ensure the flexibility and scalability of the query, while ensuring the maximization of execution efficiency.

[0026] Step 203: Obtain a data processing result by executing a data processing operation corresponding to the query plan representation, where the data processing operation includes a data writing operation and a data query operation.

[0027] In a specific implementation, an execution engine can be called to execute a data processing operation corresponding to the query plan representation to obtain the data processing result output by the execution engine. Among them, such as Figure 3As shown in the figure, the execution engine 330 may include Velox, Spark, Ray, Postgre, Flink, etc. Here, the application scenarios of different execution engines are different. For example, Velox focuses on the query engine and is suitable for large-scale data analysis; Spark has a powerful big data analysis framework and is suitable for batch processing and stream processing; Ray focuses on distributed computing, especially machine learning and deep learning tasks; PostgreSQL is a powerful relational database and is suitable for complex database applications; Flink focuses on stream processing and is suitable for real-time data computing and event-driven applications. Since the embodiments of the present application convert the query language input by the user into a unified query plan representation through query clauses and execution operators, this unified query plan representation can be adapted to multiple execution engines.

[0028] Through the above steps, converting the query language into a data query request representation and generating a query plan representation according to the data query request representation can decouple the query language from the execution logic, enabling the same type of query language to be executed using different execution logics, or the same execution logic to handle different types of query languages, thereby improving the reusability of query execution.

[0029] In some embodiments, in step 201 above, converting the obtained query language into a data query request representation according to at least one query clause includes: Obtain the query language; generate a structured representation corresponding to the query language by parsing the query language; map the structured representation to at least one query clause to generate a data query request representation.

[0030] In a specific implementation, obtain the query language input by the user through the interface layer 310 of the data management system. The interface layer 310 includes interfaces such as SQL, GraphQL, and natural language; by parsing the query language, represent the query language in a structured manner. For example, represent the query language as an expression tree, including function calls, table references, constants, and various operator operations such as filtering, projection, sorting, joining, aggregation, window functions, shuffling / repartitioning, etc.; then, map the structured representation to query clauses such as select_fields, from_tables, where_conditions, etc. to generate a data query request representation, which includes the basic information of the query, such as select fields, where conditions, orderBy sorting, and limit restrictions.

[0031] In some possible implementation manners, the above-mentioned generating a structured representation corresponding to the query language by parsing the query language includes: Obtain multiple clauses in the query language; convert the query elements in each clause into an intermediate representation of the interface; according to the dependency relationships between the multiple clauses, reorganize the intermediate representation of the interface to generate a structured representation corresponding to the query language.

[0032] In a specific implementation, the query language includes multiple clauses, such as SELECT, FROM, WHERE, JOIN, etc. Obtain multiple clauses in the query language, convert query elements such as fields, aggregate functions, and expressions in each clause into an intermediate representation IR of the interface, and according to the dependency relationships between the multiple clauses, reorganize the interface IR to generate a structured representation corresponding to the query language. Among them, the mapping method from the query language to the interface IR is shown in the following table: Table 1. Correspondence table of the mapping method from the query language to the interface IR

[0033] In some embodiments, in step 202 above, a query plan representation is generated according to the data query request representation and at least one execution operator, including: Generate a query logical plan according to the target query clause in the data query request representation; wherein, the target query clause includes a first query clause for indicating the query intention, a second query clause for indicating the context information, and a third query clause for indicating the query structure; generate an executable physical plan by optimizing the execution efficiency of the query logical plan; convert the executable physical plan into a query plan representation according to at least one execution operator.

[0034] In a specific implementation, such as Figure 3As shown in the figure, the query engine 320 includes Calcite, Orca, Presto, Postgre, Flink, etc. The query engine receives the data query request representation in step 201 above. First, it extracts from the data query request representation a first query clause for indicating the query intent, a second query clause for indicating the context information, and a third query clause for indicating the query structure, such as select_fields, join_conditions, where_conditions, order_by, etc., and generates a query logical plan; based on the time cost and preset rules, it optimizes the execution efficiency of the query logical plan to generate an executable physical plan; furthermore, according to execution operators such as Scan, Sort, Limit, etc., it converts the executable physical plan into a query plan representation, that is, the query plan IR. The query plan IR describes the execution manner and optimization strategy of the query. The plan IR consists of a series of operators, and these operators represent the execution steps of the database query. The plan IR is defined using json and includes the following content: root: The root operator of the query (the final output) operators: Each operator involved in the query execution type: Operator type (such as Scan, Filter, Sort, Limit, etc.) input: The upstream operator that this operator depends on (i.e., the data source) output_fields: The fields output by this operator conditions: Conditions such as filtering and joining order_by: Sorting method limit: The row limit of the query return cost: The estimated execution cost of the operator (optional) parallelism: The parallelism of the operator (optional) Among them, the operator types mainly include:

[0035] In some embodiments, in step 203 above, by performing data processing operations corresponding to the query plan representation, the data processing result is obtained, including: Obtain at least one operator node in the query plan representation, and assign an execution task to each operator node; select a target storage engine that matches the execution task from a preset plurality of storage engines, and adapt the execution task to an operation request of the target storage engine; by calling the target storage engine to execute the data processing operation corresponding to the operation request, the data processing result is obtained.

[0036] In specific implementations, such as Figure 3 shown, the execution engine 330 can have different choices according to different scenarios, including Velox, spark, Ray, Postgre, Flink, etc. The execution engine 330 is the core component in the data management system responsible for actually executing the query plan and returning the results. It first parses the operator nodes from the query plan IR and assigns execution tasks to each operator. Optionally, the execution engine 330 can manage the execution order and dependency relationships of tasks through a task scheduler, divide the tasks into multiple subtasks and allocate them to different computing nodes or threads for parallel processing; each task sequentially executes operators such as scan, filter, join, sort, aggregate, limit, etc., and generates intermediate results by reading data from the storage engine and gradually transforming and processing; the execution engine 330 is used to manage memory and intermediate results, and for large-scale data processing, the intermediate results may be written to disk cache to avoid memory overflow.

[0037] In specific applications, in order to improve efficiency, the above-mentioned execution engine 330 can adopt vectorized execution technology, process batch data operations as a unit, and can combine strategies such as JIT compilation, partition parallelism, and pipeline optimization for performance optimization. During the execution process, the execution engine 330 can also detect and handle runtime errors, and provide a fault tolerance mechanism for automatic retry or degradation processing. Finally, the execution engine 330 merges and formats the intermediate results of all subtasks to obtain the data processing result, and returns the data processing result to the user or the upper-layer application in the output format specified by the query plan IR.

[0038] In this way, it can ensure the efficient execution and stability of the query, and at the same time has good scalability and flexibility. By extending the execution engine, it supports partially or fully distributing execution tasks (such as projection, aggregation, sorting, encoding and decoding, etc.) to heterogeneous hardware such as GPUs, FPGAs, and DPUs, etc., and can easily support new heterogeneous hardware.

[0039] Among them, such as Figure 3As shown, the storage engine 340 can include different components according to different scenarios of executing tasks, such as DuckDB, RocksDB, SQLite, LanceDB, Parquet, etc. It is the core component in the data management system responsible for the persistent storage and efficient access of data. The storage engine 340 supports multiple storage structures, such as row storage, column storage, and key-value storage, and provides corresponding optimization strategies according to the characteristics of different storage engines. Furthermore, a storage adaptation layer is added between the execution engine 330 and the storage engine 340. This storage adaptation layer is used to select a target storage engine that matches the execution task from a preset plurality of storage engines, and adapt the execution task into an operation request of the target storage engine. The storage engine 340 obtains a data processing result by calling the target storage engine to execute a data processing operation corresponding to the operation request. Optionally, the storage engine 340 can also improve data access performance through index structures, batch operations, and data compression technologies, and ensure the ability of data persistence and fault recovery. For example, data reliability is achieved through Write-Ahead Logging (WAL), transaction logs, snapshots, and checkpoint mechanisms. To support multi-user concurrent access, the storage engine 340 can also provide a lock mechanism and Multi-Version Concurrency Control (MVCC) to ensure the isolation and consistency of transactions. In this way, not only is efficient data storage and management capabilities provided, but also flexible replacement and expansion of the storage engine are supported, ensuring high performance and scalability of the system in different application scenarios.

[0040] In some possible implementation manners, the above-mentioned selecting a target storage engine that matches the execution task from a preset plurality of storage engines includes: Determining a task scenario that matches each storage engine according to the characteristic information of the preset plurality of storage engines; and selecting a target storage engine that matches the scenario where the execution task is located from the plurality of storage engines according to the matching relationship between the plurality of storage engines and the task scenarios.

[0041] In a specific implementation, the storage adaptation layer is an abstraction layer between the execution engine and the underlying storage engine. Its main function is to provide a unified data access interface for the execution engine, thereby supporting seamless integration of different types of storage engines (e.g., DuckDB, SQLite, RocksDB, LanceDB, etc.). The core responsibilities of the storage adaptation layer include: storage engine abstraction, data format conversion, interface standardization, optimization strategy adaptation, and resource management. First, the storage adaptation layer defines standardized storage engine interfaces, such as methods like CreateTable, DropTable, Scan, IndexScan, Read, Write, Update, Delete, etc., and provides specific adapter implementations for each storage engine. Second, it is responsible for adapting the requests of the execution engine to the data formats and APIs of the underlying storage engines, including data encoding and decoding, metadata management, and table structure mapping. The storage adaptation layer also provides an optimization strategy adaptation function, allowing the execution engine to utilize the characteristic information of the storage engine to determine the task scenarios that match each storage engine. For example, RocksDB is suitable for key-value indexing, DuckDB is suitable for column storage, and SQLite is suitable for transaction support. Furthermore, based on the matching relationships between multiple storage engines and task scenarios, the target storage engine that matches the scenario where the task is executed can be selected from multiple storage engines to improve query performance.

[0042] Optionally, to support concurrent execution and efficient data access, the storage adaptation layer can also provide functions such as a caching mechanism, parallel I / O scheduling, and memory management, and handle exceptions and errors that may be thrown by the underlying storage engine. In this way, the execution engine can uniformly access and control different storage engines, thereby achieving efficient and flexible storage engine adaptation and integration.

[0043] By combining different components, the data management requirements of different scenarios can be met. For example, in a transaction processing scenario, SQL is used as the interface, PostgreSQL is used as the query and execution engine, and SQLite is used as the storage engine; for a data analysis scenario, GraphQL is used as the interface, Orca or Velox is used as the query and execution engine respectively, and DuckDB is used as the storage engine; in a large model application scenario, natural language and Pandas are used as the interface, Calcite or Ray is used as the query and execution engine respectively, and Parquet is used as the storage engine, as shown in the following table: Table 2. Combination Relationship Table of Different Components in the Data Management System

[0044] In this way, through the combination of different components, the expansion of the data management software can be achieved, thereby meeting the requirements of different scenarios.

[0045] Figure 4 FIG. shows a schematic structural diagram of a data processing system provided by some embodiments of the present application. The data processing system can implement all or part of the content in the embodiments as Figure 2 shown. The data processing system 400 includes: An interface intermediate representation layer module 410, configured to convert the obtained query language into a data query request representation according to at least one query clause; A plan intermediate representation layer module 420, configured to generate a query plan representation according to the data query request representation and at least one execution operator; A result acquisition module 430, configured to obtain a data processing result by executing a data processing operation corresponding to the query plan representation, where the data processing operation includes a data write operation and a data query operation.

[0046] In some embodiments, when the interface intermediate representation layer module 410 is configured to convert the obtained query language into a data query request representation according to at least one query clause, it is specifically configured to: Obtain the query language; Generate a structured representation corresponding to the query language by parsing the query language; Map the structured representation to at least one query clause to generate a data query request representation.

[0047] In some possible implementation manners, when the interface intermediate representation layer module 410 is configured to generate a structured representation corresponding to the query language by parsing the query language, it is specifically configured to: Obtain multiple clauses in the query language; Convert the query elements in each clause into an interface intermediate representation; Recombine the interface intermediate representation according to the dependency relationship between the multiple clauses to generate a structured representation corresponding to the query language.

[0048] In some embodiments, when the plan intermediate representation layer module 420 is configured to generate a query plan representation according to the data query request representation and at least one execution operator, it is specifically configured to: Generate a query plan representation according to the data query request representation and at least one execution operator, including: Generate a query logical plan according to the target query clause in the data query request; wherein, the target query clause includes a first query clause for indicating the query intention, a second query clause for indicating context information, and a third query clause for indicating the query structure; Generate an executable physical plan by optimizing the execution efficiency of the query logical plan; Convert the executable physical plan into a query plan representation according to at least one execution operator.

[0049] In some embodiments, the result acquisition module 430 includes: An execution engine layer module, configured to obtain at least one operator node in the query plan representation and assign an execution task to each operator node; A storage adaptation layer module, configured to select a target storage engine that matches the execution task from a plurality of preset storage engines and adapt the execution task to an operation request of the target storage engine; A storage engine layer module, configured to perform a data processing operation corresponding to the operation request by calling the target storage engine to obtain a data processing result.

[0050] In an exemplary embodiment, as Figure 5 shown, the embodiment of the present application further provides an assembled data management system, and the assembled data management system includes: An interface layer 510, configured to obtain a query language input by a user; An interface intermediate representation 520, configured to convert the obtained query language into a data query request representation according to at least one query clause; A query engine 530, configured to generate a query plan representation according to the data query request representation and at least one execution operator; A plan intermediate representation 540, configured to obtain the query plan representation generated by the query engine 530; An execution engine 550, configured to obtain at least one operator node in the query plan and assign an execution task to each operator node; A storage adaptation layer 560, configured to select a target storage engine that matches the execution task from a plurality of preset storage engines and adapt the execution task to an operation request of the target storage engine; A storage engine layer 570, configured to perform a data processing operation corresponding to the operation request by calling the target storage engine to obtain a data processing result.

[0051] In the embodiments of the present application, the data management system is decomposed into a series of reusable components, including an interface layer, an interface intermediate representation (IR), a query engine, a plan intermediate representation (IR), an execution engine, a storage adapter, and a storage engine. These components interact through clearly defined interfaces (APIs), thereby achieving system modularization and decoupling. This modular design not only improves development efficiency but also reduces the system's maintenance cost, while providing users with a more consistent experience. Through the unified interface IR and plan IR, different system interfaces can generate a unified interface IR, which is then optimized into a unified plan IR by the query engine, and different execution engines can execute these plan IRs. This decoupling enables the system interfaces, query engine, execution engine, and storage engine to develop independently, while also supporting cross-system query optimization and execution. This architecture not only supports various workloads, such as from online transaction processing to online analytical processing, from stream processing to machine learning, but also allows developers to select and combine different components according to their needs, thus quickly building a data management system that meets specific requirements. Through componentization and standardization, the data management system will be able to better adapt to the rapidly changing technical environment, adapt to new hardware accelerators such as GPUs and FPGAs, promote the co-evolution between hardware and software, and fully utilize the functional and performance advantages of new hardware.

[0052] The PostgreSQL database extends its functions through a plugin mechanism (Extension), allowing users to flexibly add new features and capabilities without modifying the core code of the database. Plugins can implement data type extensions. For example, by installing PostGIS, geospatial data can be supported, or fuzzy search functions can be provided through pg_trgm. Plugins also support the definition of custom index methods, functions, and operators to optimize query performance. For example, the btree_gin plugin enhances the B-tree index. However, this plugin mechanism usually depends on the internal extensions and external function packages of the database. The extension ability is often tightly coupled to specific database versions and architectures and cannot be cross-platform shared and migrated between different database systems, bringing challenges in terms of compatibility, maintenance, and performance. Compared with the way of database extension through the plugin mechanism in PostgreSQL, the embodiments of the present application perform extensions through the interface IR, plan IR, and storage adapter layer. The interface IR and plan IR structurally separate the query semantics from the execution plan, enabling the extension logic to be shared between different database systems, with better flexibility, maintainability, and heterogeneous system compatibility, and is more suitable for building a modular, multimodal, and sustainable evolving data processing platform.

[0053] Figure 6The figure shows a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Referring to this figure, at the hardware level, the electronic device 600 includes a processor 610. Optionally, it includes an internal bus 620, a network interface 630, and a memory. Among them, the memory may include a memory 641, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory 642 (non-volatile memory), such as at least one disk memory, etc. Of course, the electronic device 600 may also include other hardware required for other services.

[0054] The processor 610, the network interface 630, and the memory can be interconnected through the internal bus 620. The internal bus 620 can be an Advanced Microcontroller Bus Architecture bus, a Wishbone bus, an Open Core Protocol (OCP) bus, an Avalon bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a bidirectional arrow is used in this figure, but it does not mean that there is only one bus or one type of bus.

[0055] The memory stores programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include a memory 641 and a non-volatile memory 642, and provide instructions and data to the processor 610.

[0056] The processor 610 reads the corresponding computer program from the non-volatile memory 642 into the memory and then runs it, forming a device for locating the target user at the logical level. The processor 610 executes the program stored in the memory and specifically executes: Figure 2 The method disclosed in the illustrated embodiment and realizes the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be elaborated here.

[0057] The above is as described in the present application Figure 2The method disclosed in the illustrated embodiment can be applied to or implemented by the processor 610. The processor 610 may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in the processor 610. The above-mentioned processor 610 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0058] This computer device can also execute the various methods described in the foregoing method embodiments and achieve the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be elaborated here.

[0059] Of course, in addition to the software implementation, the electronic device 600 of the present application does not exclude other implementation manners, such as a logic device or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and may also be hardware or a logic device.

[0060] The embodiments of the present application also propose a computer-readable storage medium. The computer-readable medium stores one or more programs. When the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device is caused to execute Figure 2 the method disclosed in the illustrated embodiment and achieve the functions and beneficial effects of the various methods described in the foregoing method embodiments, which will not be elaborated here.

[0061] Among them, the computer-readable storage medium includes a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, an optical disc, etc.

[0062] Furthermore, the embodiment of the present application also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the following process is implemented: Figure 2 The method disclosed in the illustrated embodiment realizes the functions and beneficial effects of each method described in the foregoing method embodiments, and will not be elaborated herein.

[0063] The embodiment of the present application can be applied to various scenarios of electronic device cooperation or interconnection, including: cooperation and interconnection between a mobile phone and a laptop / tablet computer; cooperation and interconnection between a mobile terminal and a smart TV / display; cooperation and interconnection between a mobile phone or a tablet computer and an in-vehicle entertainment system; cooperation and interconnection between a mobile terminal and a smart conference system, etc. Thus, it meets the diverse scenario requirements of users in smart home, smart office, smart travel, etc.

[0064] In summary, the above are only the preferred embodiments of the present application, and do not limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0065] The system, device, module or unit illustrated in the above embodiment can be specifically implemented by a computer chip or entity, or by a product with a certain function. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0066] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0067] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0068] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

Claims

1. A data processing method, characterized in that, including: converting the obtained query language into a data query request representation according to at least one query clause; generating a query plan representation according to the data query request representation and at least one execution operator; obtaining a data processing result by performing a data processing operation corresponding to the query plan representation, wherein the data processing operation includes a data writing operation and a data query operation.

2. The method according to claim 1, wherein The converting the obtained query language into a data query request representation according to at least one query clause includes: obtaining a query language; generating a structured representation corresponding to the query language by parsing the query language; mapping the structured representation to at least one query clause to generate a data query request representation.

3. The method according to claim 2, wherein The generating a structured representation corresponding to the query language by parsing the query language includes: obtaining multiple clauses in the query language; converting query elements in each clause into an interface intermediate representation; reorganizing the interface intermediate representation according to the dependency relationship between the multiple clauses to generate a structured representation corresponding to the query language.

4. The method according to claim 1, wherein The generating a query plan representation according to the data query request representation and at least one execution operator includes: generating a query logical plan according to a target query clause in the data query request representation; wherein the target query clause includes a first query clause for indicating a query intention, a second query clause for indicating context information, and a third query clause for indicating a query structure; generating an executable physical plan by optimizing the execution efficiency of the query logical plan; converting the executable physical plan into a query plan representation according to at least one execution operator.

5. The method according to any one of claims 1 to 4, characterized in that, The obtaining a data processing result by performing a data processing operation corresponding to the query plan representation includes: obtaining at least one operator node in the query plan representation and assigning an execution task to each operator node; selecting a target storage engine matching the execution task from a plurality of preset storage engines and adapting the execution task to an operation request of the target storage engine; obtaining a data processing result by calling the target storage engine to perform a data processing operation corresponding to the operation request.

6. The method according to claim 5, wherein The selecting a target storage engine matching the execution task from a plurality of preset storage engines includes: determining a task scenario matching each storage engine according to characteristic information of a plurality of preset storage engines; selecting a target storage engine matching the scenario where the execution task is located from the plurality of storage engines according to the matching relationship between the plurality of storage engines and the task scenarios.

7. A data processing system, characterized in that, including: an interface intermediate representation layer module for converting the obtained query language into a data query request representation according to at least one query clause; a plan intermediate representation layer module for generating a query plan representation according to the data query request representation and at least one execution operator; a result obtaining module for obtaining a data processing result by performing a data processing operation corresponding to the query plan representation, wherein the data processing operation includes a data writing operation and a data query operation.

8. The system according to claim 7, wherein The result acquisition module includes: An execution engine layer module, configured to obtain at least one operator node in the query plan representation and allocate an execution task to each of the operator nodes; A storage adaptation layer module, configured to select a target storage engine that matches the execution task from a plurality of preset storage engines and adapt the execution task to an operation request of the target storage engine; A storage engine layer module, configured to execute a data processing operation corresponding to the operation request by invoking the target storage engine to obtain a data processing result.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory, and the memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium, characterized in that, A program or instruction is stored on the computer-readable storage medium. When the program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Hybrid query optimization method and device based on big data

    CN111221860A

  • Data query method and query engine

    CN117785910A

  • Unified query engine for graphics and relational data

    CN118210951A

  • Method of converting query plans to native code

    US20140280030A1

  • Multi-language fusion query method and multi-model database system

    US20220075780A1