Pandas-based cross-database query method, system and device and medium
Through the cross-database query method based on pandas, the problems of high complexity, insufficient real-time and performance development, lack of automation support for data consistency and type conversion in cross-database query technology are solved, and an efficient, simple and maintainable cross-database query solution is realized, ensuring data consistency and integrity.
Patent Information
- Application Number
- CN202510250664.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-03
AI Technical Summary
The current cross-database query technology has problems such as high development complexity, insufficient real-time and performance, and lack of automation support for data consistency and type conversion, especially in heterogeneous database environments.
Using a cross-database query method based on pandas, we can configure database connection information, parse and decompose cross-database query requests, and use pandas' data processing capabilities and flexible interfaces to execute queries and integrate results to ensure data consistency and integrity.
It realizes an efficient, simple and maintainable cross-database query solution, improves query performance, reduces development difficulty and cost, and ensures data consistency and integrity.
Smart Images

Figure CN120086274A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and relates to a cross-database query method, system, device, and medium based on pandas. Background Art
[0002] Currently, in today's data processing field, the complexity of data storage and management has increased significantly with the in-depth development of enterprise digital transformation. Data is no longer limited to a single database or homogeneous system, but is widely distributed in various types of database management systems (DBMS). These systems include both traditional relational databases (such as MySQL, PostgreSQL, Oracle, etc.), non-relational databases (such as MongoDB, Cassandra, Redis, etc.), as well as emerging time-series databases (such as InfluxDB), graph databases (such as Neo4j), etc. For example, in the financial industry, transaction data may be stored in Oracle for high-concurrency transaction processing, while user behavior logs may be stored in an unstructured manner through MongoDB; in the Internet of Things field, a large amount of time-series data generated by devices is usually managed by InfluxDB, while business metadata may rely on PostgreSQL. This heterogeneity of data storage, although meeting the performance and flexibility requirements in different scenarios, has brought huge challenges to cross-database data integration and joint query.
[0003] The popularization of the microservices architecture has further exacerbated the complexity of this problem. In the microservices mode, business systems are split into multiple independent services, and each service may use different database technologies to meet its specific needs. Taking an e-commerce platform as an example, the order service may use MySQL to ensure transaction consistency, the product recommendation service relies on Elasticsearch to achieve efficient search, and the user behavior analysis service stores semi-structured logs through MongoDB. Although this architecture improves the scalability and maintainability of the system, when facing cross-service data query requirements (such as generating user portraits that require combining order data, browsing records, and social behavior data), developers have to face problems such as multi-database syntax differences, incompatible connection methods, and data consistency guarantee. As a tool for quickly building enterprise applications, the core advantage of the low-code platform is to reduce the development threshold through visual configuration. However, when it comes to cross-database queries, it still relies on complex underlying script writing, which contradicts its "low-code" concept.
[0004] Currently, the implementation of cross-database queries mainly relies on the following two types of solutions: one is offline batch processing based on ETL (Extract-Transform-Load) tools, such as using Apache NiFi or Talend to regularly extract data into a unified data warehouse and then perform queries; the other is to utilize federated query technology (Federated Query), such as using Presto or Apache Calcite to achieve cross-database queries in a virtual data layer. However, these solutions all have significant limitations. ETL tools need to pre-define data flow rules, making it difficult to support scenarios with high real-time requirements, and redundant data storage increases storage costs and management burdens. Although federated query technology can achieve "query on demand", its performance is limited by network latency and differences in optimizer of heterogeneous databases. Especially when dealing with complex joins or aggregation operations, the response time may increase exponentially. In addition, federated query systems usually require custom adapters for different databases, with high maintenance costs and limited support for non-relational databases.
[0005] From the perspective of technical implementation details, the differences in syntax and interfaces of different databases are one of the core factors hindering efficient cross-database queries. Taking SQL dialects as an example, the `LIMIT` clause in MySQL needs to be replaced with `FETCH FIRST n ROWS ONLY` in PostgreSQL, while the query language of MongoDB is completely based on the JSON format and cannot be directly compatible with SQL. In terms of connection methods, relational databases mostly achieve programmatic access through JDBC or ODBC drivers, while non-relational databases rely on exclusive client libraries (such as `pymongo` for MongoDB). This fragmented technical ecosystem requires developers to master the programming interfaces of multiple databases and write a large amount of glue code to coordinate the interaction of different systems. For example, for an application that needs to jointly query data from MySQL and MongoDB, developers not only need to use the `pymysql` and `pymongo` libraries to establish connections respectively, but also need to manually handle the data type conversion of the two result sets (such as converting the BSON date in MongoDB to a `datetime` object in Python) and resolve field naming conflicts (such as `user_id` in MySQL and `uid` in MongoDB referring to the same entity). This process is not only time-consuming and laborious, but also prone to data misalignment or type conversion errors due to human negligence.
[0006] Another pain point of existing solutions lies in the balance between data consistency and performance. In a distributed system, the existence of the CAP theorem (consistency, availability, partition tolerance) makes it extremely difficult to implement cross-database transactions. For example, inventory deduction (stored in MySQL) and order creation (stored in MongoDB) on an e-commerce platform need to ensure atomicity, but the lack of cross-database transaction support may lead to overselling or data inconsistency. In addition, even for non-transactional queries, due to differences in lock mechanisms and isolation levels of different databases, there may be problems such as dirty reads or phantom reads in the query results. In terms of performance, cross-database queries often involve multiple network round-trips and data serialization / deserialization operations. Especially when dealing with large-scale datasets, the consumption of memory and computing resources may become a bottleneck. According to a Gartner report in 2022, more than 60% of enterprises failed to achieve the expected goals due to performance issues when implementing cross-database query projects.
[0007] Although the Pandas library has become one of the core tools in the field of data science and performs excellently in cleaning, transforming, and analyzing single-source data (such as CSV files, single database tables), its application in cross-database scenarios is still in the exploratory stage. The core advantage of Pandas lies in its efficient DataFrame structure and rich built-in functions (such as `merge`, `groupby`), which can simplify the data operation process. However, directly using Pandas to implement cross-database queries faces the following obstacles: First, the database connections natively supported by Pandas rely on functions such as `read_sql`, which require pre-configuring database drivers and lack direct support for non-relational databases; Second, the data types returned by different databases (such as `DECIMAL` in MySQL and `NUMERIC` in PostgreSQL) may not be automatically mapped to the `dtype` of Pandas and need to be processed additionally; Third, cross-database association operations (such as JOIN) require developers to manually coordinate the merging logic of multiple DataFrames, making it difficult to automate complex join conditions. Therefore, currently, Pandas is more used for offline data analysis rather than real-time scenarios of online cross-database queries.
[0008] Industry practice shows that enterprises' demand for cross-database queries is growing at an annual rate of 23% (IDC, 2023), but the limitations of the existing technology stack have led to extended development cycles and rising costs. For example, a leading e-commerce platform once disclosed that its data middle-end team needs to invest 30% of development resources to maintain cross-database query scripts, and the query performance is more than 40% lower than that of single-database operations. At the same time, emerging technologies such as Data Lake and Data Mesh attempt to improve data integration problems through unified storage architecture or decentralized governance, but these solutions are still in the proof-of-concept stage and have limited support for real-time queries. In this context, there is an urgent need for a cross-database query solution that is both efficient, compatible and easy to use to fill the gaps in existing technologies and help enterprises unleash the potential value of multi-source data.
[0009] In summary, the current technical challenges in the cross-database query field can be summarized as follows: differences in syntax and connections between heterogeneous databases lead to high development complexity, existing tools are insufficient in real-time and performance, and data consistency and type conversion lack automated support. The potential of the Pandas library has not yet been fully explored in this scenario, and how to combine its powerful data processing capabilities with cross-database query needs has become a technical direction worth exploring.
[0010] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0011] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.
[0012] The embodiments of the present disclosure provide a pandas-based cross-database query method, system, device and medium to overcome the shortcomings of existing cross-database query methods, and utilize pandas' data processing capabilities and flexible interfaces to implement an efficient, simple and maintainable cross-database query solution while ensuring data consistency and integrity.
[0013] In some embodiments, the method comprises:
[0014] Configure connection information for the database and store the connection information in a configuration file or a database; the connection information includes database type, address, port, user name and password;
[0015] Receive a cross-database query request from a user or a system, parse the request and decompose it into sub-query requests for different databases; the request includes database identifiers involved, query conditions, table and field information, table connection methods, and aggregation, sorting, or grouping operation information;
[0016] According to the sub-query requests and the corresponding database connection information, establish connections with different databases respectively, use the read_sql function of pandas or custom functions to execute the queries, and store the results as DataFrame objects;
[0017] Handle issues of inconsistent data types and field names through custom mapping rules or pandas' data conversion functions, and use pandas' merge or join functions to associate multiple DataFrame objects according to the table connection method in the request to generate the final result DataFrame;
[0018] Output the integrated result DataFrame in a specified format, including storing it as a file, writing it to a database, or converting it to JSON format.
[0019] Preferably, for relational databases, use pymysql or psycopg2 drivers, and for non-relational databases, use the corresponding driver programs.
[0020] Preferably, the custom mapping rules include field name mapping, data type coercion, and null value filling rules.
[0021] Preferably, the specified format includes CSV, Excel, or writing to the target database through pandas' to_sql function.
[0022] A cross-database query system based on pandas includes the following modules:
[0023] Configuration module: used to store and manage connection information of different types of databases, supporting dynamic update and maintenance;
[0024] Request processing module: used to receive and parse cross-database query requests, split them into sub-query requests and distribute them;
[0025] Query execution module: contains multiple query execution units adapted to different database types, used to establish database connections, execute sub-queries, and store the results as DataFrame objects;
[0026] Data integration module: used to perform data cleaning, type conversion, and association operations on multiple DataFrame objects to generate the integrated result DataFrame;
[0027] Result output module: used to output the integration result in the format specified by the user or store it at the target location.
[0028] Preferably, the query execution module supports query adaptation for relational databases and non-relational databases, and realizes compatibility by dynamically loading the corresponding database drivers.
[0029] Preferably, the data integration module automatically corrects field name conflicts and data type differences through preset mapping rules.
[0030] Preferably, the result output module supports converting the result DataFrame into a JSON array and returning it through an API interface.
[0031] In some embodiments, the apparatus includes: a processor and a memory storing program instructions, wherein the processor is configured to execute the cross-database query method based on pandas when running the program instructions.
[0032] In some embodiments, the computer-readable storage medium, wherein a computer program is stored thereon, and when the program is executed by a processor, the cross-database query method based on pandas is implemented.
[0033] A cross-database query method, system, apparatus, and medium based on pandas provided by the embodiments of the present disclosure can achieve the following technical effects:
[0034] High efficiency: Utilize the efficient data processing ability of pandas to quickly process and integrate query results, reduce the conversion time of data between different formats, and improve the overall performance of cross-database queries.
[0035] Flexibility: Support multiple database types. By configuring different database connection information, data can be conveniently queried from different types of databases without being restricted by the database types.
[0036] Ease of use: Developers do not need to write complex multi-database operation codes. Only by simple configuration and request submission, cross-database queries can be achieved, reducing the development difficulty and cost.
[0037] Scalability: New database types can be conveniently added. Only by adding the corresponding connection drivers and query execution units for the new database types, the system can be easily extended and maintained.
[0038] Data consistency and integrity: In the data integration process, through clear rules and pandas' data processing functions, data consistency and integrity are ensured, avoiding data loss and inconsistency problems.
[0039] The above general description and the following description are merely exemplary and explanatory, and are not intended to limit this application. Description of the Drawings
[0040] One or more embodiments are exemplarily illustrated by corresponding drawings. These exemplary illustrations and the drawings do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation, and wherein:
[0041] Figure 1 is a schematic diagram of the method flow of the present invention;
[0042] Figure 2 is a schematic diagram of the method steps provided by an embodiment of the present disclosure;
[0043] Figure 3 is a schematic diagram of the system structure provided by an embodiment of the present disclosure;
[0044] Figure 4 is a schematic diagram of a cross-database query device based on pandas provided by an embodiment of the present disclosure. Detailed Embodiments
[0045] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the drawings. The attached drawings are for reference and illustration purposes only, and are not intended to limit the embodiments of the present disclosure. In the following technical description, for the sake of explanation, numerous details are provided to give a thorough understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be shown in a simplified manner.
[0046] In the specification and claims of the embodiments of the present disclosure and the above drawings, terms such as "first" and "second" are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so as to implement the embodiments of the present disclosure described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.
[0047] Unless otherwise specified, the term "plurality" means two or more.
[0048] In the embodiments of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.
[0049] The term "and / or" is an associative relationship describing an object, indicating that three relationships may exist. For example, A and / or B means: A or B, or, the three relationships of A and B.
[0050] The term "corresponding" can refer to an association relationship or a binding relationship. That A corresponds to B means that there is an association relationship or a binding relationship between A and B.
[0051] Embodiment 1
[0052] A cross-database query method based on pandas, which utilizes the data processing capabilities and flexible interfaces of pandas to implement an efficient, simple, and maintainable cross-database query solution, while ensuring data consistency and integrity.
[0053] The method includes:
[0054] S1. Database connection configuration: Configure connection information for the database and store the connection information in a configuration file or the database; the connection information includes database type, address, port, username, and password.
[0055] S2. Cross-database query request reception and parsing: Receive a cross-database query request from a user or a system, parse the request and decompose it into sub-query requests for different databases; the request includes database identifiers involved, query conditions, table and field information, table connection methods, and aggregation, sorting, or grouping operation information.
[0056] S3. Query execution: According to the sub-query requests and the corresponding database connection information, establish connections with different databases respectively, use the read_sql function or custom functions of pandas to execute the queries, and store the results as DataFrame objects.
[0057] S4. Query execution: Process issues of inconsistent data types and field names through custom mapping rules or data conversion functions of pandas, and use the merge or join functions of pandas to associate multiple DataFrame objects according to the table connection method in the request to generate the final result DataFrame.
[0058] S5. Result output and processing: Output the integrated result DataFrame in a specified format, including storing it as a file, writing it to the database, or converting it to JSON format.
[0059] As a refinement of the above embodiment, in step S1, for different databases, different connection drivers can be used (such as pymysql, psycopg2, etc. for relational databases, and corresponding driver programs for non-relational databases). These connection information can be stored using a dictionary or a configuration class to facilitate subsequent database connection operations.
[0060] As a refinement of the above embodiment, the custom mapping rules include field name mapping, data type coercion, and null value filling rules.
[0061] As a refinement of the above embodiment, the specified format includes CSV, Excel, or writing to the target database through the to_sql function of pandas.
[0062] Embodiment 2
[0063] A cross-database query system based on pandas includes the following modules:
[0064] Configuration module: used to store and manage connection information of different types of databases, supporting dynamic update and maintenance;
[0065] Request processing module: used to receive and parse cross-database query requests, split them into sub-query requests and distribute them;
[0066] Query execution module: contains multiple query execution units adapted to different database types, used to establish database connections, execute sub-queries and store the results as DataFrame objects;
[0067] Data integration module: used to perform data cleaning, type conversion and association operations on multiple DataFrame objects to generate an integrated result DataFrame;
[0068] Result output module: used to output the integrated result in the format specified by the user or store it in the target location.
[0069] As a refinement of the above embodiment, the query execution module supports query adaptation for relational databases and non-relational databases, and realizes compatibility by dynamically loading the corresponding database drivers. This module contains query execution units for different database types to ensure compatibility with different databases.
[0070] As a refinement of the above embodiment, the data integration module automatically corrects field name conflicts and data type differences through preset mapping rules.
[0071] As a refinement of the above embodiment, the result output module supports converting the result DataFrame into a JSON array and returning it through the API interface.
[0072] The following takes the development of a low-code platform as an example to refine the above system implementation. In this embodiment, the user table in the user management microservice stores information such as user names, ages, regions, user IDs, etc. There is an order table in the application microservice. The order table records user IDs, order IDs, order placement dates, etc., and it is necessary to query user orders according to the user name.
[0073] 1. Configure the connection information of the database tables in the user management microservice and the order table database in the application management microservice in the configuration module.
[0074] 2. The request module parses the query of user orders by username into two sub - queries: querying the user table by username and querying all order tables.
[0075] 3. In the execution module, according to the configuration information in 1, establish a database connection, use the read_sql function of pandas to read data from the database, and store the results in DataFrame_user and DataFrame_order respectively.
[0076] 4. The data integration module associates DataFrame_user and DataFrame_order through the user ID, and performs data merging through the merge function to obtain the result DataFrame_result.
[0077] 5. The result output module calls the to_json method provided by Pandas and returns it as a json array.
[0078] Combined Figure 4 As shown, the embodiment of the present disclosure provides a cross - database query device 300 based on pandas, including a processor 304 and a memory 301. Optionally, the device may further include a communication interface 302 and a bus 303. Among them, the processor 304, the communication interface 302, and the memory 301 can complete mutual communication through the bus 303. The communication interface 302 can be used for information transmission. The processor 304 can call the logical instructions in the memory 301 to execute the cross - database query method based on pandas in the above - mentioned embodiment.
[0079] In addition, when the logical instructions in the above - mentioned memory 301 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer - readable storage medium.
[0080] The memory 301, as a computer - readable storage medium, can be used to store software programs and computer - executable programs, such as the program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 304 executes functional applications and data processing by running the program instructions / modules stored in the memory 301, that is, implements the cross - database query method based on pandas in the above - mentioned embodiment.
[0081] The memory 301 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the terminal device and the like. In addition, the memory 301 may include a high-speed random access memory and may also include a non-volatile memory.
[0082] Embodiments of the present disclosure provide a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are configured to execute the above-mentioned cross-database query method based on pandas.
[0083] The above-mentioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transient computer-readable storage medium.
[0084] The technical solution of the embodiments of the present disclosure may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present disclosure. The foregoing storage medium may be a non-transient storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, and other media that can store program codes, or may also be a transient storage medium.
[0085] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. Embodiments merely represent possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terms used in this application are only for describing embodiments and do not limit the claims. As used in the description of embodiments and claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to also include the plural forms. Similarly, as used in this application, the term "and / or" refers to any and all possible combinations of one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" etc. mean the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, or apparatus comprising the element. Herein, what each embodiment focuses on may be the differences from other embodiments, and the same or similar parts among the embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method parts disclosed in the embodiments, the relevant parts may refer to the description of the method parts.
[0086] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner may depend on the specific application and design constraints of the technical solution. The technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure. The technician can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0087] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units can be merely a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. Additionally, in the embodiments of the present disclosure, the various functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to the embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks can also occur in a different order than disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A cross-database query method based on pandas, characterized in that: The following steps are involved: Configure connection information for the database and store the connection information in a configuration file or a database; the connection information includes database type, address, port, user name and password; Receiving a cross-database query request from a user or system, parsing the request and decomposing it into sub-query requests for different databases; the request includes the database identifier involved, query conditions, table and field information, table connection method, and aggregation, sorting or grouping operation information; According to the subquery request and the corresponding database connection information, establish connections with different databases respectively, use the pandas read_sql function or a custom function to execute the query, and store the results as a DataFrame object; Use custom mapping rules or pandas data conversion functions to handle data type and field name inconsistencies, and use pandas merge or join functions to associate multiple DataFrame objects according to the table connection method in the request to generate the final result DataFrame. Output the resulting DataFrame in a specified format, including storing it as a file, writing it to a database, or converting it to JSON format.
2. The cross-database query method according to claim 1, characterized in that: Relational databases use pymysql or psycopg2 drivers, and non-relational databases use corresponding drivers.
3. The cross-database query method according to claim 1, characterized in that: The custom mapping rules include field name mapping, data type mandatory conversion, and null value filling rules.
4. The cross-database query method according to claim 1, characterized in that: The specified format includes CSV, Excel, or writing to the target database through the to_sql function of pandas.
5. A cross-database query system based on pandas, characterized in that: Includes the following modules: Configuration module: used to store and manage connection information of different types of databases, supporting dynamic update and maintenance; Request processing module: used to receive and parse cross-database query requests, split them into sub-query requests and distribute them; Query execution module: contains query execution units adapted to multiple database types, used to establish database connections, execute subqueries, and store the results as DataFrame objects; Data integration module: used to perform data cleaning, type conversion and association operations on multiple DataFrame objects to generate the integrated result DataFrame; Result output module: used to output or store the integration results to the target location in the format specified by the user.
6. The cross-database query system according to claim 5, characterized in that: The query execution module supports query adaptation of relational databases and non-relational databases, and achieves compatibility by dynamically loading corresponding database drivers.
7. The cross-database query system according to claim 5, characterized in that: The data integration module automatically corrects field name conflicts and data type differences through preset mapping rules.
8. The cross-database query system according to claim 5, characterized in that: The result output module supports converting the result DataFrame into a JSON array and returning it through an API interface.
9. A cross-database query device based on pandas, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the pandas-based cross-database query method according to any one of claims 1 to 4 when running the program instructions.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the pandas-based cross-database query method as described in any one of claims 1 to 4 above is implemented.
Citation Information
Cited By
Dynamic form generation and cross-database adaptation method based on metadata driving
CN120743967A