Multi-domain collaborative query method and system for multiple bins, computer equipment and medium
By building independent business domain data sets in the data warehouse and using query engines and consistent dimensions to associate data, the problems of high maintenance costs and resource waste in traditional data warehouses are solved, and efficient and low-coupling data queries are achieved.
Patent Information
- Application Number
- CN202510638126.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-09-26
AI Technical Summary
When existing data warehouses use single-granularity data sets to integrate multiple business domains, there are problems such as high maintenance costs, waste of resources, and low fault tolerance.
Adopting the approach of multi-domain decoupling and independent construction, by building independent data sets based on different business domains, and using the query engine to build SQL statements according to business query requests, the target indicator data is extracted from each independent data set, and associated through the consistency dimension, and the asynchronous IO and caching mechanism are used to optimize data processing.
It reduces system coupling, avoids the high maintenance costs and resource waste brought by the traditional large and wide table architecture, improves the accuracy of data extraction and query efficiency, and adapts to different business needs.
Smart Images

Figure CN120705225A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of data warehouse technology, and in particular to a multi-domain collaborative query method, system, computer device, and medium for data warehouse. Background Art
[0002] In modern enterprise data management and analysis, data warehouses are critical infrastructure for efficient data processing and analysis. Data warehouses integrate various types of data within an enterprise to support decision-making. Self-service analytics are a crucial component of data warehouse applications, allowing users to perform data analysis, manipulate datasets, create dashboards, conduct interactive analysis, perform chart analysis, and perform query controls based on their needs. The core of self-service analytics is the dataset, a collection of data used for analysis, model building, or prediction. A dataset can include a variety of data types, including numerical and textual data.
[0003] In the existing data warehouse architecture, a single-granularity data set is usually used to integrate indicators from multiple business domains. For example, a driver-granularity data set may include indicators such as the number of completed orders in the transaction domain, the amount of activity rewards in the marketing domain, and the number of bans in the management and control domain. Although this design meets the needs of multi-business domain data integration to a certain extent, it has many problems in actual application. Ultimately, the data set queried by the query engine when splicing SQL (Structured Query Language) is usually one or more large wide tables. This large wide table design will have problems such as high maintenance costs, waste of resources, and low fault tolerance. Summary of the Invention
[0004] In response to the above-mentioned deficiencies or shortcomings, the present application provides a multi-domain collaborative query method, system, computer equipment and medium for a data warehouse. By decoupling and independently constructing multiple domains and adopting the method of constructing independent data sets by business domain, the high maintenance costs and resource waste brought about by the traditional large and wide table architecture are avoided, while reducing the system coupling.
[0005] According to a first aspect, the present application provides a multi-domain collaborative query method for a data warehouse. The method is based on a data warehouse, and the data warehouse includes multiple pre-built independent data sets corresponding to different business domains, each independent data set containing corresponding business domain indicator data. The method includes:
[0006] In response to receiving the business query request, a SQL statement is constructed based on the request content of the business query request through a set query engine.
[0007] Extract target indicator data from each independent data set based on SQL statements.
[0008] Create an association query logic according to the preset consistency dimension, and based on the association query logic, associate the target indicator data extracted from each independent data set to obtain the target query data.
[0009] In some embodiments, the target indicator data extracted from each independent data set is associated based on the association query logic, including:
[0010] Use SQL JOIN (Structured Query Language JOIN) operations to associate target indicator data based on consistency dimensions. SQL JOIN operations include at least one of INNER JOIN, LEFT JOIN, RIGHT JOIN, or FULL OUTER JOIN.
[0011] In some embodiments, when extracting target indicator data, the method further includes:
[0012] An asynchronous input / output (IO) mechanism is used to extract target indicator data.
[0013] A cache mechanism is introduced to store the target indicator data to obtain the target cache data. The cache mechanism includes the least recently used LRU (Least Recently Used) cache strategy and the ALL cache strategy.
[0014] In some embodiments, after obtaining the target query data, the method further includes:
[0015] In response to receiving a business query request with the same conditions, target result data is obtained from the target cache data. Alternatively, in response to receiving a new business query request, a new SQL statement is constructed based on the content of the new business query request.
[0016] "Same conditions" means that the service query request currently initiated by the user uses the same query parameters or filter conditions as the service query request, and the target result data is the query result data obtained based on the target cache data. "New service query request" means that the service query request currently initiated by the user uses query parameters or filter conditions that are inconsistent with the service query request.
[0017] In some embodiments, before constructing the SQL statement, the method further includes:
[0018] Select the time period, entity attribute, or spatial range corresponding to each business domain's business scenario as the core granularity for that business domain. Set the core granularity for each independent dataset to the same size. Each core granularity is the most basic unit used to record events or business actions in the data warehouse.
[0019] In some embodiments, the data dimensions of the consistency dimension include the operator, city, and age of the driver. Each business domain includes a marketing business domain, a transaction business domain, and a gross profit calculation business domain. The independent data set corresponding to the marketing business domain includes marketing indicators, which include the amount of activity rewards and the number of orders completed by participating in the activity. The independent data set corresponding to the transaction business domain includes transaction indicators, which include the number of responses and the number of completed orders. The independent data set corresponding to the gross profit calculation business domain includes a transaction data set and a financial data set. The transaction data set includes order quantity information and amount information. The financial data set includes the driver's commission amount and the commission-free amount. The order quantity information includes order issuance information, matching information, response information, and order completion information. The amount information includes the total order amount and travel fee. The core granularity of the marketing business domain, the transaction business domain, and the gross profit calculation business domain are atomic events or indivisible business behaviors in the marketing business, the transaction business, and the gross profit calculation business processes, respectively.
[0020] In some embodiments, the data warehouse pre-synchronizes each independent data set to a predetermined OLAP (Online Analytical Processing) database.
[0021] According to a second aspect, the present application provides a multi-domain collaborative query system for a data warehouse. The system is based on a data warehouse, and the data warehouse includes multiple pre-built independent data sets corresponding to different business domains, each independent data set containing corresponding business domain indicator data; the system includes:
[0022] The SQL statement splicing module is used to respond to receiving a business query request and construct an SQL statement based on the request content of the business query request through a set query engine.
[0023] The indicator data extraction module is used to extract target indicator data from each independent data set based on SQL statements.
[0024] The query data exposure module is used to create associated query logic according to the preset consistency dimension, and to associate the target indicator data extracted from each independent data set based on the associated query logic to obtain the target query data.
[0025] According to a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the multi-domain collaborative query method of a data warehouse in any one of the above embodiments are implemented.
[0026] According to a fourth aspect, the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When executed by the processor, the steps of the multi-domain collaborative query method for a data warehouse in any one of the above-mentioned embodiments are implemented.
[0027] The multi-domain collaborative query method of the data warehouse in the above embodiment is based on a data warehouse, which is divided into multiple independent data sets, each data set corresponding to a specific business domain and containing the indicator data of the business domain. However, in the traditional large wide table architecture, the data of all business domains are centrally stored in a wide table. The data structure is complex and highly coupled, resulting in extremely high maintenance costs. Any data change may trigger a chain reaction, affecting the stability and availability of the entire system. In contrast, the multi-domain collaborative query method makes the data structure of each business domain clear and independent by independently constructing data sets. When business needs change or data needs to be updated, it is only necessary to adjust the data sets of the relevant business domains without the need for large-scale reconstruction of the entire data warehouse. In addition, the method uses a set query engine to construct SQL statements based on the request content of the business query request, and extracts target indicator data from each independent data set. This process not only achieves accurate data extraction, but also creates associated query logic through preset consistency dimensions, effectively associates data from different business domains, and ultimately obtains the target query data.
[0028] Therefore, by decoupling and independently building multiple domains and building independent data sets by business domain, we avoid the high maintenance costs and resource waste brought about by the traditional large and wide table architecture, while reducing system coupling. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A schematic diagram of an application environment for a multi-domain collaborative query method for a data warehouse in one or more embodiments of the present application;
[0030] Figure 2 This is a flow chart of a multi-domain collaborative query method for a data warehouse in one or more embodiments of the present application;
[0031] Figure 3 A flow chart of a method for extracting target cache data in one or more embodiments of the present application;
[0032] Figure 4 A flowchart of a method for configuring an independent data set in one or more embodiments of the present application;
[0033] Figure 5 This is a schematic diagram of the structure of a multi-domain collaborative query system for a data warehouse in one or more embodiments of the present application;
[0034] Figure 6This is a schematic diagram of the internal structure of a computer device in one or more embodiments of the present application. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0036] The present application provides a multi-domain collaborative query method for data warehouses according to the first aspect, which can be applied to Figure 1 In the cloud environment shown, the user terminal (101) can first send a data query request to the cloud platform server (102), so that the cloud platform server (102) instructs the data warehouse management system of the data warehouse (103) to execute the SQL query task. Then, when the cloud platform server (102) receives the target query data transmitted by the data warehouse (103), it sends the target query data to the user terminal (101).
[0037] In some exemplary embodiments of the present application, the method is based on Figure 1 The data warehouse shown can be applied to a cloud platform server. The data warehouse includes multiple pre-built independent data sets corresponding to different business domains, and each independent data set contains corresponding business domain indicator data. For example, for an independent data set that is a marketing domain data set, the marketing domain data set stores business domain indicator data such as customer behavior and promotion effect. The transaction domain data set records business domain indicator data such as transaction amount and frequency. Next, the gross profit calculation domain data set contains business domain indicator data such as cost, revenue and profit. Figure 2 As shown, the method includes the following steps:
[0038] Step 201: In response to receiving a business query request, a SQL statement is constructed based on the request content of the business query request through a set query engine.
[0039] Specifically, after receiving a business query request from a user, the cloud platform server first uses the query engine deployed in the data warehouse to instruct the query engine to assemble an SQL statement according to the request content of the business query request.
[0040] For example, the business query request may include a marketing business query request, a transaction business query request, or a gross profit calculation business query request. The query engine analyzes the key information in the request content, determines the query target table, field, and conditions, and then constructs a complete SQL statement according to the SQL syntax format.
[0041] Step 202: Extract target indicator data from each independent data set according to the SQL statement.
[0042] Specifically, the query engine can send the assembled SQL statements to the database management system in the data warehouse. Based on the table name, field, and conditions in the SQL statement, the database management system retrieves matching data records from the corresponding independent data set, extracts target indicator data, such as marketing data, transaction data, or gross profit data, and returns the results to the cloud platform server for subsequent processing and display.
[0043] Step 203: Create an association query logic according to the preset consistency dimension, and associate the target indicator data extracted from each independent data set based on the association query logic to obtain target query data.
[0044] Specifically, the cloud platform server identifies the associated dimensions between independent data sets based on business needs and data models, such as time, customer, or business type. Based on these dimensions, the query engine constructs associated query logic and integrates the target indicator data extracted from different data sets through SQL associated query statements (such as INNER JOIN and LEFT JOIN). For example, marketing data and transaction data can be associated by customer ID (Identifier), or gross profit data and cost data can be associated by date. Ultimately, the integrated data serves as the target query data, providing consistency and integrity support for subsequent business analysis and decision-making.
[0045] Therefore, since the above embodiment decouples and independently constructs multiple domains and adopts the method of constructing independent data sets by business domain, it avoids the high maintenance costs and resource waste caused by the traditional large and wide table architecture, and reduces system coupling.
[0046] In some embodiments, the target indicator data extracted from each independent data set is associated based on the association query logic, including:
[0047] Use SQL JOIN operations to associate target indicator data based on consistency dimensions. SQL JOIN operations include at least one of INNER JOIN, LEFT JOIN, RIGHT JOIN, or FULL OUTERJOIN.
[0048] For example, when a business requirement requires analyzing the relationship between customer marketing behavior and transaction amounts, the query engine uses an INNER JOIN operation based on the consistent dimension of customer ID to link and integrate customer behavior data in the marketing domain dataset with transaction amount data in the transaction domain dataset. Alternatively, when analyzing gross profit measurement data and cost data, a LEFT JOIN operation is used based on the consistent dimension of date to link profit data in the gross profit measurement domain dataset with cost data. This approach enables efficient integration of multi-domain data, providing accurate and consistent data support for subsequent business analysis and decision-making.
[0049] In some embodiments, when extracting target indicator data, such as Figure 3 As shown, the method includes the following steps:
[0050] Step 301: Use an asynchronous input and output (IO) mechanism to extract target indicator data.
[0051] Specifically, the query engine sends SQL statements to the database management system in the data warehouse through an asynchronous IO mechanism, triggering the data retrieval task. The asynchronous IO mechanism allows the cloud platform server to continue processing other tasks while waiting for data retrieval results, improving system resource utilization and response speed.
[0052] For example, when processing a marketing business query request, the query engine extracts customer behavior data from the marketing domain dataset through an asynchronous IO mechanism, while processing other business query requests in parallel.
[0053] Step 302: Introduce a cache mechanism to store the target indicator data to obtain target cache data. The cache mechanism includes a least recently used (LRU) cache strategy and an ALL cache strategy.
[0054] Specifically, after extracting the target metric data, the query engine stores it in the cache. The LRU cache strategy eliminates the least recently used data based on the frequency and chronological order of data access, ensuring that frequently accessed data is always stored in the cache, thereby improving data access efficiency. The ALL cache strategy, on the other hand, stores all target metric data in the cache and is suitable for scenarios with small data volumes and the need for fast access to all data. By introducing a caching mechanism, the pressure on the data warehouse for repeated queries can be significantly reduced, improving the system's response speed and overall performance.
[0055] In some embodiments, after obtaining the target query data, the method further includes:
[0056] In response to receiving a business query request with the same conditions, target result data is obtained from the target cache data. Alternatively, in response to receiving a new business query request, a new SQL statement is constructed based on the content of the new business query request.
[0057] "Same conditions" means that the service query request currently initiated by the user uses the same query parameters or filter conditions as the service query request, and the target result data is the query result data obtained based on the target cache data. "New service query request" means that the service query request currently initiated by the user uses query parameters or filter conditions that are inconsistent with the service query request.
[0058] Here, "same conditions" means that the business query request currently initiated by the user uses the same query parameters or filter conditions as the previous business query request, and the target result data is the query result data obtained based on the target cache data. For example, the user's first query request is "Query the transaction amount of a certain customer in 2024". The query engine constructs the SQL statement and extracts data from the data warehouse, and then stores the results in the cache. When the user initiates a query request with the same conditions again, the system directly obtains the results from the target cache data without having to re-query the data warehouse, thereby significantly improving the response speed.
[0059] A "new business query request" refers to a business query request initiated by a user that uses different query parameters or filter conditions than the previous one. For example, if a user's first query is "a certain customer's transaction amount in 2024," and the new query request is "a certain customer's marketing behavior data in 2024," the system will re-analyze the request content due to the different query parameters, construct a new SQL statement through the query engine, extract the new target indicator data from the data warehouse, and update the target cache data.
[0060] In this way, the system can flexibly respond to different query requirements, while using the caching mechanism to optimize the efficiency of repeated queries, reduce the query pressure of the data warehouse, and improve overall performance and user experience.
[0061] In some embodiments, before constructing the SQL statement, such as Figure 4 As shown, the method includes the following steps:
[0062] Step 401: Select the time period, entity attribute or spatial range of each business domain corresponding to the business scenario as the core granularity of the corresponding business domain.
[0063] Specifically, in the multi-domain collaborative query approach, core granularity refers to the key dimension used to define the scope and precision of business domain data queries. For example, in the marketing business domain, core granularity might be the time period (e.g., the first quarter of 2024) and customer entity attributes (e.g., customer ID or region); in the transaction business domain, core granularity might be the transaction time period (e.g., January 2024) and transaction spatial scope (e.g., transaction location).
[0064] Step 402: Set the core granularity corresponding to each independent data set to the same size.
[0065] Each core granularity is the most basic unit used to record events or business behaviors in the data warehouse and is a key dimension that defines the scope and precision of data queries. For example, in a marketing domain dataset, the core granularity might be set to customer ID and time period (e.g., monthly); in a transaction domain dataset, the core granularity might be set to transaction ID and time period (also monthly).
[0066] In some embodiments, the data dimensions of the consistency dimension include the operator, city, and age of the driver. Each business domain includes a marketing business domain, a transaction business domain, and a gross profit calculation business domain. The independent data set corresponding to the marketing business domain includes marketing indicators, which include the amount of activity rewards and the number of orders completed by participating in the activity. The independent data set corresponding to the transaction business domain includes transaction indicators, which include the number of responses and the number of completed orders. The independent data set corresponding to the gross profit calculation business domain includes a transaction data set and a financial data set. The transaction data set includes order quantity information and amount information. The financial data set includes the driver's commission amount and the commission-free amount. The order quantity information includes order issuance information, matching information, response information, and order completion information. The amount information includes the total order amount and travel fee. The core granularity of the marketing business domain, the transaction business domain, and the gross profit calculation business domain are atomic events or indivisible business behaviors in the marketing business, the transaction business, and the gross profit calculation business processes, respectively.
[0067] For example, the core granularity of the marketing business domain is participation in a single marketing activity, the core granularity of the transaction business domain is a single transaction, and the core granularity of the gross profit measurement business domain is the financial settlement of a single transaction. For another example, in the marketing business domain, the core granularity is "single activity participation," which records the reward amount and number of orders completed for each activity. In the transaction business domain, the core granularity is "single transaction behavior," which records the number of responses and orders completed for each transaction. In the gross profit measurement business domain, the core granularity is "single transaction financial settlement," which records the total order amount, trip fee, driver commission amount, and commission-free amount for each transaction. This approach allows the system to flexibly respond to query requirements from different business domains while efficiently integrating cross-domain data through consistent dimensions (such as the driver's carrier, city, and age), providing accurate and consistent data support for subsequent business analysis and decision-making.
[0068] In another embodiment of a multi-domain collaborative query method for a data warehouse, the method is based on a data warehouse, and the data warehouse is divided into multiple independent data sets, each data set corresponds to a specific business domain and contains indicator data of the business domain, and the structure of each independent data set is consistent with the core granularity setting.
[0069] The first independent dataset corresponds to a specific marketing business domain. The independent dataset corresponding to this business domain stores metrics such as customer behavior and promotion effectiveness at the business layer or cache layer. The core granularity of this independent dataset is set to the customer ID and campaign period for the marketing business, recording atomic events such as the campaign reward amount and the number of completed orders from participating in the campaign. The second independent dataset corresponds to the transaction business domain. This independent dataset records the number of responses and completed orders. The core granularity of this independent dataset is the transaction ID and timestamp. The third independent dataset corresponds to the gross profit measurement business domain. This independent dataset includes both the transaction and financial datasets. The core granularity of this independent dataset is the financial settlement of a single transaction, such as the transaction ID and settlement date, and covers financial metrics such as the total order amount and the driver's commission amount. Therefore, the core granularity of each dataset is set to the same monthly period to ensure dimension alignment during cross-domain correlation. That is, both the marketing and transaction domains aggregate data on a monthly basis, achieving seamless correlation through the customer ID and time dimensions.
[0070] The query engine then constructs SQL statements based on the business request, extracts target metric data from each independent dataset, and correlates them using consistent dimensions. These consistent dimensions include customer ID, driver operator, monthly or quarterly time period, and city or regional spatial scope. Next, taking the analysis of the relationship between customer marketing effectiveness and transaction conversion as an example, during the data extraction phase, the data warehouse management system extracts customer activity reward amounts from the marketing domain dataset and completed order volume from the transaction domain dataset. Then, during the correlation logic phase, the two datasets are combined into a joint view using the SQL INNER JOIN operation, with customer ID and monthly time period as the correlation keys. This ultimately generates target query data containing customer ID, monthly reward amount, and monthly completed order volume to support business analysis. This process decouples independent datasets, avoiding the redundant storage of traditional wide-table architectures. It also leverages SQL JOIN operations (such as LEFT JOIN and FULL OUTER JOIN) to flexibly adapt to different business scenarios.
[0071] Finally, the data warehouse management system uses an asynchronous IO mechanism to improve data extraction efficiency. That is, the query engine sends multiple SQL requests to the data warehouse in parallel, and can process other tasks while waiting for a response, thereby reducing query latency. Among them, the LRU strategy automatically eliminates low-frequency access data and prioritizes the retention of high-frequency query results of monthly transaction summaries. Therefore, this embodiment achieves efficient, low-coupling cross-domain query analysis through domain-independent data sets, consistent dimension association, asynchronous IO and cache optimization mechanisms, significantly improving the query performance and business adaptability of the data warehouse, and is suitable for large-scale, multi-business domain enterprise-level data warehouse scenarios.
[0072] According to the second aspect, the present application provides a multi-domain collaborative query system for a data warehouse. The system is based on a data warehouse, which includes multiple pre-built independent data sets corresponding to different business domains. Each independent data set contains corresponding business domain indicator data, such as Figure 5 As shown, the system includes:
[0073] The SQL statement splicing module 110 is configured to construct an SQL statement based on the request content of the business query request through a set query engine in response to receiving the business query request.
[0074] The indicator data extraction module 120 is used to extract target indicator data from each independent data set according to the SQL statement.
[0075] The query data transmission module 130 is used to create an associated query logic according to a preset consistency dimension, and associate the target indicator data extracted from each independent data set based on the associated query logic to obtain target query data.
[0076] In some embodiments, the query data exposure module 130 is further configured to use an SQL JOIN operation to associate the target indicator data based on the consistency dimension. The SQL JOIN operation includes at least one of an INNER JOIN, a LEFT JOIN, a RIGHT JOIN, or a FULL OUTER JOIN.
[0077] In some embodiments, when extracting the target indicator data, the system is further configured to: extract the target indicator data using an asynchronous input / output (IO) mechanism, introduce a cache mechanism to store the target indicator data, and obtain target cache data, wherein the cache mechanism includes a least recently used (LRU) cache strategy and an all cache strategy.
[0078] In some embodiments, after obtaining the target query data, the system is further configured to: obtain target result data from the target cache data in response to receiving a business query request with the same conditions. Alternatively, in response to receiving a new business query request, construct a new SQL statement based on the content of the new business query request. The term "same conditions" means that the business query request currently initiated by the user uses query parameters or filter conditions that are consistent with the business query request, and the target result data is query result data obtained based on the target cache data; and the term "new business query request" means that the business query request currently initiated by the user uses query parameters or filter conditions that are inconsistent with the business query request.
[0079] In some embodiments, before constructing the SQL statement, the system is further configured to: select a time period, entity attribute, or spatial range corresponding to a business scenario in each business domain as a core granularity for the corresponding business domain. The core granularities corresponding to each independent data set are set to the same size. Each core granularity is the most basic unit for recording events or business behaviors in the data warehouse.
[0080] According to a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the multi-domain collaborative query method of a data warehouse in any one of the above embodiments are implemented.
[0081] According to a fourth aspect, the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When executed by the processor, the steps of the multi-domain collaborative query method for a data warehouse in any one of the above-mentioned embodiments are implemented.
[0082] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data related to multi-domain collaborative query analysis of the data warehouse. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements any of the above-mentioned multi-domain collaborative query methods of the data warehouse.
[0083] Among them, any reference to memory, storage, database or other media used in the various embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (RamCUs), direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0084] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0085] The above embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
[0086] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
Claims
1. A multi-domain collaborative query method for a data warehouse, characterized by: The method is based on a data warehouse, which includes a plurality of pre-built independent data sets corresponding to different business domains, each of which contains corresponding business domain indicator data; the method includes: In response to receiving a business query request, constructing an SQL statement based on the request content of the business query request through a set query engine; Extracting target indicator data from each of the independent data sets according to the SQL statement; An association query logic is created according to a preset consistency dimension, and target indicator data extracted from each of the independent data sets are associated based on the association query logic to obtain target query data.
2. The method according to claim 1, characterized in that Associating the target indicator data extracted from each of the independent data sets based on the association query logic includes: Using SQL JOIN operation to associate the target indicator data according to the consistency dimension; The SQL JOIN operation includes at least one of an INNER JOIN, a LEFT JOIN, a RIGHT JOIN, or a FULL OUTER JOIN.
3. The method according to claim 2, characterized in that When extracting the target indicator data, the method further includes: Using asynchronous input and output (IO) mechanism to extract the target indicator data; A cache mechanism is introduced to store the target indicator data to obtain target cache data. The cache mechanism includes a least recently used (LRU) cache strategy and an ALL cache strategy.
4. The method according to claim 3, characterized in that After obtaining the target query data, the method further includes: In response to receiving a service query request with the same conditions, obtaining the target result data from the target cache data; Alternatively, in response to receiving a new business query request, constructing a new SQL statement again based on the content of the new business query request; Among them, the same conditions mean that the business query request currently initiated by the user uses query parameters or filtering conditions that are consistent with the business query request, and the target result data is query result data obtained based on the target cache data; the new business query request means that the business query request currently initiated by the user uses query parameters or filtering conditions that are inconsistent with the business query request.
5. The method according to claim 1, wherein Before constructing the SQL statement, the method further includes: Selecting the time period, entity attribute or spatial range of each business domain corresponding to the business scenario as the core granularity of the corresponding business domain; The core granularities corresponding to the independent data sets are set to the same size; each core granularity is the most basic unit for recording events or business behaviors in the data warehouse.
6. The method according to claim 5, characterized in that The data dimensions of the consistency dimension include the operator, city, and age of the driver; each of the business domains includes a marketing business domain, a transaction business domain, and a gross profit calculation business domain; The independent data set corresponding to the marketing business domain includes marketing indicators, including the amount of activity rewards and the number of completed orders participating in the activity; The independent data set corresponding to the transaction business domain includes transaction indicators, including response volume and order completion volume; The independent data sets corresponding to the gross profit calculation business domain include transaction data sets and financial data sets. The transaction data sets include order quantity information and amount information. The financial data sets include the driver's commission amount and commission-free amount. The order quantity information includes order issuance information, matching information, response information, and order completion information. The amount information includes the total order amount and travel fee. The core granularities of the marketing business domain, the transaction business domain, and the gross profit calculation business domain are atomic events or indivisible business behaviors in the marketing business, transaction business, and gross profit calculation business processes, respectively.
7. The method according to claim 1, characterized in that The data warehouse pre-synchronizes each of the independent data sets to a predetermined online analytical processing (OLAP) database.
8. A multi-domain collaborative query system for data warehouses, characterized by: The system is based on a data warehouse, which includes multiple pre-built independent data sets corresponding to different business domains, each of which contains corresponding business domain indicator data; the system includes: An SQL statement splicing module is used to construct an SQL statement based on the request content of the business query request through a set query engine in response to receiving the business query request; An indicator data extraction module, configured to extract target indicator data from each of the independent data sets according to the SQL statement; The query data transmission module is used to create an associated query logic according to a preset consistency dimension, and associate the target indicator data extracted from each of the independent data sets based on the associated query logic to obtain target query data.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.