Analysis method, device and system for database sharing analysis resources
By sharing the analysis data clusters among multiple transaction data clusters and adopting a consistent hash cache, efficient utilization of analysis resources is achieved, and the problem of excessive resource coupling in traditional databases is solved, and the separate scaling and dynamic allocation of resources is supported, which improves resource utilization.
Patent Information
- Application Number
- CN202510745223.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
In traditional database engines, the coupling of analytical computing resources and transactional computing resources is too high, and analytical resources cannot be shared and it is difficult to expand and scale separately between the two types of resources, resulting in low resource utilization.
By setting up a shared analytics data cluster to share among multiple transaction data clusters, the data query results are cached using a consistent hashing method, and the analysis coordinator of the analytics data cluster can achieve separate scaling and dynamic allocation of resources, supporting the separate deployment of transaction data and analytics resources.
It improves the utilization rate of analytical resources, reduces the resource vacancy rate, realizes efficient execution of analytical operations, and solves the problems of excessive resource coupling and difficulty in scaling alone.
Smart Images

Figure CN120256470A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis, and particularly to an analysis method, device, and system for a database that shares analysis resources. Background Art
[0002] In a database system, transactional operations and analytical operations are two different types of data operations; transactional operations mainly focus on processing and managing individual transactions or deals. Transactional operations usually require fast read and write capabilities to ensure high throughput and low latency. Analytical operations mainly focus on performing complex analysis, summarization, and reporting on data to support decision-making and business analysis. Analytical operations usually do not require low latency but need to be able to process large datasets for efficient data analysis. Due to the significant differences in the implementation principles of transactional operations and analytical operations at the engine end, traditional database engines are usually only good at one type of operation, or have made simple splicing but cannot efficiently achieve data interconnection and resource isolation directly between the two. In the prior art, if the transaction data engine and the analysis data engine are in a completely isolated architecture at the engine end, the data of the transaction data engine cannot be efficiently submitted to the analysis data engine, or the analysis data engine bypasses the transaction data engine and directly loads data from the storage. In this case, there will also be problems such as inaccurate data screening and possible data inconsistency; if the transaction data engine and the analysis data engine are tightly coupled, the resource isolation of the resources required by the transaction data engine and the analysis data engine is not high, and the job execution of the analysis data engine will greatly affect the operations of the transaction data engine. Moreover, when the transaction data engine and the analysis data engine need to scale resources up and down simultaneously, resources are likely to compete with each other or be wasted and left idle. If the traditional hybrid structure of the transaction data engine and the analysis data engine is adopted, the resource isolation is relatively high, but the resource utilization rate is low. Summary of the Invention
[0003] The present invention provides an analysis method, device, and system for a database that shares analysis resources to solve the technical problems in the prior art that the coupling degree between analytical computing resources and transaction processing computing resources in a database engine is too high, analytical resources cannot be shared, and it is difficult to scale resources up and down separately for the two types of resources.
[0004] According to one aspect of the present invention, an analysis method for a database that shares analysis resources is provided, which is applied to at least one transaction data cluster; including: Based on at least one transaction data engine responding to the analytical statement text input by a data analysis user, generating a task identifier corresponding to the analytical statement text; For each of the transaction data engines, querying a subset of data to be analyzed corresponding to the analytical statement text; storing the subset of data to be analyzed in a transaction result pool; the subset of data to be analyzed includes at least one piece of data to be analyzed; Generate a data analysis request based on the data storage address of the transaction result pool of each of the transaction data engines and the task identifier, and send the data analysis request to the analysis data cluster; Receive the data analysis result of the data analysis request returned by the analysis data cluster, and display the data analysis result to the data analysis user.
[0005] According to another aspect of the present invention, there is provided an analysis method for a database sharing analysis resources, which is applied to an analysis data cluster; including: Based on the analysis coordinator of the analysis data cluster, receive and process a data analysis request, and allocate at least one analysis data engine for the data analysis request; Generate a data analysis task based on at least one data storage address and a task identifier, and send the data analysis task to each of the analysis data engines; For each of the analysis data engines, generate a data pull request based on the working number of the analysis data engine and the number of engine allocations of the analysis data engine corresponding to the data analysis request, and send the data pull request to each of the data storage addresses.
[0006] According to one aspect of the present invention, there is provided an analysis device for a database sharing analysis resources, which is deployed in at least one transaction data cluster; including: A task response module, configured to generate a task identifier corresponding to the analytical statement text based on at least one transaction data engine responding to the analytical statement text input by a data analysis user; A query module, configured to, for each of the transaction data engines, query a subset of data to be analyzed corresponding to the analytical statement text; store the subset of data to be analyzed in a transaction result pool; the subset of data to be analyzed includes at least one piece of data to be analyzed; A communication module, configured to generate a data analysis request based on the data storage address of the transaction result pool of each of the transaction data engines and the task identifier, and send the data analysis request to the analysis data cluster; A display module, configured to receive the data analysis result of the data analysis request returned by the analysis data cluster, and display the data analysis result to the data analysis user.
[0007] According to one aspect of the present invention, there is provided an analysis device for a database sharing analysis resources, which is deployed in an analysis data cluster; including: An analysis coordination module, configured to receive and process a data analysis request based on the analysis coordinator of the analysis data cluster, and allocate at least one analysis data engine for the data analysis request; An allocation module, configured to generate a data analysis task based on at least one data storage address and a task identifier, and send the data analysis task to each of the analysis data engines respectively; An analysis execution module, configured to, for each of the analysis data engines, generate a data pull request based on the working number of the analysis data engine and the number of engine allocations of the analysis data engine corresponding to the data analysis request, and send the data pull request to each of the data storage addresses.
[0008] The technical solution of the embodiment of the present invention reduces the resource vacancy rate of the analysis data cluster by sharing the shared analysis data cluster among multiple transaction data clusters, and improves the utilization rate of analytical resources; by setting an analysis coordinator for the analysis data engine in the analysis data cluster, the analysis data cluster and the transaction data cluster can be separately allocated and scaled independently, so that in the scenario of an analysis job, the resources of the analysis data engine can be matched, and in the scenario where no analysis job is required, the resources of the analysis data engine are not allocated, decoupling resource utilization; and the data query results of the transaction data cluster are cached in a consistent hashing manner, waiting for the analysis data engine to pull, and the analysis data engine can concurrently pull non-duplicate data, and due to the consistent hashing algorithm, it can be ensured that no re-positioning is required after the data is pulled and the analysis can be directly performed, greatly improving the analysis efficiency. It solves the technical problems in the prior art that the coupling degree between the analytical computing resources and the transaction processing computing resources in the database engine is too high, the analytical resources cannot be shared, and it is difficult to scale the two types of resources independently. The present invention can support the separate deployment of transaction data and analytical resources, share the analytical resources among multiple transaction data clusters, and support independent scaling; dynamically allocate analytical resources for tasks to improve resource utilization.
[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 It is a flowchart of an analysis method for a database sharing analytical resources provided by an embodiment of the present invention; Figure 2The present invention provides a flowchart of an analysis method for a database sharing analysis resources in an embodiment of the present invention; Figure 3 The present invention provides a flowchart of an analysis method for a database sharing analysis resources in an embodiment of the present invention; Figure 4 The present invention provides a flowchart of an analysis method for a database sharing analysis resources in an embodiment of the present invention; Figure 5 The present invention provides a schematic structural diagram of an analysis device for a database sharing analysis resources in an embodiment of the present invention; Figure 6 The present invention provides a schematic structural diagram of another analysis device for a database sharing analysis resources in an embodiment of the present invention; Figure 7 The present invention provides a schematic structural diagram of an analysis system for a database sharing analysis resources in an embodiment of the present invention. Detailed implementation manners
[0012] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.
[0013] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0014] Figure 1 The present invention provides a flowchart of an analysis method for a database sharing analysis resources in an embodiment of the present invention. This embodiment is applicable to the situation where a transaction data cluster responds to a user's data analysis requirement. This embodiment can be executed by a transaction data cluster. The analysis device of the database sharing analysis resources can be implemented in the form of hardware and / or software, and the analysis device of the database sharing analysis resources can be configured in the transaction data cluster. AsFigure 1 As shown, the method includes: S110. Based on at least one transaction data engine, analyze the analytical statement text input by the user in response to data, and generate a task identifier corresponding to the analytical statement text.
[0015] Among them, the transaction data engine can be a database that processes and manages individual transactions or transactions and supports transactional operations (OLTP - Online Transaction Processing, TP) for daily business activities.
[0016] Optionally, the components related to transactional operations in the transaction data engine may include a hybrid transaction analysis query plan parser, a transaction processing executor, and a transaction result pool; the hybrid transaction analysis query plan parser is used to parse the data request of the data analysis user and compile it into an execution plan corresponding to the data request; the transaction processing executor is responsible for transaction queries and outputs the data corresponding to the data request; the transaction result pool is responsible for caching the data corresponding to the data request; if the data request is an analytical statement text, after the transaction result pool caches the data, it waits for the analysis data engine to pull the data.
[0017] Among them, the analytical statement text can be a statement text for complex analysis of data; for example, the analytical statement text can be used for social network analysis, recommendation systems, and path analysis.
[0018] Optionally, the transactional statement text can be a statement text for transaction processing of data. Exemplarily, the transactional statement text can be used to handle operations such as adding, deleting, and modifying nodes and relationships in a graph.
[0019] Among them, the data analysis user can be a user who uses the database client or driver corresponding to the transaction data cluster; it should be noted that the data analysis user can send the analytical statement text and / or the transactional statement text to the transaction data cluster, and the transaction data cluster can parse the analytical statement text and the transactional statement text, compile them into execution plans corresponding to the analytical statement text and the transactional statement text, and generate corresponding task ids as task identifiers; among them, the task identifier corresponding to the analytical statement text is called an analysis task identifier; the task identifier corresponding to the transactional statement text is called a transaction task identifier.
[0020] Optionally, the transaction data cluster can be a cluster composed of at least one transaction data engine and can collectively process transactional operations.
[0021] Optionally, the data analysis user can input the analytical statement text to the transaction data cluster through the database client page provided by the database client, and then send the analytical statement text to the transaction data cluster.
[0022] Specifically, after the transaction data cluster receives the analytical statement text, at least one transaction data engine in the transaction data cluster responds to the analytical statement text input by the data analysis user, processes the analytical statement text, compiles it into an execution plan corresponding to the analytical statement text, and generates a task identifier corresponding to the analytical statement text.
[0023] S120. For each of the transaction data engines, query the subset of data to be analyzed corresponding to the analytical statement text; store the subset of data to be analyzed in the transaction result pool; the subset of data to be analyzed includes at least one piece of data to be analyzed.
[0024] Among them, the subset of data to be analyzed can be the query result constructed from the data stored in the transaction data engine that conforms to the analytical statement text. Exemplarily, the data stored in the transaction data engine can be the user profiles of a certain video website; the analytical statement text can be to analyze the age distribution of users of the video website in a certain region. Then, the transaction data engine queries all users and user ages in that region, and the query result is all users and user ages in that region; all users and user ages in that region are the subset of data to be analyzed corresponding to the analytical statement text.
[0025] Among them, the transaction result pool can be set by the transaction data engine to cache the subset of data to be analyzed. It should be noted that each transaction data engine sets a transaction result pool.
[0026] Among them, the data to be analyzed can be the query data obtained based on the analytical statement text in the transaction data engine. It should be noted that when the transaction data engine queries the analytical statement text, it queries one by one the data to be analyzed in the transaction data engine that conforms to the analytical statement text, and stores the data to be analyzed in the transaction result pool one by one, and then forms the subset of data to be analyzed.
[0027] Optionally, the subset of data to be analyzed includes at least one piece of data to be analyzed. Exemplarily, the data stored in the transaction data engine can be the user profiles of a certain video website; the analytical statement text can be to analyze the age distribution of users of the video website in a certain region. Then, the transaction data engine queries all users and user ages in that region. When one user and user age in that region are queried, it is the data to be analyzed. After the query is completed, the query result is all users and user ages in that region; all users and user ages in that region are the subset of data to be analyzed corresponding to the analytical statement text.
[0028] Optionally, when the transaction data cluster receives the analytical statement text, since the data corresponding to the analytical statement text may be stored in different transaction data engines of the transaction data cluster, data queries are performed on each transaction data engine based on the analytical statement text, and the query results are stored in the transaction result pools of the respective transaction data engines.
[0029] Specifically, when each transaction data engine receives the analytical statement text, for each transaction data engine, the subset of data to be analyzed corresponding to the analytical statement text is queried, and the subset of data to be analyzed is stored in the transaction result pool.
[0030] Optionally, the transaction result pool includes at least one storage data slot, and the storage data slot can be used to store each subset of data to be analyzed. After the subset of data to be analyzed is queried, the subset of data to be analyzed is stored in a storage data slot; it should be noted that the transaction result pool is divided into any number of storage data slots, and the storage data slots can be numbered. Each storage data slot has a corresponding data slot number. Each subset of data to be analyzed can fall into any one storage data slot or into an assigned storage data slot, and only one analytical data engine will pull data from each storage data slot. Exemplarily, the transaction result pool can be divided into 1024 or 65526 storage data slots, and the data slot numbers start from 0, 1, 2 until the total number of slots minus one, that is, 1023 or 65525.
[0031] Optionally, in another optional embodiment of the present invention, the step of storing the subset of data to be analyzed in the transaction result pool includes: For each subset of data to be analyzed, identify the target data slot corresponding to the subset of data to be analyzed in the storage data slot; sequentially store each subset of data to be analyzed in the target data slot of each subset of data to be analyzed.
[0032] Among them, the target data slot can be the storage data slot into which the subset of data to be analyzed needs to be stored. Optionally, when the subset of data to be analyzed is stored or cached, a target data slot can be matched for the subset of data to be analyzed, and the subset of data to be analyzed is stored in the target data slot.
[0033] Optionally, each subset of data to be analyzed has a target data slot, but multiple subsets of data to be analyzed can have the same target data slot.
[0034] Optionally, after each transaction data engine queries each subset of data to be analyzed, identify each subset of data to be analyzed, match a storage data slot as the target data slot for each subset of data to be analyzed, and sequentially store each subset of data to be analyzed in the target data slot of each subset of data to be analyzed.
[0035] Specifically, for each piece of data to be analyzed, identify the target data slot corresponding to the data to be analyzed in the storage data slots; sequentially store each piece of data to be analyzed into the target data slot of each piece of data to be analyzed.
[0036] Optionally, in another alternative embodiment of the present invention, the identifying the target data slot corresponding to the data to be analyzed in the storage data slots includes: Perform a hash calculation on the data to be analyzed through a preset hash function to obtain a hash calculation value corresponding to the data to be analyzed; Obtain the data slot numbers of each of the storage data slots, and determine the storage data slot with the same data slot number as the hash calculation value as the target data slot of the data to be analyzed.
[0037] Among them, the preset hash function can be a consistent hash function.
[0038] Among them, the hash calculation value can be the hash value obtained by calculating the data to be analyzed based on the hash function.
[0039] Optionally, for each piece of data to be analyzed, perform a hash calculation on the data to be analyzed through a preset hash function to obtain a hash calculation value corresponding to the data to be analyzed, and use the data slot number that is the same as the hash calculation value as the target data slot.
[0040] Specifically, perform a hash calculation on the data to be analyzed through a preset hash function to obtain a hash calculation value corresponding to the data to be analyzed; obtain the data slot numbers of each storage data slot, and determine the storage data slot with the same data slot number as the hash calculation value as the target data slot of the data to be analyzed.
[0041] S130. Generate a data analysis request based on the data storage address of the transaction result pool of each transaction data engine and the task identifier, and send the data analysis request to the analysis data cluster.
[0042] Among them, the data storage address can be the storage address where the transaction result pool of the transaction data engine stores the subset of data to be analyzed. It should be noted that the data storage address of the transaction data engine points to the transaction result pool, and the analysis data engine can obtain the subset of data to be analyzed stored in the transaction result pool by accessing the data storage address.
[0043] Among them, the data analysis request can be the request notification information sent by the transaction data cluster to the analysis data cluster. The data analysis request consists of the data storage addresses of each transaction data engine that responds to the analytical statement text, the task identifier, and the analysis requirements corresponding to the analytical statement text.
[0044] Optionally, the transaction data cluster can also evaluate the amount of data to be analyzed in the analytical statement text and put the amount of data to be analyzed into the data analysis request to notify the analytical data cluster.
[0045] Optionally, the transaction data cluster parses the analytical statement text to obtain the analysis requirements of the analytical statement text.
[0046] Optionally, the transaction data cluster can also set the data return policy for each transaction data engine and put the data return policy set for the transaction data engine into the data analysis request to notify the analytical data cluster. Among them, the data return policy is the policy for returning the data analysis result to the transaction data cluster.
[0047] Specifically, the transaction data cluster generates a data analysis request based on the data storage address and task identifier of the transaction result pool of each transaction data engine, and sends the data analysis request to the analytical data cluster.
[0048] S140. Receive the data analysis result of the data analysis request returned by the analytical data cluster, and display the data analysis result to the data analysis user.
[0049] Among them, the data analysis result can be the analysis result obtained by the analytical data cluster based on the analytical statement text and the data to be analyzed. It should be noted that after the analytical data cluster obtains the data to be analyzed of each transaction data engine in the transaction data cluster, it performs a joint analysis on the data to be analyzed based on the analytical statement text to obtain the data analysis result.
[0050] Optionally, after the analytical data cluster analyzes the data to be analyzed to obtain the data analysis result, it returns the data analysis result to the transaction data cluster, and the transaction data cluster displays the data analysis result in the database client of the data analysis user.
[0051] Optionally, if it is the way of synchronous return of the data analysis result, during the process of waiting for the data analysis result, the transaction data cluster displays a waiting analysis animation in the database client of the data analysis user, and after obtaining the data analysis result, it ends the waiting analysis animation and displays the data analysis result.
[0052] Optionally, if it is the way of asynchronous return of the data analysis result, the transaction data cluster does not display a waiting analysis animation in the database client of the data analysis user, and gradually loads the data analysis result in a streaming manner until the analytical data cluster completes all analyses and loads the data analysis result.
[0053] Specifically, the transaction data cluster receives the data analysis result of the data analysis request returned by the analytical data cluster, and displays the data analysis result to the data analysis user.
[0054] The technical solution of the embodiment of the present invention reduces the resource vacancy rate of the analysis data cluster by setting a shared analysis data cluster to be shared among multiple transaction data clusters, thereby improving the utilization rate of analytical resources; by setting an analysis coordinator of the analysis data engine in the analysis data cluster, the analysis data cluster and the transaction data cluster can be separately allocated and scaled independently, so that in the scenario of analysis jobs, the resources of the analysis data engine can be matched, and in the scenario where analysis jobs are not required, the resources of the analysis data engine are not allocated, decoupling resource utilization; and the data query results of the transaction data cluster are cached in a consistent hashing manner, waiting for the analysis data engine to pull, and the analysis data engine can concurrently pull non-duplicate data, and due to the consistent hashing algorithm, it can be ensured that the data does not need to be repositioned after being pulled and can be directly analyzed, greatly improving the analysis efficiency. This solves the technical problems in the prior art that the coupling degree between analytical computing resources and transaction processing computing resources in the database engine is too high, analytical resources cannot be shared, and it is difficult to scale these two types of resources independently. The present invention can support the separate deployment of transaction data and analytical resources, share analytical resources among multiple transaction data clusters, and support independent scaling; dynamically allocate analytical resources for tasks to improve resource utilization rate.
[0055] Figure 2 The flowchart of an analysis method for a database sharing analytical resources provided by an embodiment of the present invention is applicable to the situation where an analysis data cluster responds to data analysis requests from each transaction data cluster. This embodiment can be executed by the analysis data cluster. The analysis device of this database sharing analytical resources can be implemented in the form of hardware and / or software, and the analysis device of this database sharing analytical resources can be configured in the analysis data cluster. As Figure 2 shown, the method includes: S210. Receive and process a data analysis request based on the analysis coordinator of the analysis data cluster, and allocate at least one analysis data engine for the data analysis request.
[0056] Optionally, the analysis data cluster can be used to execute large-scale data analysis tasks. The analysis data cluster consists of at least one analysis data engine and an analysis coordinator. The analysis coordinator can be the data analysis request of the transaction data cluster in the analysis data cluster. The analysis coordinator can be a specific computing device with computing capabilities; any one of the analysis data engines can also be designated as the analysis coordinator; the analysis data engine can pull data, execute the analysis tasks corresponding to the analytical statement text, and write the data analysis results to the specified location in the transaction data cluster.
[0057] Optionally, the analysis coordinator of the analysis data cluster can communicate with multiple transaction data clusters to share the analysis data cluster among multiple transaction data clusters.
[0058] Optionally, the analysis data engine may be composed of an input controller, a task executor, and a result controller. The input controller may be used to pull data, the task executor may be used to execute the analysis task corresponding to the analytical statement text, and the result controller may be used to write the data analysis result to a specified location in the transaction data cluster.
[0059] Specifically, after receiving a data analysis request, the analysis coordinator of the analysis data cluster processes the data analysis request, evaluates the number of analysis data engines required to process the data analysis request, and allocates at least one analysis data engine to the data analysis request.
[0060] S220. Generate a data analysis communication message based on at least one data storage address and a task identifier, and send the data analysis communication message to each of the analysis data engines.
[0061] Among them, the data analysis communication message may be a communication message used to notify the analysis data engine to process the data analysis request. It should be noted that the data analysis communication message is composed of a task identifier, the number of engine allocations of the analysis data engine for processing the data analysis request, the working number of each analysis data engine for processing the data analysis request, and the data storage address of each transaction data engine.
[0062] Optionally, when the number of analysis data engines allocated by the analysis coordinator of the analysis data cluster for a data analysis request is greater than 1, the analysis coordinator of the analysis data cluster will set a working number for each allocated analysis data engine and record the number of engine allocations of the analysis data engine. The working number may be the number corresponding to the data analysis request when the analysis data engine processes a data analysis request, and the number of engine allocations may be the total number of analysis data engines allocated by the analysis coordinator of the analysis data cluster for the data analysis request. It should be noted that in the same data analysis request, the working number uniquely corresponds to one analysis data engine, and in different data analysis requests, the working numbers of the analysis data engines may be different.
[0063] Specifically, generate a data analysis communication message based on at least one data storage address and a task identifier, and send the data analysis communication message to each of the analysis data engines.
[0064] S230. For each of the analysis data engines, generate a data pull request based on the working number of the analysis data engine and the number of engine allocations of the analysis data engine corresponding to the data analysis request, and send the data pull request to each of the data storage addresses.
[0065] Among them, the data pulling request can be a communication request for the analysis data engine to request a transaction data engine to pull the data to be analyzed in the transaction result pool of the transaction data engine. It should be noted that the analysis data engine generates a data pulling request based on the job number and the number of engines allocated for the data analysis request by the analysis coordinator of the analysis data cluster.
[0066] Optionally, when the analysis data engine generates a data pulling request, a task identifier can also be added to the data pulling request to inform the transaction data engine of the data to be analyzed that needs to be pulled.
[0067] Optionally, since the analysis data engine can synchronously pull data from at least one transaction data engine when pulling data, the analysis data engine can generate a data pulling request for each transaction data engine and send the data pulling request to each transaction data engine that needs to pull data.
[0068] Optionally, the score data engine can also generate only one data pulling request and send the data pulling request to each transaction data engine that needs to pull data.
[0069] Specifically, for each analysis data engine, a data pulling request is generated based on the job number of the analysis data engine and the number of engines allocated for the analysis data engine corresponding to the data analysis request, and the data pulling request is sent to each data storage address.
[0070] Optionally, in another optional embodiment of the present invention, each of the analysis data engines receives at least one piece of data to be analyzed sent by at least one transaction data engine; after each of the analysis data engines receives the data sending completion identifier sent by all of the transaction data engines, data analysis is performed on at least one piece of the data to be analyzed based on the data analysis task to obtain a data analysis result; the data analysis result is returned to at least one transaction data cluster corresponding to the data analysis request.
[0071] Among them, the data sending completion identifier can be a data query completion identifier generated by the transaction data engine after querying the data to be analyzed in the transaction data engine that conforms to the analytical statement text. The completion identifier can be used to inform the analysis data engine that no new data to be analyzed will be stored in the transaction result pool of this transaction data engine.
[0072] Optionally, since the analysis data engine pulls the data to be analyzed in the transaction result pool of the transaction data engine in a streaming manner, after receiving a completion identifier of a transaction data engine, it stops pulling data from this transaction data engine.
[0073] Optionally, after each analysis data engine receives the data sending completion flag sent by each transaction data engine, it is considered that each analysis data engine has pulled all the data to be analyzed, and each analysis data engine needs to perform data analysis on all the pulled data to be analyzed.
[0074] Specifically, each analysis data engine receives at least one data to be analyzed sent by at least one transaction data engine; after each analysis data engine receives the data sending completion flag sent by all transaction data engines, it performs data analysis on at least one data to be analyzed based on the data analysis task to obtain a data analysis result; and returns the data analysis result to at least one transaction data cluster corresponding to the data analysis request.
[0075] The technical solution of the embodiment of the present invention reduces the resource vacancy rate of the analysis data cluster and improves the utilization rate of analysis resources by setting a shared analysis data cluster to be shared among multiple transaction data clusters; by setting an analysis coordinator for the analysis data engine in the analysis data cluster, the analysis data cluster and the transaction data cluster can be separately allocated and scaled independently, so that in the scenario of analysis jobs, the resources of the analysis data engine can be matched, and in the scenario where analysis jobs are not required, the resources of the analysis data engine are not allocated, decoupling resource utilization; and the data query results of the transaction data cluster are cached in a consistent hashing manner, waiting for the analysis data engine to pull, and the analysis data engine can concurrently pull non-repeated data, and due to the consistent hashing algorithm, it can be ensured that the data does not need to be re-positioned after being pulled and can be directly analyzed, greatly improving the analysis efficiency. It solves the technical problems in the prior art that the coupling degree between the analysis computing resources and the transaction processing computing resources in the database engine is too high, the analysis resources cannot be shared, and it is difficult to scale the two types of resources independently. The present invention can support the separate deployment of transaction data and analysis resources, share the analysis resources among multiple transaction data clusters, and support independent scaling; dynamically allocate analysis resources for tasks and improve resource utilization.
[0076] Figure 3 FIG. is a flowchart of an analysis method for a database sharing analysis resources according to an embodiment of the present invention. This embodiment is applicable to describing the entire process of a transaction data cluster and an analysis data cluster processing an analytical statement text. As Figure 3 shown, the method includes: S310. Based on at least one transaction data engine responding to an analytical statement text input by a data analysis user, generate a task identifier corresponding to the analytical statement text.
[0077] S320. For each of the transaction data engines, query the subset of data to be analyzed corresponding to the text of the analytical statement; store the subset of data to be analyzed in the transaction result pool; the subset of data to be analyzed includes at least one piece of data to be analyzed.
[0078] S330. Generate a data analysis request based on the data storage address of the transaction result pool of each transaction data engine and the task identifier, and send the data analysis request to the analysis data cluster.
[0079] S340. The analysis coordinator of the analysis data cluster receives and processes the data analysis request, and allocates at least one analysis data engine for the data analysis request.
[0080] S350. Generate a data analysis communication message based on at least one data storage address and the task identifier, and send the data analysis communication message to each of the analysis data engines.
[0081] S360. For each of the analysis data engines, generate a data pull request based on the working number of the analysis data engine and the number of engine allocations of the analysis data engine corresponding to the data analysis request, and send the data pull request to each of the data storage addresses.
[0082] S370. For each of the transaction data engines, respond to and process at least one received data pull request, and determine the working number of the analysis data engine corresponding to each data pull request and the number of engine allocations.
[0083] Optionally, when the transaction data engine of the transaction data cluster receives the data pull requests of each analysis data engine, if there is a task identifier in the data pull request, identify the task identifier and determine the data to be scored corresponding to the task identifier.
[0084] Specifically, when the transaction data engine of the transaction data cluster receives the data pull requests of each analysis data engine, respond to and process the data pull requests of each analysis data engine, and determine the working number of the analysis data engine corresponding to each data analysis request and the number of engine allocations of all analysis data engines.
[0085] S380. Allocate an analysis data engine for each of the storage data slots based on the working number and the number of engine allocations of the analysis data engine, and send the data to be analyzed in the storage data slot to the corresponding analysis data engine in the form of a data stream.
[0086] Optionally, when the analysis data engine fetches the data to be analyzed in each storage data slot, a unique analysis data engine is assigned to each storage data slot based on the working number of the analysis data engine and the engine allocation quantity, which improves the efficiency of data fetching and prevents conflicts in data fetching.
[0087] Optionally, during the data fetching process, the transaction data engine is still querying the data to be analyzed corresponding to the analytical statement text, putting the data to be analyzed into the target data slot corresponding to the data to be analyzed, and the analysis data engine fetches the data to be analyzed in the form of a data stream from the storage data slot, so as to achieve synchronous querying and fetching.
[0088] Optionally, if the analysis data engine queries that there is no data to be analyzed in a storage data slot, it chooses to wait for the transaction data engine to deposit the data to be analyzed into this storage data slot until the transaction data engine notifies the analysis data engine of the data sending completion flag, and then stops fetching the data in the storage data slot.
[0089] Specifically, an analysis data engine is assigned to each storage data slot based on the working number of the analysis data engine and the engine allocation quantity, and the data to be analyzed in the storage data slot is sent to the corresponding analysis data engine of the storage data slot in the form of a data stream.
[0090] Optionally, in another optional embodiment of the present invention, the step of assigning an analysis data engine to each storage data slot based on the working number of the analysis data engine and the engine allocation quantity includes: Performing a modulo operation on each data slot number and the engine allocation quantity respectively to determine the data fetching number corresponding to each storage data slot; Assigning an analysis data engine to each storage data slot based on the data fetching number and the working number.
[0091] Optionally, in the present invention, the number of storage data slots is greater than the engine allocation quantity of the analysis data engines. When performing a modulo operation on each data slot number and the engine allocation quantity to determine the data fetching number corresponding to each storage data slot, the storage data slots are evenly allocated to each analysis data engine. Each analysis data engine can be assigned multiple storage data slots, and the difference in the number of storage data slots assigned to each analysis data engine is extremely small.
[0092] Optionally, after assigning the data fetching number to the analysis data engine, the transaction data engine finds the analysis data engine equal to the data fetching number and restores the data to be analyzed in this storage data slot to this analysis data engine.
[0093] Specifically, perform a modulo operation on each data slot number and the number of engines allocated respectively to determine the data pull number corresponding to each storage data slot; allocate an analysis data engine for each storage data slot based on the data pull number and the job number.
[0094] S390. Receive at least one data to be analyzed sent by at least one transaction data engine through each of the analysis data engines.
[0095] S3100. When each of the analysis data engines has received the data sending completion identifier sent by all the transaction data engines, perform data analysis on at least one of the data to be analyzed based on the data analysis task to obtain a data analysis result. S3110. Return the data analysis result to at least one transaction data cluster corresponding to the data analysis request.
[0096] S3120. Receive the data analysis result of the data analysis request returned by the analysis data cluster, and display the data analysis result to the data analysis user.
[0097] The technical solution of the embodiment of the present invention reduces the resource vacancy rate of the analysis data cluster and improves the utilization rate of analytical resources by setting a shared analysis data cluster to be shared among multiple transaction data clusters; by setting an analysis coordinator for the analysis data engine in the analysis data cluster, the analysis data cluster and the transaction data cluster can be allocated separately and scaled independently, so that in the scenario of an analysis job, the analysis data engine resources can be matched, and in the scenario where no analysis job is required, no analysis data engine resources are allocated, decoupling resource utilization; and the data query results of the transaction data cluster are cached in a consistent hashing manner, waiting for the analysis data engine to pull, and the analysis data engine can concurrently pull non-duplicate data, and due to the consistent hashing algorithm, it can be ensured that the data does not need to be repositioned after being pulled and can be directly analyzed, greatly improving the analysis efficiency. It solves the technical problems in the prior art that the coupling degree between the analytical computing resources and the transaction processing computing resources in the database engine is too high, the analytical resources cannot be shared, and it is difficult to scale the two types of resources independently. The present invention can support the separate deployment of transaction data and analytical resources, share the analytical resources among multiple transaction data clusters, and support independent scaling; dynamically allocate analytical resources for tasks and improve resource utilization.
[0098] Figure 4 It is a flowchart of an analysis method for a database sharing analytical resources provided by an embodiment of the present invention; as Figure 4As shown in the figure, the database client sends the analytical statement text to the transaction data cluster. The transaction data engines 401 and 402 of the transaction data cluster query and store the subset of data to be analyzed based on the analytical statement text. The transaction data cluster notifies the analysis data cluster to fetch the data and sends a data analysis request to the analysis data cluster. The analysis coordinator of the analysis data cluster receives and processes the data analysis request, notifies the analysis data engine 403, and the analysis data engine 403 pulls the data of the subset of data to be analyzed from the transaction data engines 401 and 402 respectively, performs data analysis, and returns the data analysis result to the transaction data cluster.
[0099] Figure 5 As shown in the figure, it is a schematic structural diagram of an analysis device for a database that shares analysis resources provided by an embodiment of the present invention. Figure 5 As shown in the figure, the device includes: a task response module 510, a query module 520, a communication module 530, and a display module 540; among them, The task response module 510 is configured to generate a task identifier corresponding to the analytical statement text based on at least one transaction data engine in response to the analytical statement text input by the data analysis user. The query module 520 is configured to, for each of the transaction data engines, query the subset of data to be analyzed corresponding to the analytical statement text; store the subset of data to be analyzed in the transaction result pool; the subset of data to be analyzed includes at least one piece of data to be analyzed. The communication module 530 is configured to generate a data analysis request based on the data storage address of the transaction result pool of each of the transaction data engines and the task identifier, and send the data analysis request to the analysis data cluster. The display module 540 is configured to receive the data analysis result of the data analysis request returned by the analysis data cluster and display the data analysis result to the data analysis user.
[0100] The technical solution of the embodiment of the present invention reduces the resource vacancy rate of the analysis data cluster by setting a shared analysis data cluster to be shared among multiple transaction data clusters, and improves the utilization rate of analytical resources; by setting an analysis coordinator of the analysis data engine in the analysis data cluster, the analysis data cluster and the transaction data cluster can be separately allocated and scaled independently, so that in the scenario of analysis jobs, the resources of the analysis data engine can be matched, and in the scenario where analysis jobs are not required, the resources of the analysis data engine are not allocated, decoupling resource utilization; and the data query results of the transaction data cluster are cached in a consistent hashing manner, waiting for the analysis data engine to pull, and the analysis data engine can concurrently pull non-duplicate data, and due to the consistent hashing algorithm, it can be ensured that after the data is pulled, it can be directly analyzed without re-changing the position, greatly improving the analysis efficiency. It solves the technical problems in the prior art that the coupling degree between the analytical computing resources and the transaction processing computing resources in the database engine is too high, the analytical resources cannot be shared, and it is difficult to scale the two types of resources independently. The present invention can support the separate deployment of transaction data and analytical resources, share the analytical resources among multiple transaction data clusters, and support independent scaling; dynamically allocate analytical resources for tasks to improve resource utilization rate.
[0101] Optionally, the query module 520 is specifically configured to: the transaction result pool includes at least one storage data slot; For each piece of data to be analyzed, identify the target data slot corresponding to the data to be analyzed in the storage data slot; sequentially store each piece of data to be analyzed into the target data slot of each piece of data to be analyzed.
[0102] Optionally, the query module 520 is further specifically configured to: perform a hash calculation on the data to be analyzed through a preset hash function to obtain a hash calculation value corresponding to the data to be analyzed; Obtain the data slot numbers of each storage data slot, and determine the storage data slot with the same data slot number as the hash calculation value as the target data slot of the data to be analyzed.
[0103] Optionally, the device further includes: a data slot module and a data pulling module; wherein, The allocation module is configured to, for each transaction data engine, respond to and process at least one received data pulling request, and determine the working number and the engine allocation quantity of the analysis data engine corresponding to each data pulling request; The data pulling module is configured to allocate an analysis data engine for each storage data slot based on the working number and the engine allocation quantity of the analysis data engine, and send the data to be analyzed in the storage data slot to the analysis data engine corresponding to the storage data slot in the form of a data stream.
[0104] The data slot module is specifically configured to: Perform modulo calculations on each of the data slot numbers and the engine allocation quantities respectively to determine the data pull numbers corresponding to each of the storage data slots; and allocate analysis data engines to each of the storage data slots based on the data pull numbers and the work numbers.
[0105] The analysis device of the database for sharing analysis resources provided by the embodiments of the present invention can execute the analysis method of the database for sharing analysis resources provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0106] Figure 6 It is a schematic structural diagram of another analysis device of the database for sharing analysis resources provided by the embodiments of the present invention, as Figure 6 shown. The device includes: an analysis coordination module 610, an allocation module 620, and an analysis execution module 630; wherein, The analysis coordination module 610 is configured to receive and process a data analysis request based on the analysis coordinator of the analysis data cluster, and allocate at least one analysis data engine for the data analysis request; The allocation module 620 is configured to generate a data analysis communication message based on at least one data storage address and a task identifier, and send the data analysis communication message to each of the analysis data engines respectively; The analysis execution module 630 is configured to, for each of the analysis data engines, generate a data pull request based on the work number of the analysis data engine and the engine allocation quantity of the analysis data engine corresponding to the data analysis request, and send the data pull request to each of the data storage addresses.
[0107] The technical solution of the embodiment of the present invention reduces the resource vacancy rate of the analysis data cluster by setting a shared analysis data cluster to be shared among multiple transaction data clusters, thereby improving the utilization rate of analytical resources; by setting an analysis coordinator of the analysis data engine in the analysis data cluster, the analysis data cluster and the transaction data cluster can be separately allocated and scaled independently, so that in the scenario of an analysis job, the resources of the analysis data engine are matched, and in the scenario where no analysis job is required, the resources of the analysis data engine are not allocated, decoupling resource utilization; and the data query results of the transaction data cluster are cached in a consistent hashing manner, waiting for the analysis data engine to pull, and the analysis data engine can concurrently pull non-duplicate data. And due to the consistent hashing algorithm, it can be ensured that the data does not need to be re-positioned after being pulled and can be directly analyzed, greatly improving the analysis efficiency. It solves the technical problems in the prior art that the coupling degree between the analytical computing resources and the transaction processing computing resources in the database engine is too high, the analytical resources cannot be shared, and it is difficult to scale the two types of resources independently. The present invention can support the separate deployment of transaction data and analytical resources, share the analytical resources among multiple transaction data clusters, and support independent scaling; dynamically allocate analytical resources for tasks and improve resource utilization rate.
[0108] Optionally, the device further includes: a data receiving module, a data analysis module, and a data reply module; wherein, The data receiving module is configured to receive at least one piece of data to be analyzed sent by at least one transaction data engine through each of the analysis data engines; The data analysis module is configured to, when each of the analysis data engines has received a data sending completion flag sent by all the transaction data engines, perform data analysis on at least one piece of data to be analyzed based on the data analysis task to obtain a data analysis result; The data reply module is configured to return the data analysis result to at least one transaction data cluster corresponding to the data analysis request.
[0109] The analysis device of the database for sharing analysis resources provided by the embodiment of the present invention can execute the analysis method of the database for sharing analysis resources provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0110] Figure 7 It is a schematic structural diagram of an analysis system of a database for sharing analysis resources provided by an embodiment of the present invention. As Figure 7 shown: it includes at least one transaction data cluster and one analysis data cluster; the transaction data cluster includes at least one transaction data engine, and the analysis data cluster includes at least one analysis coordinator and at least one analysis data engine; wherein, The transaction analysis cluster includes transaction analysis clusters 701, 702, and 703; the transaction analysis cluster 1 includes 5 transaction analysis engines, the transaction analysis cluster 2 includes 3 transaction analysis engines, and the transaction analysis cluster 3 includes 2 transaction analysis engines; each transaction data cluster can execute the database access method provided by any embodiment of the present invention.
[0111] The analysis data cluster includes an analysis coordinator and 4 analysis data engines; the analysis data cluster can execute the database access method provided by any embodiment of the present invention.
[0112] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and this is not limited herein.
[0113] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An analysis method for a database that shares analysis resources, characterized in that, Applied to at least one transaction data cluster; including: Based on at least one transaction data engine, analyze the analytical statement text input by the user in response to data analysis, and generate a task identifier corresponding to the analytical statement text; For each of the transaction data engines, query the subset of data to be analyzed corresponding to the analytical statement text; store the subset of data to be analyzed in the transaction result pool; the subset of data to be analyzed includes at least one piece of data to be analyzed; Generate a data analysis request based on the data storage address and the task identifier of the transaction result pool of each of the transaction data engines, and send the data analysis request to the analysis data cluster; Receive the data analysis result of the data analysis request returned by the analysis data cluster, and display the data analysis result to the data analysis user.
2. The method according to claim 1, wherein The transaction result pool includes at least one storage data slot; The storing the subset of data to be analyzed in the transaction result pool includes: For each piece of data to be analyzed, identify the target data slot corresponding to the data to be analyzed in the storage data slot; Sequentially store each piece of data to be analyzed in the target data slot of each piece of data to be analyzed.
3. The method according to claim 2, wherein The identifying the target data slot corresponding to the data to be analyzed in the storage data slot includes: Perform a hash calculation on the data to be analyzed through a preset hash function to obtain a hash calculation value corresponding to the data to be analyzed; Obtain the data slot number of each storage data slot, and determine the storage data slot with the same data slot number as the hash calculation value as the target data slot of the data to be analyzed.
4. The method according to claim 3, wherein Before receiving the data analysis result of the data analysis request returned by the analysis data cluster based on the data analysis request and displaying the data analysis result to the data analysis user, it includes: For each of the transaction data engines, respond to and process at least one data pull request received, and determine the working number and the number of engines allocated for the analysis data engine corresponding to each data pull request; Allocate an analysis data engine for each storage data slot based on the working number and the number of engines allocated for the analysis data engine, and send the data to be analyzed in the storage data slot to the analysis data engine corresponding to the storage data slot in the form of a data stream.
5. The method according to claim 4, characterized in that, The allocating an analysis data engine for each storage data slot based on the working number and the number of engines allocated for the analysis data engine includes: Perform a modulo calculation on each of the data slot numbers and the number of engines allocated respectively to determine the data pull number corresponding to each storage data slot; Allocate an analysis data engine for each storage data slot based on the data pull number and the working number.
6. An analysis method for a database that shares analysis resources, characterized in that, Applied to the analysis data cluster; including: Based on the analysis coordinator of the analysis data cluster, receive and process the data analysis request, and allocate at least one analysis data engine for the data analysis request; Generate a data analysis communication message based on at least one data storage address and a task identifier, and send the data analysis communication message to each of the analysis data engines respectively; For each of the analysis data engines, a data pull request is generated based on the working number of the analysis data engine and the number of engine allocations of the analysis data engine corresponding to the data analysis request, and the data pull request is sent to each of the data storage addresses.
7. The method according to claim 6, wherein Further included are: Receiving, by each of the analysis data engines, at least one piece of data to be analyzed sent by at least one transaction data engine; When each of the analysis data engines has received the data transmission completion identifier sent by all of the transaction data engines, performing data analysis on at least one piece of the data to be analyzed based on the data analysis task to obtain a data analysis result; Returning the data analysis result to at least one transaction data cluster corresponding to the data analysis request.
8. An analysis device for a database that shares analysis resources, characterized in that, Deployed in at least one transaction data cluster; including: A task response module, configured to generate a task identifier corresponding to the analytical statement text based on at least one transaction data engine responding to the analytical statement text input by the data analysis user; A query module, configured to, for each of the transaction data engines, query a subset of the data to be analyzed corresponding to the analytical statement text; storing the subset of the data to be analyzed in a transaction result pool; the subset of the data to be analyzed includes at least one piece of data to be analyzed; A communication module, configured to generate a data analysis request based on the data storage address of the transaction result pool of each of the transaction data engines and the task identifier, and send the data analysis request to the analysis data cluster; A display module, configured to receive the data analysis result of the data analysis request returned by the analysis data cluster, and display the data analysis result to the data analysis user.
9. An analysis device for a database that shares analysis resources, characterized in that, Deployed in the analysis data cluster; including: An analysis coordination module, configured to receive and process the data analysis request based on the analysis coordinator of the analysis data cluster, and allocate at least one analysis data engine for the data analysis request; An allocation module, configured to generate a data analysis communication message based on at least one data storage address and the task identifier, and send the data analysis communication message to each of the analysis data engines; An analysis execution module, configured to, for each of the analysis data engines, generate a data pull request based on the working number of the analysis data engine and the number of engine allocations of the analysis data engine corresponding to the data analysis request, and send the data pull request to each of the data storage addresses.
10. An analysis system for a database that shares analysis resources, characterized in that, Including at least one transaction data cluster and an analysis data cluster, the transaction data cluster includes at least one transaction data engine, and the analysis data cluster includes at least one analysis coordinator and at least one analysis data engine; wherein, The transaction data cluster is used to execute the analysis method of the database for sharing analysis resources according to any one of claims 1-5; The analysis data cluster is used to execute the analysis method of the database for sharing analysis resources according to any one of claims 6-7.
Citation Information
Patent Citations
Data query method and device, computer equipment and storage medium
CN112650759A
Cross-multi-engine routing processing method and device, equipment and storage medium
CN114020446A
Data blood relationship analysis system and analysis method
CN118467175A
Heterogeneous database processing archetypes for hybrid system
US20160098425A1
Adaptive disk spill based on memory usage of shared storage
US20240403292A1