Data query method, device, equipment and readable storage medium
By constructing query conditions and determining the number of concurrent nodes, the tree-like data structure is queried in parallel layer by layer, which solves the problem of low data query efficiency and realizes rapid query analysis of data.
Patent Information
- Application Number
- CN202010742022.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-07-28
AI Technical Summary
In the prior art, data query efficiency is low, which affects the effectiveness of data analysis.
By constructing query conditions and determining the number of concurrent nodes, the tree data structure is queried in parallel layer by layer, and the query conditions and number of concurrent nodes are updated until all data layers are queried.
Improve data query efficiency and promote rapid data query analysis.
Smart Images

Figure CN114003616B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis, and in particular to a data query method, device, equipment and readable storage medium. Background Art
[0002] With the development of information technology, all walks of life pay more and more attention to the digital and indicator processing of information. Various types of information are queried and analyzed in the form of data to obtain various indicators in various dimensions, reflecting the development of all walks of life in an all-round way.
[0003] The amount of basic data for data query analysis is usually huge, and the efficiency of data query will inevitably affect data analysis. The higher the query efficiency, the more conducive it is to analysis, and vice versa. Therefore, how to improve data query efficiency is a technical problem that needs to be solved urgently. Summary of the invention
[0004] The main purpose of the present invention is to provide a data query method, device, equipment and readable storage medium, aiming to solve the technical problem of how to improve data query efficiency in the prior art.
[0005] To achieve the above object, the present invention provides a data query method, which comprises the following steps:
[0006] Construct query conditions and determine the number of concurrent nodes;
[0007] According to the query condition and the number of concurrent nodes, query the data layer corresponding to the query condition in the tree data structure to obtain a query result;
[0008] According to the query result, the query condition and the number of concurrent nodes are updated, and based on the updated query condition and the number of concurrent nodes, the step of querying the data layer corresponding to the query condition in the tree data structure according to the query condition and the number of concurrent nodes is executed, until all data layers in the tree data structure are queried and the final query result is obtained.
[0009] Optionally, the step of constructing the query condition and determining the number of concurrent nodes includes:
[0010] Determine whether there is a cache data layer, and if so, build query conditions based on the cache data layer and determine the number of concurrent nodes;
[0011] If the cache data layer does not exist, the query conditions are constructed according to the received dimension keywords, and the number of concurrent nodes is determined.
[0012] Optionally, the step of constructing a query condition according to the cache data layer and determining the number of concurrent nodes includes:
[0013] Obtain the target data layer at the next level below the cache data layer in the tree - like data structure, and the data volume of the cache result corresponding to the cache data layer;
[0014] Construct the query condition according to the target data layer, and determine the number of concurrent nodes according to the total number of queries to be made and the data volume.
[0015] Optionally, before the step of obtaining the target data layer at the next level below the cache data layer in the tree - like data structure, it includes:
[0016] Judge whether each data layer in the tree - like data structure carries a query identifier. If all carry the query identifier, use the cache data corresponding to the cache data layer as the final query result and end the query;
[0017] If not all carry the query identifier, execute the step of obtaining the target data layer at the next level below the cache data layer in the tree - like data structure.
[0018] Optionally, the step of constructing the query condition according to the received dimensional keywords and determining the number of concurrent nodes includes:
[0019] Determine the trial level according to the number of received dimensional keywords, and construct the query condition according to the dimensional keywords and the trial level;
[0020] Initiate a trial query on the tree - like data structure according to the query condition and the trial level to obtain a trial query result;
[0021] Determine the number of concurrent nodes according to the total number of queries to be made and the result quantity of the trial query result.
[0022] Optionally, after the step of querying the data layer corresponding to the query condition in the tree - like data structure to obtain a query result, it includes:
[0023] Cache the data layer corresponding to the query condition in the tree - like data structure and the query result in the form of key - value pairs.
[0024] Optionally, before the step of updating the query condition and the number of concurrent nodes according to the query result, it includes:
[0025] Judge whether the cached data volume corresponding to the cached query result is greater than or equal to the total number of queries to be made. If it is greater than or equal to the total number of queries to be made, use the cached query result as the final query result and end the query;
[0026] If the amount of cached data is less than the total number of queries to be made, perform the step of updating the query condition and the number of concurrent nodes according to the query result.
[0027] Optionally, after the step of querying the data layer corresponding to the query condition in the tree - like data structure to obtain a query result, the following steps are included:
[0028] Add a query identifier to the data layer corresponding to the query condition in the tree - like data structure.
[0029] Optionally, after the step of querying the data layer corresponding to the query condition in the tree - like data structure to obtain a query result, the following steps are included:
[0030] Add a result identifier to the query result according to the query condition, the dimension keyword corresponding to the query condition, and the query keyword.
[0031] Optionally, after the step of querying the data layer corresponding to the query condition in the tree - like data structure to obtain a query result, the following steps are included:
[0032] Add a count identifier to the data layer corresponding to the query condition in the tree - like data structure.
[0033] Optionally, before the step of constructing the query condition and determining the number of concurrent nodes, the following steps are included:
[0034] Obtain multiple pieces of data that support querying, and clean the multiple pieces of data to generate valid data;
[0035] Construct the valid data into the tree - like data structure.
[0036] Furthermore, to achieve the above object, the present invention also provides a data query device, and the data query device includes:
[0037] A construction module, configured to construct a query condition and determine the number of concurrent nodes;
[0038] A query module, configured to query the data layer corresponding to the query condition in the tree - like data structure according to the query condition and the number of concurrent nodes to obtain a query result;
[0039] An update module, configured to update the query condition and the number of concurrent nodes according to the query result, and based on the updated query condition and the number of concurrent nodes, perform the step of querying the data layer corresponding to the query condition in the tree - like data structure according to the query condition and the number of concurrent nodes until each data layer in the tree - like data structure has been queried to obtain a final query result.
[0040] Optionally, the building block further includes:
[0041] A judgment unit for judging whether there is a cache data layer. If there is the cache data layer, query conditions are constructed according to the cache data layer, and the number of concurrent nodes is determined;
[0042] A construction unit for constructing query conditions according to the received dimensional keywords and determining the number of concurrent nodes if there is no such cache data layer.
[0043] Optionally, the judgment unit is further configured to:
[0044] Obtain the target data layer at the next level of the cache data layer in the tree-like data structure, and the data volume of the cache result corresponding to the cache data layer;
[0045] Construct the query conditions according to the target data layer, and determine the number of concurrent nodes according to the total number of queries to be queried and the data volume.
[0046] Optionally, the judgment unit is further configured to:
[0047] Judge whether each data layer in the tree-like data structure carries a query identifier. If all carry the query identifier, use the cache data corresponding to the cache data layer as the final query result and end the query;
[0048] If not all carry the query identifier, execute the step of obtaining the target data layer at the next level of the cache data layer in the tree-like data structure.
[0049] Optionally, the construction unit is further configured to:
[0050] Determine a tentative level according to the number of received dimensional keywords, and construct query conditions according to the dimensional keywords and the tentative level;
[0051] Initiate a tentative query on the tree-like data structure according to the query conditions and the tentative level, and obtain a tentative query result;
[0052] Determine the number of concurrent nodes according to the total number of queries to be queried and the result quantity of the tentative query result.
[0053] Optionally, the data query device further includes:
[0054] A cache module for caching the data layer corresponding to the query conditions in the tree-like data structure and the query result in the form of key-value pairs.
[0055] Optionally, the data query device further includes:
[0056] A judgment module, configured to judge whether the amount of cached data corresponding to the cached query result is greater than or equal to the total number of queries to be made. If it is greater than or equal to the total number of queries to be made, the cached query result is used as the final query result, and the query is ended.
[0057] An execution module, configured to, if the amount of cached data is less than the total number of queries to be made, execute the steps of updating the query conditions and the number of concurrent nodes according to the query result.
[0058] Furthermore, to achieve the above object, the present invention also provides a data query device, which includes a memory, a processor, and a data query program stored on the memory and executable on the processor. When the data query program is executed by the processor, the steps of the data query method as described above are implemented.
[0059] Furthermore, to achieve the above object, the present invention also provides a readable storage medium, on which a data query program is stored. When the data query program is executed by a processor, the steps of the data query method as described above are implemented.
[0060] In the data query method, device, equipment, and readable storage medium of the present invention, first, query conditions are constructed and the number of concurrent nodes is determined; then, according to the query conditions and the number of concurrent nodes, the data layer corresponding to the query conditions in the tree-like data structure is queried to obtain a query result; furthermore, according to the query result, the query conditions and the number of concurrent nodes are updated, and based on the updated query conditions and the number of concurrent nodes, the data layer corresponding to the query conditions in the tree-like data structure is queried again, and so on in a loop to update the query until each data layer in the tree-like data structure has been queried to obtain the final query result. By updating the query conditions and the number of concurrent nodes, each data layer in the data structure is queried layer by layer in a parallel manner, and multiple data nodes are simultaneously and parallelly queried for each data layer, improving the data query efficiency and facilitating the rapid query and analysis of data. Description of the Drawings
[0061] Figure 1 It is a schematic structural diagram of the hardware operating environment related to the embodiment of the data query device of the present invention;
[0062] Figure 2 It is a schematic flowchart of the first embodiment of the data query method of the present invention;
[0063] Figure 3 It is a schematic diagram of the functional modules of the preferred embodiment of the data query device of the present invention;
[0064] Figure 4 It is a schematic diagram of the tree-like data structure in the data query method of the present invention.
[0065] The realization, functional features, and advantages of the present invention will be further described in conjunction with embodiments with reference to the accompanying drawings. Detailed implementation manners
[0066] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0067] The present invention provides a data query device. Refer to Figure 1 , Figure 1 which is a schematic structural diagram of the hardware operating environment involved in the embodiment solution of the data query device of the present invention.
[0068] As Figure 1 shown, the data query device may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage data query device independent of the aforementioned processor 1001.
[0069] Those skilled in the art can understand that Figure 1 the hardware structure of the data query device shown in
[0070] As Figure 1 shown, in the memory 1005, as a readable storage medium, there may be included an operating system, a network communication module, a user interface module, and a data query program. Among them, the operating system is a program for managing and controlling the hardware and software resources of the data query device, and supports the operation of the network communication module, the user interface module, the data query program, and other programs or software; the network communication module is used to manage and control the network interface 1004; the user interface module is used to manage and control the user interface 1003.
[0071] In Figure 1In the shown hardware structure of the data query device, the network interface 1004 is mainly used to connect to the background server and conduct data communication with the background server; the user interface 1003 is mainly used to connect to the client (user side) and conduct data communication with the client; the processor 1001 can call the data query program stored in the memory 1005 and perform the following operations:
[0072] Construct a query condition and determine the number of concurrent nodes;
[0073] According to the query condition and the number of concurrent nodes, query the data layer corresponding to the query condition in the tree - like data structure to obtain a query result;
[0074] According to the query result, update the query condition and the number of concurrent nodes, and based on the updated query condition and the number of concurrent nodes, execute the step of querying the data layer corresponding to the query condition in the tree - like data structure according to the query condition and the number of concurrent nodes until all data layers in the tree - like data structure have been queried to obtain the final query result.
[0075] Further, the step of constructing the query condition and determining the number of concurrent nodes includes:
[0076] Judge whether there is a cached data layer. If there is the cached data layer, construct a query condition according to the cached data layer and determine the number of concurrent nodes;
[0077] If there is no such cached data layer, construct a query condition according to the received dimensional keywords and determine the number of concurrent nodes.
[0078] Further, the step of constructing a query condition according to the cached data layer and determining the number of concurrent nodes includes:
[0079] Obtain the target data layer in the tree - like data structure at the next level below the cached data layer, and the data volume of the cached result corresponding to the cached data layer;
[0080] According to the target data layer, construct the query condition, and determine the number of concurrent nodes according to the total number to be queried and the data volume.
[0081] Further, before the step of obtaining the target data layer in the tree - like data structure at the next level below the cached data layer, the processor 1001 can call the data query program stored in the memory 1005 and perform the following operations:
[0082] Judge whether each data layer in the tree - like data structure carries a query identifier. If all carry the query identifier, use the cached data corresponding to the cached data layer as the final query result and end the query;
[0083] If none of them carry the query identifier, perform the step of obtaining the target data layer at the next level below the cache data layer in the tree - like data structure.
[0084] Further, the steps of constructing a query condition according to the received dimensional keywords and determining the number of concurrent nodes include:
[0085] Determine a tentative level according to the number of received dimensional keywords, and construct a query condition according to the dimensional keywords and the tentative level;
[0086] Initiate a tentative query on the tree - like data structure according to the query condition and the tentative level, and obtain a tentative query result;
[0087] Determine the number of concurrent nodes according to the total number to be queried and the number of results in the tentative query result.
[0088] Further, after the step of querying the data layer corresponding to the query condition in the tree - like data structure to obtain a query result, the processor 1001 may call the data query program stored in the memory 1005 and perform the following operations:
[0089] Cache the data layer corresponding to the query condition in the tree - like data structure and the query result in the form of key - value pairs.
[0090] Further, before the step of updating the query condition and the number of concurrent nodes according to the query result, the processor 1001 may call the data query program stored in the memory 1005 and perform the following operations:
[0091] Judge whether the cached data volume corresponding to the cached query result is greater than or equal to the total number to be queried. If it is greater than or equal to the total number to be queried, use the cached query result as the final query result and end the query;
[0092] If the cached data volume is less than the total number to be queried, perform the step of updating the query condition and the number of concurrent nodes according to the query result.
[0093] Further, after the step of querying the data layer corresponding to the query condition in the tree - like data structure to obtain a query result, the processor 1001 may call the data query program stored in the memory 1005 and perform the following operations:
[0094] Add a query identifier to the data layer corresponding to the query condition in the tree - like data structure.
[0095] Further, after the step of querying the data layer corresponding to the query condition in the tree - shaped data structure to obtain a query result, the processor 1001 may call the data query program stored in the memory 1005 and perform the following operations:
[0096] According to the query condition, as well as the dimension keyword and the query keyword corresponding to the query condition, add a result identifier to the query result.
[0097] Further, after the step of querying the data layer corresponding to the query condition in the tree - shaped data structure to obtain a query result, the processor 1001 may call the data query program stored in the memory 1005 and perform the following operations:
[0098] Add a count identifier to the data layer corresponding to the query condition in the tree - shaped data structure.
[0099] Further, before the step of constructing the query condition and determining the number of concurrent nodes, the processor 1001 may call the data query program stored in the memory 1005 and perform the following operations:
[0100] Obtain multiple pieces of data that support querying, and clean the multiple pieces of data to generate valid data;
[0101] Construct the valid data into the tree - shaped data structure.
[0102] The specific implementation manner of the data query device of the present invention is basically the same as that of each embodiment of the following data query method, and will not be elaborated here.
[0103] The present invention also provides a data query method.
[0104] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the data query method of the present invention.
[0105] The embodiments of the present invention provide embodiments of the data query method. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here. Specifically, in the data query method of this embodiment, the data query method includes:
[0106] Step S10, construct a query condition and determine the number of concurrent nodes;
[0107] The data query method in this embodiment is applied to a data analysis system, and is suitable for constructing query conditions through a data analysis system to query the data therein. The analysis system may only include multiple data indicators of a certain aspect, such as sales data, growth data, etc. of a certain enterprise, or gender data, employment data, etc. of personnel in a certain region, or may include multiple data indicators of multiple aspects, such as sales data, growth data of multiple enterprises, or gender data, employment data, etc. of personnel in multiple regions. Whether it is multiple aspects or a single aspect, each aspect of data exists in the form of a tree data structure, and the data of one aspect forms a tree data structure. The tree data structure contains multiple hierarchical nodes, and each hierarchical node forms a dimension under the aspect represented by the tree data structure, such as the hierarchical nodes formed by various departments under an enterprise, or the hierarchical nodes formed by various cities and counties under a region. This implementation constructs query conditions through various dimensions, and determines the number of concurrent nodes to realize layer-by-layer query of the tree data structure. Among them, the number of concurrent nodes represents the number of nodes that are queried simultaneously on the tree data structure.
[0108] Furthermore, the step of constructing query conditions and determining the number of concurrent nodes includes:
[0109] Step a1, obtaining multiple data supporting the query, and cleaning the multiple data to generate valid data;
[0110] Step a2, constructing the valid data into the tree data structure.
[0111] Understandably, each industry has a lot of data on each indicator, and each data includes data that supports query and data that does not support query. For the data that supports query, a tree data structure is constructed for query; for the data that does not support query, it can be added to the tree data structure or not. Among them, different query permissions can be set for the data that does not support query added to the tree data structure, so as to limit the query of the data that does not support query through different query permissions.
[0112] Furthermore, the data supporting the query may contain noise data, such as duplicate data, invalid data that is difficult to identify, etc. After obtaining multiple data supporting the query, this type of data is identified, and the identification results are removed from the multiple data to clean the multiple data, generate valid data, and then construct the valid data into a tree data structure. Among them, because the tree data structure contains various levels, the level where the valid data is located is first identified during construction, and then added to the corresponding level to form a tree data structure. The specific tree data structure can be referred to Figure 4 , Figure 4In the tree - like data structure, root is the root node, x1 and x2 are nodes at the first level, and y1, y1, and y3 are nodes at the second level. A node at a level represents a dimension, and each dimension corresponds to its own data index. For example, x1 and x2 represent provincial - level regions and have their own employment - number indexes. y1, y2, and y3 represent municipal - level regions under the provincial region x2 and also have their own employment - number indexes.
[0113] Step S20: Query the data layer corresponding to the query condition in the tree - like data structure according to the query condition and the number of concurrent nodes to obtain a query result.
[0114] Furthermore, after constructing the query condition and determining the number of concurrent nodes, query the data layer corresponding to the query condition in the tree - like data structure according to the query condition and the number of concurrent nodes to obtain a query result. Among them, the query condition varies according to the different dimensions to be queried, and the tree - like data structure is queried in a hierarchical manner. If there is a need to query the first level as described above Figure 4 in the first level, the query condition includes keywords of the first - level dimension, such as provincial - level keywords. And the keywords included in the query condition can be divided into basic keywords and filtering keywords. For example, if querying without distinguishing specific provinces, the provincial - level keywords are formed as basic keywords. If querying by distinguishing specific provinces, in addition to forming the provincial - level keywords as basic keywords, the specific provinces are also formed as filtering keywords to filter out the data of specific provinces from the data obtained by querying with the basic keywords as the query result. At the same time, for the data layer corresponding to the query condition, multiple nodes located in this data layer are queried simultaneously according to the number of concurrent nodes. For example, if the data layer is the first layer and the number of concurrent nodes is 2, then x1 and x2 are queried simultaneously to obtain a query result.
[0115] It can be understood that for the hierarchical query mechanism of the tree - like data structure, after the first query, some nodes in the tree - like data structure have been queried and some have not; and the results obtained from each query are different. Therefore, for the sake of easy distinction, this embodiment sets up a mechanism for adding identifiers to the query results obtained for each layer of query. Specifically, after the step of querying the data layer corresponding to the query condition in the tree - like data structure to obtain a query result, it includes:
[0116] Step b1: Add a query identifier to the data layer corresponding to the query condition in the tree - like data structure.
[0117] Step b2: Add a result identifier to the query result according to the query condition, the dimension keywords corresponding to the query condition, and the query keywords.
[0118] Step b3, add a counting identifier to the data layer in the tree - like data structure corresponding to the query condition.
[0119] Specifically, after querying the data layer in the tree - like data structure corresponding to the query condition to obtain a query result, add a query identifier to this data layer to indicate that this data layer has been queried and distinguish it from the data layers that have not been queried. At the same time, according to the current query condition, the dimension keyword corresponding to the query condition, and the query keyword, add a result identifier to the query result to distinguish the query results. The dimension keyword corresponding to the query condition represents the data layer targeted by the current query. For example, querying the y1 layer in Figure 4 ; the query keyword corresponding to the query condition represents the data index targeted by the current query, such as querying the sales volume, etc.
[0120] Furthermore, to ensure that all nodes in each data layer are queried, in addition to adding a query identifier to the data layer in this embodiment, a counting identifier is also added. The query situation of the nodes and their child nodes corresponding to the data layer is reflected by the counting identifier. When only querying the node itself corresponding to the data layer, add a counting identifier with a value of 1 to the node corresponding to the data layer. When all child nodes under the corresponding node have been queried, add a corresponding numerical counting identifier to the corresponding node according to the number of child nodes. For example, Figure 4 for the node x2 in , if only x2 has been queried, its counting identifier is 1. If all its child nodes y1, y2, and y3 have been queried, the counting identifier of x2 is 4, and the counting identifiers of y1, y2, and y3 are all 1.
[0121] Step S30, update the query condition and the number of concurrent nodes according to the query result, and based on the updated query condition and the number of concurrent nodes, execute the step of querying the data layer in the tree - like data structure corresponding to the query condition according to the query condition and the number of concurrent nodes until all data layers in the tree - like data structure have been queried to obtain the final query result.
[0122] Furthermore, after obtaining the query result through query, the query condition and the number of concurrent nodes are updated according to the query result. It is judged whether the data volume of the query result obtained through query reaches the requirement of the data volume to be queried. If it reaches the requirement, the query is stopped; if it does not reach the requirement, the data layer in the tree - shaped data structure that has not been queried is determined, and the dimension corresponding to the data layer that has not been queried is added to the query condition to update the query condition, so as to continue querying the data layer that has not been queried through the updated query condition. Meanwhile, the data volume of the subsequent query is estimated according to the data volume of the query result, and the difference between the data volume of the query result and the data volume to be queried is determined; furthermore, the number of concurrent nodes is updated according to the difference and the estimated data volume, so as to query multiple nodes that have not been queried in the tree - shaped data structure simultaneously through the updated number of concurrent nodes.
[0123] Further, after updating the query condition and the number of concurrent nodes, the data layer corresponding to the query condition in the tree - shaped data structure is queried according to the updated query condition and the number of concurrent nodes; that is, according to the number of concurrent nodes, the multiple nodes that have not been queried represented by the updated query condition are continuously queried. In this way, the query result is obtained to update the query condition and the number of concurrent nodes again, and based on the update, a new query result is obtained through continuous query, and so on in a cycle until all nodes have been queried. Among them, a query identifier is added to the data layer being queried each time a query is made. After query identifiers are added to the nodes of each data layer in the tree - shaped data structure, it means that all nodes have been queried, and then the query is stopped. Furthermore, the query result obtained previously is used as the final query result to complete the layer - by - layer parallel query of the tree - shaped data structure.
[0124] The data query method of the present invention first constructs a query condition and determines the number of concurrent nodes; then, according to the query condition and the number of concurrent nodes, the data layer corresponding to the query condition in the tree - shaped data structure is queried to obtain a query result; furthermore, according to the query result, the query condition and the number of concurrent nodes are updated, and based on the updated query condition and the number of concurrent nodes, the data layer corresponding to the query condition in the tree - shaped data structure is queried again, and so on in a cycle to update the query until all data layers in the tree - shaped data structure have been queried to obtain the final query result. By updating the query condition and the number of concurrent nodes, each data layer in the data structure is queried layer - by - layer in a parallel manner, and multiple data nodes are queried in parallel at the same time for each data layer, which improves the data query efficiency and is conducive to the rapid query and analysis of data.
[0125] Further, based on the first embodiment of the data query method of the present invention, a second embodiment of the data query method of the present invention is proposed.
[0126] The difference between the second embodiment of the data query method and the first embodiment of the data query method is that the step of constructing the query condition and determining the number of concurrent nodes includes:
[0127] Step S11: Determine whether there is a cached data layer. If there is such a cached data layer, construct a query condition based on the cached data layer and determine the number of concurrent nodes.
[0128] In this embodiment, the query condition has different construction methods depending on whether cached data is obtained through the query, and the number of concurrent nodes is determined based on the query result obtained through the query. Specifically, when there is a query requirement, enter the dimensional keyword representing the query requirement in the query box of the analysis system. The analysis system constructs a query condition based on the dimensional keyword for querying, obtains a query result, and caches the data layer information targeted by this query. Thereafter, update the constructed query condition based on the cache to continue the query. Therefore, in this embodiment, the query condition used for each query is different depending on whether it is cached. Therefore, when constructing the query condition, first determine whether there is a cached data layer, which includes the query result obtained through the query, the data layer targeted by the previous query, and other content. If it is determined that there is a cached data layer in the analysis system, update the query condition and determine the number of concurrent nodes based on the cached data layer. Among them, the updated query condition is the query condition constructed for subsequent queries. Specifically, the steps of constructing the query condition based on the cached data layer and determining the number of concurrent nodes include:
[0129] Step S111: Obtain the target data layer at the next level of the cached data layer in the tree - like data structure, and the data volume of the cache result corresponding to the cached data layer.
[0130] Step S112: Construct the query condition based on the target data layer, and determine the number of concurrent nodes according to the total number of items to be queried and the data volume.
[0131] Furthermore, the cached data layer includes the data layer targeted by the previous query. Based on this, obtain the data layer at the next level of the targeted data layer in the tree - like data structure as the target data layer. At the same time, obtain the query result included in the cached data layer as the cache result corresponding to the cached data layer, and count the data volume of the query result to obtain the data volume of the cache result. Thereafter, construct a query condition based on the target data layer, and add the keyword representing the dimension of the target data layer to the query condition of the previous query to construct a new query condition.
[0132] In addition, the number of concurrent nodes is determined based on the total number of queries to be made and the amount of data. Here, the total number of queries to be made represents the quantity required to be obtained from the query, and the amount of data is the quantity obtained from the previous query. The amount of data obtained from the previous query and the amount of data obtained from previous queries and cached are accumulated to obtain the amount of data that has been queried. Then, by comparing the total number of queries to be made with the amount of data that has been queried, the difference between the two is obtained, which represents the amount of data that still needs to be queried. Thereafter, based on the proportional relationship between the amount of data that each node can obtain from the query represented by the amount of data that has been queried and the amount of data that needs to be queried, the number of nodes that need to be queried is determined, and this number of nodes is the number of concurrent nodes. In a specific embodiment, the number of each node that has been queried is p, the amount of data that has been queried is R, and the total number of queries to be made is D. Then, the number of concurrent nodes m for the next simultaneous query is m = (D - R) / (R / p). If the calculated number of concurrent nodes m is a non-integer, then m is taken as the integer greater than the calculated m, and the integer taken has an adjacent relationship with the calculated value; for example, if it is calculated that m = 2.2, then m is taken as 3.
[0133] Understandably, as the queries are made layer by layer, each node in the tree-like data structure is queried in sequence until all nodes have been queried. Therefore, before obtaining the target data layer at the next level below the cached data layer, it is first determined whether all nodes in the tree-like data structure have been queried. If all have been queried, there is no target data layer; otherwise, there is a target data layer. Specifically, before the step of obtaining the target data layer at the next level below the cached data layer in the tree-like data structure, it includes:
[0134] Step c1, determining whether each data layer in the tree-like data structure carries a query identifier. If all carry a query identifier, then the cached data corresponding to the cached data layer is used as the final query result, and the query is ended;
[0135] Step c2, if not all carry a query identifier, then execute the step of obtaining the target data layer at the next level below the cached data layer in the tree-like data structure.
[0136] Furthermore, for each node in the tree-like data structure that has been queried, a query identifier is set. Therefore, it can be determined whether all nodes have been queried by determining whether each node carries a query identifier. If it is determined that all data layers carry a query identifier, it indicates that all nodes in the tree-like data structure have been queried. At this time, the cached data corresponding to the cached data layer forms the final query result, and the query operation is ended. Among them, the cached data corresponding to the cached data layer includes, in addition to the cached result obtained from the previous query, the query results of each previous query, and is the cache of the query results of each query.
[0137] Further, if it is determined that not all data layers carry query identifiers, indicating that there are data layers that have not been queried, obtain the data layer at the next level below the cached data layer in the tree - like data structure as the target data layer, construct a query condition to continue the query until all data layers in the tree - like data structure have been queried, and obtain the final query result.
[0138] Step S12, if there is no such cached data layer, construct a query condition according to the received dimensional keywords and determine the number of concurrent nodes.
[0139] Further, for the case where it is determined that there is no cached data layer in the analysis system, initiate a trial query through the received dimensional keywords. From the results of the trial query, construct a query condition and determine the number of concurrent nodes. The trial query determines the trial data layer based on the number of input dimensional keywords, initiates a query for the trial data layer to explore the amount of data that can be obtained from querying this data layer. From this amount of data, estimate the data layer targeted by the subsequent query, and then construct a query condition from the data layer and determine the number of concurrent nodes. Specifically, the steps of constructing a query condition according to the received dimensional keywords and determining the number of concurrent nodes include:
[0140] Step S121, determine the trial level according to the number of received dimensional keywords, and construct a query condition according to the dimensional keywords and the trial level;
[0141] Step S122, initiate a trial query for the tree - like data structure according to the query condition and the trial level, and obtain the trial query result;
[0142] Step S123, determine the number of concurrent nodes according to the total number to be queried and the number of results of the trial query result.
[0143] Furthermore, when the number of dimensional keywords is different, the dimensions to be queried are different, the data layer targeted by the query is different, and the amount of data in the query result obtained can also be different. Therefore, in order to determine the amount of data that each node can query, determine the trial level based on the number of dimensional keywords. When the number of dimensional keywords is larger and the data layer targeted by the query is deeper, the corresponding trial level is also deeper. For example, if the dimensions to be queried are 2 - 5 dimensions and the number of dimensional keywords is 2 - 5, the first layer in the tree - like data structure can be used as the trial level; as the number of dimensions increases, the selected trial level increases accordingly, which can be the second layer or the third layer, etc.
[0144] Further, according to the dimension keywords and the probing level, query conditions are constructed, and the keywords corresponding to the probing level in each dimension keyword are constructed into query conditions. A probing query is initiated for the probing level through the query conditions, and a probing query result is obtained. Thereafter, according to the proportional relationship between the number of results of the probing query and the total number of queries to be made, the number of concurrent nodes is determined. In a specific embodiment, if the number of results obtained from the probing query is R and the total number of queries to be made is D, then the number of concurrent nodes m for simultaneous query is m=(D - R) / R.
[0145] It should be noted that the probing query can be performed on a single node or on multiple nodes. For a single node, the number of results of the probing query represents the amount of data that can be obtained by querying a single node. Then, according to the proportional relationship between the total number of queries to be made and the number of results, the number of concurrent nodes is determined. For multiple nodes, the number of results of the probing query represents the amount of data that can be obtained by querying multiple nodes. First, according to the proportional relationship between the number of results and the multiple nodes, the average amount of data per single node is determined. Then, according to the proportional relationship between the total number of queries to be made and the average amount of data, the number of concurrent nodes can be determined.
[0146] This embodiment constructs query conditions and determines the number of concurrent nodes in different ways based on the existence of a cache data layer. If there is a cache data layer, query conditions are constructed and the number of concurrent nodes is determined based on the cache data layer, so that the constructed query conditions and the determined number of concurrent nodes are related to the data layer cached in the most recent query, and the constructed query conditions and the determined number of concurrent nodes are more accurate. If there is no cache data layer, query conditions are constructed based on the dimension keywords to initiate a probing query, and the number of concurrent nodes is determined by the amount of data that can be obtained by probing each node; this improves the accuracy of the determined number of concurrent nodes and is conducive to achieving accurate queries based on the number of concurrent nodes.
[0147] Further, based on the first or second embodiment of the data query method of the present invention, a third embodiment of the data query method of the present invention is proposed.
[0148] The difference between the third embodiment of the data query method and the first or second embodiment of the data query method is that after the step of querying the data layer corresponding to the query conditions in the tree - like data structure to obtain a query result, the following steps are included:
[0149] Step d1, caching the data layer corresponding to the query conditions in the tree - like data structure and the query result in the form of key - value pairs.
[0150] In this embodiment, it is determined whether to end the query by judging whether the data volume obtained from each query meets the total number of queries to be queried. Specifically, after the query result is obtained each time, the data layer corresponding to the query condition in the tree-like data structure, that is, the data layer targeted by this query, and the query result obtained from the query are cached in the form of key-value pairs. The query results obtained from each query are quickly determined by the key-value pairs, and whether to end the query is determined by the size relationship between the data volume of each query result and the total number of queries to be queried.
[0151] Further, before the step of updating the query condition and the number of concurrent nodes according to the query result, it includes:
[0152] Step e1, judge whether the cached data volume corresponding to the cached query result is greater than or equal to the total number of queries to be queried. If it is greater than or equal to the total number of queries to be queried, then use the cached query result as the final query result and end the query;
[0153] Step e2, if the cached data volume is less than the total number of queries to be queried, then execute the step of updating the query condition and the number of concurrent nodes according to the query result.
[0154] Furthermore, before starting the next query after the query result is obtained in the current query, first judge whether the cached data volume obtained after the current query meets the requirements of the total number of queries to be queried. Specifically, the query results obtained from each query are cached, and the query result obtained from the current query is also cached accordingly, and together with the previously cached query results, they form the cached query result. The data volume of the cached query result is statistically analyzed to obtain the cached data volume of each cache. Then, the cached data volume is compared with the total number of queries to be queried to judge whether the cached data volume is greater than or equal to the total number of queries to be queried. If it is determined to be greater than or equal to the total number of queries to be queried, it means that the data volume obtained from each query meets the query requirements. Therefore, each cached query result is used as the final query result and the query is ended.
[0155] Further, if it is judged that the cached data volume is less than the total number of queries to be queried, it means that the data volume obtained from each query does not meet the query requirements. Then, the query condition and the number of concurrent nodes are updated according to the query result, and the query is continued based on the updated query condition and the number of concurrent nodes until the total data volume obtained from the query meets the total number of queries to be queried, and the query operation is received.
[0156] In this embodiment, the data layer of the query and the query result are cached in the form of key-value pairs, so as to quickly determine the data layer from which each query result comes. At the same time, according to the size relationship between the data volume of each query and the total number of queries to be queried, it is determined whether to continue the query; while meeting the query requirements, excessive query is avoided, query resources are saved, and query efficiency is improved.
[0157] The present invention also provides a data query device. The data query device includes:
[0158] A construction module, configured to construct a query condition and determine the number of concurrent nodes;
[0159] A query module, configured to query a data layer corresponding to the query condition in a tree - like data structure according to the query condition and the number of concurrent nodes, and obtain a query result;
[0160] An update module, configured to update the query condition and the number of concurrent nodes according to the query result, and based on the updated query condition and the number of concurrent nodes, execute the step of querying a data layer corresponding to the query condition in the tree - like data structure according to the query condition and the number of concurrent nodes, until all data layers in the tree - like data structure have been queried to obtain a final query result.
[0161] Further, the construction module further includes:
[0162] A judgment unit, configured to judge whether there is a cached data layer. If there is the cached data layer, construct a query condition according to the cached data layer and determine the number of concurrent nodes;
[0163] A construction unit, configured to if there is no cached data layer, construct a query condition according to the received dimensional keywords and determine the number of concurrent nodes.
[0164] Further, the judgment unit is further configured to:
[0165] Obtain a target data layer at the next level of the cached data layer in the tree - like data structure, and the data volume of the cached result corresponding to the cached data layer;
[0166] Construct the query condition according to the target data layer, and determine the number of concurrent nodes according to the total number to be queried and the data volume.
[0167] Further, the judgment unit is further configured to:
[0168] Judge whether each data layer in the tree - like data structure carries a query identifier. If all carry the query identifier, use the cached data corresponding to the cached data layer as the final query result and end the query;
[0169] If not all carry the query identifier, execute the step of obtaining a target data layer at the next level of the cached data layer in the tree - like data structure.
[0170] Further, the construction unit is further configured to:
[0171] Determine the probing level according to the number of received dimensional keywords, and construct a query condition based on the dimensional keyword and the probing level;
[0172] Initiate a probing query on the tree-like data structure according to the query condition and the probing level, and obtain a probing query result;
[0173] Determine the number of concurrent nodes according to the total number of queries to be retrieved and the number of results of the probing query result.
[0174] Furthermore, the data query device further includes:
[0175] A cache module for caching the data layer corresponding to the query condition in the tree-like data structure and the query result in the form of key-value pairs.
[0176] Furthermore, the data query device further includes:
[0177] A judgment module for judging whether the amount of cached data corresponding to the cached query result is greater than or equal to the total number of queries to be retrieved. If it is greater than or equal to the total number of queries to be retrieved, use the cached query result as the final query result and end the query;
[0178] An execution module for, if the amount of cached data is less than the total number of queries to be retrieved, execute the steps of updating the query condition and the number of concurrent nodes according to the query result.
[0179] The specific implementation manner of the data query device of the present invention is basically the same as that of the above-mentioned embodiments of the data query method, and will not be elaborated here.
[0180] In addition, an embodiment of the present invention also provides a readable storage medium.
[0181] A data query program is stored on the readable storage medium. When the data query program is executed by a processor, the steps of the data query method described above are implemented.
[0182] The readable storage medium of the present invention can be a computer-readable storage medium. Its specific implementation manner is basically the same as that of the above-mentioned embodiments of the data query method, and will not be elaborated here.
[0183] The embodiments of the present invention are described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can still make many forms without departing from the spirit and scope protected by the present invention and the claims. All those equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, are within the scope of protection of the present invention.
Claims
1. A data query method, characterized in that, The data query method includes the following steps: Determine whether there is a cache data layer. If there is no such cache data layer, construct a query condition according to the received dimensional keywords, and determine the number of concurrent nodes; If there is such a cache data layer, determine whether each data layer in the tree-like data structure carries a query identifier; if all carry the query identifier, use the cache data corresponding to the cache data layer as the final query result and end the query; if not all carry the query identifier, obtain the target data layer at the next level below the cache data layer in the tree-like data structure, and the data volume of the cache result corresponding to the cache data layer; construct the query condition according to the target data layer, and determine the number of concurrent nodes according to the total number of queries to be made and the data volume; Query the data layer corresponding to the query condition in the tree-like data structure according to the query condition and the number of concurrent nodes to obtain a query result; Update the query condition and the number of concurrent nodes according to the query result, and based on the updated query condition and the number of concurrent nodes, execute the step of querying the data layer corresponding to the query condition in the tree-like data structure according to the query condition and the number of concurrent nodes until all data layers in the tree-like data structure have been queried to obtain the final query result.
2. The data query method according to claim 1, characterized in that The step of constructing a query condition according to the received dimensional keywords and determining the number of concurrent nodes includes: Determine a tentative level according to the number of received dimensional keywords, and construct a query condition according to the dimensional keywords and the tentative level; Initiate a tentative query on the tree-like data structure according to the query condition and the tentative level to obtain a tentative query result; Determine the number of concurrent nodes according to the total number of queries to be made and the result quantity of the tentative query result.
3. The data query method according to claim 1, wherein After the step of querying the data layer corresponding to the query condition in the tree-like data structure to obtain a query result, it includes: Cache the data layer corresponding to the query condition in the tree-like data structure and the query result in the form of key-value pairs.
4. The data query method according to claim 3, wherein Before the step of updating the query condition and the number of concurrent nodes according to the query result, it includes: Determine whether the cached data volume corresponding to the cached query result is greater than or equal to the total number of queries to be made. If it is greater than or equal to the total number of queries to be made, use the cached query result as the final query result and end the query; If the cached data volume is less than the total number of queries to be made, execute the step of updating the query condition and the number of concurrent nodes according to the query result.
5. The data query method according to claim 3, wherein After the step of querying the data layer corresponding to the query condition in the tree-like data structure to obtain a query result, it includes: Add a query identifier to the data layer corresponding to the query condition in the tree-like data structure.
6. The data query method according to claim 3, wherein After the step of querying the data layer corresponding to the query condition in the tree-like data structure to obtain a query result, it includes: Add a result identifier to the query result according to the query condition, and the dimensional keywords and query keywords corresponding to the query condition.
7. The data query method according to claim 3, wherein After the step of querying the data layer corresponding to the query condition in the tree - like data structure to obtain a query result, the following steps are included: Add a counting identifier to the data layer corresponding to the query condition in the tree - like data structure.
8. The data query method according to claim 1, wherein Before the steps of constructing the query condition and determining the number of concurrent nodes, the following steps are included: Obtain multiple pieces of data that support querying, and clean the multiple pieces of data to generate valid data; Construct the valid data into the tree - like data structure.
9. A data query device, characterized in that, The data query device includes: A construction module for constructing a query condition and determining the number of concurrent nodes; The construction module includes: A judgment unit for judging whether there is a cached data layer. If there is the cached data layer, construct a query condition and determine the number of concurrent nodes according to the cached data layer; A construction unit for, if there is no cached data layer, construct a query condition and determine the number of concurrent nodes according to the received dimensional keywords; The judgment unit is further used for: Obtain the target data layer at the next level of the cached data layer in the tree - like data structure, and the data volume of the cached result corresponding to the cached data layer; Construct the query condition according to the target data layer, and determine the number of concurrent nodes according to the total number of queries to be made and the data volume; The judgment unit is further used for: Judge whether each data layer in the tree - like data structure carries a query identifier. If all carry the query identifier, use the cached data corresponding to the cached data layer as the final query result and end the query; If not all carry the query identifier, execute the step of obtaining the target data layer at the next level of the cached data layer in the tree - like data structure; A query module for querying the data layer corresponding to the query condition in the tree - like data structure according to the query condition and the number of concurrent nodes to obtain a query result; An update module for updating the query condition and the number of concurrent nodes according to the query result, and based on the updated query condition and the number of concurrent nodes, execute the step of querying the data layer corresponding to the query condition in the tree - like data structure according to the query condition and the number of concurrent nodes until all data layers in the tree - like data structure have been queried to obtain the final query result.
10. The data query device according to claim 9, characterized in that, The construction unit is further used for: Determine a tentative level according to the number of received dimensional keywords, and construct a query condition according to the dimensional keywords and the tentative level; Initiate a tentative query on the tree - like data structure according to the query condition and the tentative level to obtain a tentative query result; Determine the number of concurrent nodes according to the total number of queries to be made and the result quantity of the tentative query result.
11. The data query device according to claim 9, wherein The data query device further includes: A cache module for caching the data layer corresponding to the query condition in the tree - like data structure and the query result in the form of key - value pairs.
12. The data query device according to claim 9, wherein, The data query device further includes: A judgment module for judging whether the cached data volume corresponding to the cached query result is greater than or equal to the total number of queries to be made. If it is greater than or equal to the total number of queries to be made, use the cached query result as the final query result and end the query; An execution module, configured to, if the amount of cached data is less than the total number of queries to be made, execute the step of updating the query condition and the number of concurrent nodes according to the query result.
13. A data query device, characterized in that, The data query device includes a memory, a processor, and a data query program stored on the memory and executable on the processor. When the data query program is executed by the processor, it implements the steps of the data query method according to any one of claims 1-8.
14. A readable storage medium, characterized in that, A data query program is stored on the readable storage medium. When the data query program is executed by the processor, it implements the steps of the data query method according to any one of claims 1-8.
Citation Information
Patent Citations
Interaction-type big-data query method and device based on distributed system, storage medium and terminal equipment
CN108241539A