A data query method, device, electronic device and storage medium
By sending data copy and query instructions to each data source terminal on the transit processing end, the data provider terminal can process data queries, solving the problem of low efficiency of cross-heterogeneous data sources, and achieving efficient query and the ability to respond to large concurrent requests.
Patent Information
- Application Number
- CN202210208059.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-04
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-04
AI Technical Summary
When queries across heterogeneous data sources, it is difficult for the existing technology to implement fast programming and efficient querying, resulting in slow query speed, especially when large concurrent query requests, the transit processing end is prone to crash.
After receiving the data query request at the transit processing end, the storage location of the target data table is determined, and data copy instructions and data query instructions are sent to each data source end that stores the target data table, so that each data provider end can obtain and query all target data tables, thereby reducing the resource occupation and calculation load of the transit processing end.
This method can make full use of the computing resources of the data provider end, reduce resource usage on the transit processing end, improve the speed of data query results, and cope with large concurrent query requests, avoiding the relay processing end crash.
Smart Images

Figure CN114564481B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data query method, device, electronic device and storage medium. Background Art
[0002] Big data is booming. In order to more comprehensively explore the value of data, it is often necessary to query data across different data sources. For example, it is necessary to analyze the behavioral habits of players of one game in another game, or to analyze the similarities and differences in data distribution between e-commerce business and cloud music business.
[0003] Usually, different data sources use different underlying storage structures or different databases for data storage. For example, one database mainly provides services to upper-level businesses through disk storage, while another database mainly provides services to upper-level businesses through memory storage. In this case, they are heterogeneous storages. For example, one data is stored in a MySQL database and another data is stored in a Hive data warehouse. In this case, they are also heterogeneous storages, and data sources are located in different regions and on different machines. Currently, although different data sources provide a Uniform Resource Identifier (URI) of domain name + port to provide data access, it is still difficult to organize multiple data sources together to support fast programming, efficient query, and cross-data source data query due to differences in data sources and underlying storage structures. Summary of the invention
[0004] In view of this, the embodiments of the present application at least provide a data query method, device, electronic device and storage medium, which can cope with a large number of concurrent query requests and improve the speed of obtaining data query results.
[0005] This application mainly includes the following aspects:
[0006] In a first aspect, an embodiment of the present application provides a data query method, which is applied to a transit processing end in a data storage system, wherein the data storage system further includes a data request end and a plurality of data providing ends; the data query method includes:
[0007] Receiving a data query request sent by the data request end, and determining a first identifier of a target data table storing the data requested to be queried based on the data query request;
[0008] Based on the first identifier of the target data table, querying the target data provider stored in the target data table from a preset mapping dictionary;
[0009] If it is found that all the target data tables are stored in at least two target data providing terminals, a data copy instruction and a data query instruction are sent to each target data providing terminal; the data copy instruction is used to instruct any target data providing terminal to copy the target data table that is not stored in itself to other target data providing terminals; the data query instruction is used to instruct to perform data query on all the target data tables;
[0010] Receive a data query result sent by any target data provider, and send the data query result to the data requester.
[0011] In a second aspect, an embodiment of the present application further provides a data query method, which is applied to a first data provider in a data storage system, wherein the data storage system further includes a data requester, a transfer processing terminal, and a second data provider; the data query method includes:
[0012] receiving a data copy instruction and a data query instruction sent by the transfer processing end; the data copy instruction is used to instruct the second data providing end to copy a first data table that is not stored in the second data providing end;
[0013] Based on the data copy instruction, obtaining the first data table from the second data provider;
[0014] Performing data query on all data tables indicated in the data query instruction to obtain data query results; all data tables include the stored second data table and the copied first data table;
[0015] In a third aspect, an embodiment of the present application further provides a data query device, which is applied to a transfer processing end in a data storage system, wherein the data storage system further includes a data request end and a plurality of data providing ends; the data query device includes:
[0016] A determination module, configured to receive a data query request sent by the data request end, and determine a first identifier of a target data table for storing the data requested to be queried based on the data query request;
[0017] A query module, configured to query a target data provider stored in the target data table from a preset mapping dictionary based on a first identifier of the target data table;
[0018] A first sending module is used for sending a data copy instruction and a data query instruction to each target data provider if it is found that all target data tables are stored in at least two target data providers; the data copy instruction is used to instruct any target data provider to copy a target data table that is not stored in the provider to other target data providers; the data query instruction is used to instruct to perform data query on all target data tables;
[0019] The second sending module is used to receive a data query result sent by any target data provider, and send the data query result to the data requester.
[0020] In a fourth aspect, an embodiment of the present application further provides a data query device, which is applied to a first data provider in a data storage system, wherein the data storage system further includes a data requester, a transfer processing terminal, and a second data provider; the data query device includes:
[0021] A receiving module, used for receiving a data copy instruction and a data query instruction sent by the transfer processing end; the data copy instruction is used for instructing to copy a first data table not stored in the second data providing end to the second data providing end;
[0022] an acquisition module, configured to acquire the first data table from the second data provider based on the data copy instruction;
[0023] A generating module, configured to perform data query on all data tables indicated in the data query instruction to obtain data query results; all data tables include the stored second data table and the copied first data table;
[0024] The third sending module is used to send the data query result to the transfer processing end.
[0025] In the fifth aspect, an embodiment of the present application also provides an electronic device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and the machine-readable instructions are executed by the processor to execute the steps of the data query method described in the first aspect or the second aspect above.
[0026] In a sixth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data query method described in the first aspect or the second aspect are executed.
[0027] The data query method, device, electronic device and storage medium provided by the embodiments of the present application, the transfer processing end in the present application issues data copy instructions and data query instructions to each data source end storing the target data table required for the query, so that each data provider end obtains all the target data tables and performs data query on them to obtain the final data query result. Compared with the prior art in which each data source end is only responsible for a part of the data query and summarizes the query results of each part to the data transfer end for unified query processing, which requires occupying a large amount of resources of the transfer processing end, and when there are a large number of concurrent query requests, the memory occupied by the transfer processing end will become larger and larger, resulting in insufficient memory, slow query speed, and crash of the transfer processing end, the present application can make full use of the computing resources of the data provider end, reduce the occupation of resources of the transfer processing end, and can cope with a large number of concurrent query requests. Moreover, since multiple data providers perform data queries in a competitive manner, the speed of obtaining data query results can be improved.
[0028] Furthermore, in the data query method provided in the embodiment of the present application, after the transit processing end receives the data query result sent by any data provider, it can avoid wasting computing resources by sending a query cancellation instruction and a copy cancellation instruction to other data providers.
[0029] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0031] Figure 1 A schematic diagram of the architecture of a data storage system provided by an embodiment of the present application is shown;
[0032] Figure 2 A flow chart of a data query method provided by an embodiment of the present application is shown;
[0033] Figure 3 A flowchart of another data query method provided in an embodiment of the present application is shown;
[0034] Figure 4 A flow chart of a data query method in a specific embodiment of the present application is shown;
[0035] Figure 5A functional module diagram of a data query device provided in an embodiment of the present application is shown;
[0036] Figure 6 A functional module diagram of another data query device provided in an embodiment of the present application is shown;
[0037] Figure 7 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0038] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of explanation and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn in real proportion. The flowchart used in this application shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowchart can be implemented out of sequence, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art, under the guidance of the content of the present application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.
[0039] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0040] The following methods, devices, electronic devices or computer-readable storage media of the embodiments of the present application can be applied to any scenario requiring data query, and are particularly suitable for scenarios of joint query of multi-source heterogeneous data. The embodiments of the present application are not limited to specific application scenarios, and any solutions using the data query methods, devices, electronic devices and storage media provided by the embodiments of the present application are within the scope of protection of this application.
[0041] It is worth noting that before the present application was proposed, each data source in the relevant scheme was only responsible for querying a part of the data, and the query results of each part were summarized to the data transfer end for unified query processing. This required a large amount of computing resources on the transfer processing end, and when there were a large number of concurrent query requests, the memory occupied by the transfer processing end would become larger and larger, resulting in insufficient memory, slow query speed, and crash of the transfer processing end.
[0042] In response to the above problems, a data query method provided in an embodiment of the present application is applied to a transfer processing end, and determines the storage location of the target data table required for the query according to the data query request. If all the target data tables required for the query are stored in at least two data providers, data copy instructions and data query instructions are issued to each data source end storing the target data table, and the query results obtained by any data provider for data query of all the target data tables are received. In this way, the transfer processing end only needs to issue instructions, and does not need to obtain and process intermediate data. The intermediate data is processed by the data provider end with higher computing power to obtain the data query results, which can make full use of the computing resources of the data provider end, reduce the occupation of resources of the transfer processing end, and can cope with large concurrent query requests. Moreover, since multiple data providers perform data queries in a competitive manner, the speed of obtaining data query results can be improved.
[0043] It should be noted that a database cluster uses at least two or more database servers to form a virtual single database logical image, which provides transparent data services to clients like a single database system; the functions, regions, and businesses of different database clusters may vary.
[0044] To facilitate the understanding of the present application, the technical solution provided by the present application is described in detail below in conjunction with specific embodiments.
[0045] In one embodiment of the present application, Figure 1 A schematic diagram of the architecture of a data storage system 100 provided in an embodiment of the present application is shown. Figure 1 As shown, the data storage system 100 includes a data request end 101, a transit processing end 102 and multiple data providing ends 103; here, the data storage system 100 can be a distributed storage system, such as a storage system that integrates multiple different data sources and different underlying storage structures.
[0046] Here, the data request end 101 can be understood as a client, through which a user can send a data query request. Usually, the client can be a browser page / software / interface, etc., and can use HTTP protocol, RPC protocol, etc.; the transfer processing end 102 can be an agent, such as Proxy, whose function is to act as an agent for network users to obtain network information, that is, a transfer station for network information; the data provider end 103 can be understood as a data cluster end, that is, a data source end, which is used to provide data involved in data query services. Usually, the data source can be represented by a URI, and the data in the data cluster can be called by the address in the URI. If different databases are installed on the machine where a data cluster is located, they can also be considered as two data sources, because the ports are different, forming two different URIs, wherein a URI is usually composed of a domain name: port (host:port), and the host can be an IP address or a domain name. For example, the domain name format is URI_1=192.168.0.1:2022; the IP address format is URI_2=www.demo.net:4091. Different data providers 103 have different storage methods for data, and the storage methods include at least one of memory storage, disk storage, and database storage.
[0047] In one embodiment of the present application, Figure 2 This is a flow chart of a data query method provided by an embodiment of the present application. Figure 2 As shown, the data query method provided in the embodiment of the present application is applied to the transit processing end in the data storage system mentioned in the above embodiment. The data storage system also includes a data request end and multiple data providing ends. The data query method includes the following steps:
[0048] S201: Receive a data query request sent by the data request end, and determine a first identifier of a target data table storing data requested for query based on the data query request.
[0049] In a specific implementation, the transit processing end receives in real time a data query request sent from a data provider. Here, the data query request is usually presented in a description statement using a structured query language (SQL), which is a statement for completing a specific query condition or business data query, that is, a description of the query requirement. The data query request can be parsed to determine in which data tables the requested data is located. Usually, the first identifier of a target data table storing the requested data can be parsed from the data query request, wherein the data corresponding to the data query request may be stored in one or more target data tables.
[0050] It should be noted that SQL is a special-purpose programming language, which can be understood as a database query and programming language. It can be used to access data as well as query, update and manage relational database systems. SQL query can be understood as obtaining a subset of data from a database. SQL query is usually applied on the client side. Each client has a fixed database, and the format of the query language used by each database may be different.
[0051] Usually, a data provider's database can store multiple data tables, and a data table contains multiple data on related topics. For example, a file contains user list data and user click behavior data. For a two-dimensional data table, it is usually organized in rows and columns. Usually each column has the same attribute, and each row is a data point.
[0052] Here, the first identifier of each data table is an identifier used to characterize the data table, such as the name and number of the data table.
[0053] Here, the process of determining the first identification of the target data table for requesting data storage based on the data query request is described, including the following steps: parsing the data query request according to a preset parsing method to determine the first identification of the target data table for the data storage requested for query. The preset parsing method includes but is not limited to a regular expression parsing method and an abstract syntax tree parsing method. As long as the parsing method can parse the first identification of the data table from the data query request, it is within the protection scope of this application.
[0054] It should be noted that regular expressions, also known as regular expressions, are usually used to retrieve and replace text that conforms to a certain rule. Many programming languages support string operations using regular expressions. Abstract Syntax Tree (AST), or syntax tree for short, is an abstract representation of the syntax structure of source code. It represents the syntax structure of a programming language in a tree form, and each node on the tree represents a structure in the source code.
[0055] S202: Based on the first identifier of the target data table, query the target data provider stored in the target data table from a preset mapping dictionary.
[0056] In a specific implementation, after the transit processing end determines the first identifier of the target data table used for the query from the data query request, it can query the storage location of the target data table from the preset mapping dictionary based on the first identifier of the target data table, that is, determine which data provider end the target data table is specifically stored in, wherein the preset mapping dictionary stores the second identifier of each data provider end and the first identifier of the data table stored by each data provider end in an associated manner.
[0057] In one example, the first identifier of a data table can be represented by the table name of the data table, and the second identifier of the data provider can be represented by a URI. The preset mapping dictionary stores the mapping between the table name and the URI, such as {table_a:URI_1, table_b:URI_2, table_c:URI_3, ...}.
[0058] It should be noted that the preset mapping dictionary is pre-generated, and there is no need to manually search and check the data tables one by one to determine their storage locations when receiving a data query request as in the prior art. This greatly improves the efficiency of determining the storage location of the data table, thereby reducing the time spent on obtaining the data query results. The generation process of the preset mapping dictionary is described below, including the following steps:
[0059] Receive a data table registration request sent by each data provider; associate and store the second identifier of each data provider with the first identifier of the data table stored by the data provider, and generate the preset mapping dictionary. The data table registration request carries the second identifier of each data provider and the first identifier of the data table stored by the data provider.
[0060] In a specific implementation, before the data storage system formally processes data queries, it is necessary to first establish an association relationship between the transit processing end and multiple data provider ends in the data storage system, and generate a preset mapping dictionary after the association relationship is established. Specifically, the transit processing end will receive a data table registration request sent by each data provider end, and the data table registration request will carry the second identifier of the data provider end and the first identifier of the data table stored by the data provider end, that is, each data provider end will report to the transit processing end which data tables are stored in each of them. In this way, the transit processing end associates and stores the second identifier of each data provider end with the first identifier of the data table stored by the data provider end, and thus generates a preset mapping dictionary.
[0061] Here, the transit processing end will also update the preset mapping dictionary in real time or periodically. Specifically, when the data stored in a data provider changes, the transit processing end receives an update request for the first identifier of the data table stored in the data provider end; or the transit processing end obtains the first identifier of the data table stored in each data provider end in real time or at regular intervals, thereby updating the preset mapping dictionary. The following is an implementation step for updating the preset mapping table, including:
[0062] Sending a request for obtaining the first identifier of the stored data table to each data provider; after obtaining the first identifier of the data table stored by each data provider, updating the preset mapping dictionary.
[0063] S203: If it is found that all the target data tables are stored in at least two target data providing ends, a data copy instruction and a data query instruction are sent to each target data providing end.
[0064] The data copy instruction is used to instruct any target data provider to copy a target data table not stored in itself to other target data providers; and the data query instruction is used to instruct to perform data query on all target data tables.
[0065] In a specific implementation, the data to be queried corresponding to the data query request may be located in multiple target data tables, and these target data tables may not be in the same target data provider. Here, after the storage location of each target data table is queried, it will be determined whether all the target data tables are stored in the same target data provider. If these data tables are stored in at least two target data providers, that is, part of the target data tables are stored in one target data provider, and the other part of the target data tables are stored in other target data providers. In this case, the transit processing end in the present application will send data copy instructions and data query instructions to each target data provider storing the target data table. Here, the copy instructions sent by the transit processing end to each target data provider instruct the target data provider to copy the target data tables that are not stored in itself to other target data providers, and the data query instructions instruct each target data provider to perform data query on all the target data tables copied and stored by itself.
[0066] Here, since the data processing speed of each data provider may be different, the present application enables each target data provider that stores a certain number of target data tables to independently execute a complete data query task, that is, to execute the data query task corresponding to the data request query. In this way, the data query results can be fed back in a competitive form through each target data provider. Compared with the prior art in which each data provider is only responsible for a part of the data query, and the query results of each part are summarized to the transit processing end for unified query processing, the data query speed is slower. The present application can improve the speed of obtaining data query results.
[0067] It should be noted that, since the data storage structures of different data providers may be different and the data types of data stored in different data tables may be different, and different types of data are incompatible, after a target data provider copies the target data tables of other target data providers, it is necessary to first uniformly modify the data types of the data in these target data tables to preset types, so as to facilitate data query for all target data tables. Specifically, a type modification rule indicating the data type of the data in the copied target data table may be carried in the data copy instruction. Furthermore, the identifier of the copied target data table may be modified, so that it is easy to distinguish whether it is the data table of the data provider itself or the data table copied from the outside.
[0068] In addition, since the table identifier (first identifier) of some target data tables is modified during the copying or storage process, different data query statements can be sent to different target data providers, and the first identifiers of the target data tables carried by different data query statements are different, that is, the data query statement in the data query request is rewritten. For example, the data query request corresponds to select price from table_1, table_2, table_3, then what is sent to target data provider 1 is select cast (price as decimal) from table_1, table_2_copy_URI_2, table_3_copy_URI_3, what is sent to target data provider 2 is select cast (price as decimal) from table_1_copy_URI_1, table_2, table_3_copy_URI_3, and what is sent to target data provider 3 is select cast (price as decimal) from table_1_copy_URI_1, table_2_copy_URL_2, table_3.
[0069] Here, after determining whether all target data tables are stored in the same target data provider, if these target data tables are stored in the same target data provider, in this case, since the target data provider stores all target data tables required for the requested query, there is no need to copy the data. Therefore, it is only necessary to send a data query request to the target data provider and only let the target data provider execute the query task.
[0070] Among them, the present application determines the number of target data providers stored in all target data tables, which refers to the number of target data providers after deduplication, that is, the minimum number of target data providers storing all target data tables is determined.
[0071] It should be noted that the transit processing end in the embodiment of the present application only needs to issue instructions, and does not need to obtain and process intermediate data. The intermediate data is processed by the data provider end which has higher computing power itself (because the data provider end is deployed on a high-performance machine cluster) to obtain data query results. In this way, the transit processing end can be deployed on any machine, and the hardware requirements of the machine are not high. It is a lightweight transit processing end.
[0072] S204: receiving a data query result sent by any target data provider, and sending the data query result to the data requester.
[0073] In a specific implementation, after receiving the data query result sent by any target data provider, the transit processing end will directly send the data query result to the data requesting end.
[0074] Here, each target data provider executes the data query instruction, but different target data providers have different data query speeds. Thus, the query speed of the data query result can be improved by adopting this form of competitive feedback query results.
[0075] In addition, after receiving the data query results sent by any target data provider, the transit processing end immediately sends a copy cancellation instruction and a query cancellation instruction to other target data providers, so that other target data providers stop performing data copy tasks and data query tasks after receiving the copy cancellation instruction and the query cancellation instruction, thereby reducing the waste of computing resources.
[0076] In an embodiment of the present application, the transit processing end determines the storage location of the target data table required for the query according to the data query request. If all the target data tables required for the query are stored in at least two data providers, data copy instructions and data query instructions are issued to each data source end storing the target data table, and the query results obtained by any data provider for data query of all the target data tables are received. The transit processing end in the present application only needs to issue instructions, and does not need to obtain and process intermediate data. The intermediate data is processed by the data provider end with higher computing power to obtain the data query results. In this way, the computing resources of the data provider end can be fully utilized, and the occupation of the resources of the transit processing end can be reduced. It can cope with a large number of concurrent query requests. Moreover, since multiple data providers perform data queries in a competitive manner, the speed of obtaining data query results can be improved.
[0077] In one embodiment of the present application, Figure 3 This is a flow chart of another data query method provided by an embodiment of the present application. Figure 3As shown, the data query method provided in the embodiment of the present application is applied to the first data provider in the data storage system mentioned in the above embodiment, and the data storage system also includes a data request end, a transfer processing end, and a second data provider. The data query method includes the following steps:
[0078] S301: receiving a data copy instruction and a data query instruction sent by the transfer processing end.
[0079] The data copy instruction is used to instruct the second data provider to copy the first data table that is not stored in the second data provider.
[0080] Here, the first data provider needs to send a data table registration request to the transit processing end in advance, wherein the data table registration request carries the second identifier of the second data table, so that the transit processing end knows which data tables are stored in each data provider, and then establishes a preset mapping dictionary.
[0081] S302: Based on the data copy instruction, obtain the first data table from the second data provider.
[0082] In a specific implementation, after receiving the data copy instruction sent by the transit processing end, the first data provider can determine which first data table to copy from which second data provider according to the first identifier of the first data table carried in the data copy instruction and the second identifier of the second data provider storing the first data table. After copying from the second data provider to the first data table, the first data provider can store the first data table, so that the first data table can be reused when receiving a data query instruction next time.
[0083] Here, in step S302, based on the data copy instruction, obtaining the first data table from the second data provider includes the following steps:
[0084] Sending a data copy request to the second data provider and receiving the first data table sent by the second data provider, wherein the data copy request carries a first identifier of the first data table requested to be copied.
[0085] Here, the data copy instruction also carries an identifier modification rule for the first data table. Then, the first data provider can modify the first identifier of the copied first data table according to the identifier modification rule carried in the data copy instruction, and store the first data table. In this way, it is easy to distinguish whether it is a data table stored by the data provider itself or a data table copied from the outside.
[0086] Here, the data copy instruction also carries type modification rules for the data types of the data in the first data table and the second data table. Then, the first data provider uniformly modifies the data types of the data in the first data table and the second data table to preset types. This makes it easier to query the data in the data tables with the uniform data type.
[0087] Furthermore, if the first data provider receives a copy cancel instruction sent by the transit processing end during the data copying process, it means that other data providers have already sent data query results to the transit processing end. At this time, the operation of obtaining the first data table from the second data provider is stopped, which can reduce the occupancy of storage resources.
[0088] Furthermore, if the first data provider receives a query cancellation instruction sent by the transit processing end during the data query process, it means that other data providers have already sent data query results to the transit processing end. At this time, the data query operation for all data tables is stopped, which can reduce the waste of computing resources.
[0089] S303: Performing data query on all data tables indicated in the data query instruction to obtain data query results.
[0090] Among them, all data tables include the stored second data table and the copied first data table.
[0091] In a specific implementation, after all the data tables indicated in the data query instruction are obtained, data query can be performed on these data tables to obtain query results.
[0092] S304: Send the data query result to the transfer processing end.
[0093] In a specific implementation, after the first data provider performs a data query on all acquired data tables according to the data query instruction to obtain a data query result, the first data provider directly sends the data query result to the transit processing end so that the transit processing end sends the data query result to the data request end.
[0094] In a specific embodiment of the present application, Figure 4 The flowchart of a data query method in a specific embodiment of the present application, wherein the client is a data request end, the proxy is a transfer processing end, and the URI is a data providing end. Figure 4 The steps are described in detail:
[0095] S401: The client submits SQL;
[0096] Here, the client submits the SQL statement to the Proxy; the client may be a browser page / software / interface, and the submission path may be the HTTP protocol or the RPC protocol, or other protocols.
[0097] S402: Proxy receives SQL;
[0098] Here, the Proxy immediately starts the subsequent processing after receiving the SQL statement.
[0099] S403: Proxy analyzes the table name used;
[0100] Here, Proxy analyzes the table name part in the SQL statement, such as table_1, table_2, table_3; the technology used in the process of analyzing the table name can be regular expressions, or it can be constructing an AST abstract syntax tree, or it can be other technologies.
[0101] S404: Proxy analysis of the URI used;
[0102] Here, Proxy searches the preset mapping dictionary MAP_TABLE_URI in memory through the parsed table name to determine which URI / URIs are used. The content of this preset mapping dictionary records the mapping of table to URI, such as {table_a:URI_1, table_b:URI_2, table_c:URI_3, ...}. In addition, Proxy will actively and regularly ask all URIs what tables they each have, thereby continuously updating this preset mapping dictionary and keeping this preset mapping dictionary MAP_TABLE_URI up to date.
[0103] S405: Determine whether only a single URI is used, if so, execute step S406, if not, execute step S408;
[0104] Here, the Proxy determines the number of URIs used by all tables after deduplication and whether only one URI is needed, that is, whether all data tables are on the same data source.
[0105] S406: Submit the SQL code directly to the URI;
[0106] Here, if only a single data source is used, the Proxy submits the SQL statement intact to this URI.
[0107] S407: This URI is calculated, the result is returned to the Proxy, and step S414 is executed;
[0108] Here, the URI executes the SQL statement and waits for the execution on the URI to complete. After the execution is completed, the URI returns the result to the Proxy.
[0109] S408: If multiple steps are used, execute step S409 and step S410 simultaneously;
[0110] Here, if multiple data sources are used, the following example uses table_1, table_2, and table_3 located in URI_1, URI_2, and URI_3 respectively.
[0111] S409: Require each URI to copy the data of other URIs to the local computer;
[0112] Here, the Proxy immediately sends instructions to the three URIs, asking them to copy the data tables of other URIs, where the copied data tables are named in a preset way. For example, URI_1 is asked to copy table_2 on URI_2 to the local computer and name it table_2_copy_URI_2, and to copy table_3 on URI_3 to the local computer and name it table_3_copy_URI_3; similarly, URI_2 is asked to copy and name table_1_copy_URI_1 and table_3_copy_URI_3; similarly, URI_3 is asked to copy and name table_1_copy_URI_1 and table_2_copy_URI_2.
[0113] S410: The Proxy sends the modified SQL statements to each URI;
[0114] Here, this step and step S409 occur simultaneously, and the Proxy modifies three sets of SQL statements, specifically, the table name needs to be modified and the data type needs to be upgraded, and the three modified SQL statements are sent to three URIs respectively. For example, if the original statement is select price from table_1, table_2, table_3, then SQL1 is select cast (price as decimal) from table_1, table_2_copy_URI_2, table_3_copy_URI_3 and is sent to URI_1; SQL2 is select cast (price as decimal) from table_1_copy_URI_1, table_2, table_3_copy_URI_3 and is sent to URI_2; SQL3 is select cast (price as decimal) from table_1_copy_URI_1, table_2_copy_URL_2, table_3 and is sent to URI_3.
[0115] S411: After each copy is completed, calculation is performed according to the SQL statement;
[0116] Here, each URI waits for execution after receiving the corresponding SQL statement. Once the table copy required on this URI is completed, the data query calculation can be started immediately.
[0117] S412: The calculation result is sent to the Proxy;
[0118] Here, once each URI is calculated, the result is returned to the Proxy.
[0119] S413: When the proxy receives the result of any URI, it requests all URIs to cancel the copy and calculation;
[0120] Here, when the Proxy receives the result of any one of the URIs, it immediately notifies all URIs that they can cancel data copying and calculation. This cancellation instruction is not mandatory. After receiving this cancellation instruction, the URI can continue copying and calculation, or stop copying or calculation immediately. The Proxy just tells the URI that I have obtained the result, and the URI can run according to its own operation plan.
[0121] S414: The Proxy returns the result to the client;
[0122] This completes the query.
[0123] S415: The Proxy periodically maintains the relationship between the table name and the URI.
[0124] Here, Proxy periodically updates the preset mapping dictionary.
[0125] Based on the same application concept, the embodiments of the present application also provide a data query device corresponding to the data query method provided in the above embodiments. Since the principle of solving the problem by the device in the embodiments of the present application is similar to the method in the above embodiments of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0126] In one embodiment of the present application, Figure 5 A schematic diagram of the structure of a data query device provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the data query device 500 includes:
[0127] A determination module 510 is configured to receive a data query request sent by the data request end, and determine a first identifier of a target data table storing the data requested to be queried based on the data query request;
[0128] A query module 520, configured to query a target data provider stored in the target data table from a preset mapping dictionary based on the first identifier of the target data table;
[0129] The first sending module 530 is used to send a data copy instruction and a data query instruction to each target data provider if it is found that all target data tables are stored in at least two target data providers; the data copy instruction is used to instruct any target data provider to copy a target data table that is not stored in the provider to other target data providers; the data query instruction is used to instruct to perform data query on all target data tables;
[0130] The second sending module 540 is used to receive a data query result sent by any target data provider, and send the data query result to the data requester.
[0131] In a possible implementation, Figure 5 As shown, the first sending module 530 is further used for:
[0132] If all the target data tables are found to be stored in the same target data provider, the data query request is forwarded to the target data provider.
[0133] In a possible implementation, Figure 5 As shown, the second sending module 540 is further used for:
[0134] Send copy cancellation instructions and query cancellation instructions to other target data providers.
[0135] In a possible implementation, Figure 5 As shown, the query module 520 is further used to generate the preset mapping dictionary according to the following steps:
[0136] Receiving a data table registration request sent by each data provider; the data table registration request carries a second identifier of each data provider and a first identifier of a data table stored in the data provider;
[0137] The second identifier of each data provider is associated with the first identifier of the data table stored in the data provider to generate the preset mapping dictionary.
[0138] In a possible implementation, Figure 5 As shown, the data query device 500 further includes an updating module 550; the updating module 550 is specifically used for:
[0139] Sending a request for obtaining a first identifier of a stored data table to each data provider;
[0140] After obtaining the first identifier of the data table stored in each data provider, the preset mapping dictionary is updated.
[0141] In a possible implementation, Figure 5 As shown, the determination module 510 is specifically used for:
[0142] Parsing the data query request according to a preset parsing method to determine a first identifier of a target data table where the data is stored;
[0143] Among them, the preset parsing methods include regular expression parsing method and abstract syntax tree parsing method.
[0144] In one embodiment of the present application, Figure 6 A schematic diagram of the structure of another data query device provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, the data query device 600 includes:
[0145] The receiving module 610 is used to receive a data copy instruction and a data query instruction sent by the transfer processing end; the data copy instruction is used to instruct the second data providing end to copy a first data table that is not stored in the second data providing end;
[0146] An acquisition module 620, configured to acquire the first data table from the second data provider based on the data copy instruction;
[0147] A generating module 630 is used to perform data query on all data tables indicated in the data query instruction to obtain data query results; all data tables include the stored second data table and the copied first data table;
[0148] The third sending module 640 is used to send the data query result to the transfer processing end.
[0149] In a possible implementation, Figure 6 As shown, the acquisition module 620 is specifically used for:
[0150] Sending a data copy request to the second data provider; the data copy request carries a first identifier of the first data table requested to be copied;
[0151] Receive the first data table sent by the second data provider.
[0152] In a possible implementation, Figure 6 As shown, the data query device 600 further includes a storage module 650; the storage module 650 is used to:
[0153] According to the identifier modification rule carried in the data copy instruction, the first identifier of the copied first data table is modified, and the first data table is stored.
[0154] In a possible implementation, Figure 6 As shown, the acquisition module 620 is also used for:
[0155] According to the type modification rule carried in the data copy instruction, the data types of the data in the first data table and the second data table are uniformly modified to preset types.
[0156] In a possible implementation, Figure 6 As shown, the receiving module 610 is also used for:
[0157] If a copy cancellation instruction sent by the transit processing end is received, the operation of obtaining the first data table from the second data providing end is stopped.
[0158] In a possible implementation, Figure 6 As shown, the receiving module 610 is also used for:
[0159] If a query cancellation instruction sent by the transfer processing end is received, the operation of performing data query on all data tables is stopped.
[0160] In a possible implementation, Figure 6 As shown, the third sending module 640 is further used for:
[0161] Sending a data table registration request to the transfer processing end;
[0162] The data table registration request carries the third identifier of the second data table.
[0163] Based on the same application concept, see Figure 7 As shown, it is a structural diagram of an electronic device 700 provided in an embodiment of the present application, including: a processor 710, a memory 720 and a bus 730, wherein the memory 720 stores machine-readable instructions executable by the processor 710, and when the electronic device 700 is running, the processor 710 communicates with the memory 720 through the bus 730, and the machine-readable instructions are executed by the processor 710 when running. Figure 2 or Figure 3 The steps of any of the data query methods described in .
[0164] Based on the same application concept, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data query method provided in the above embodiment are executed.
[0165] Specifically, the storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the storage medium is run, the above-mentioned data query method can be executed, the computing resources of the data provider end can be fully utilized, and the occupation of resources of the transit processing end can be reduced. It can cope with a large number of concurrent query requests, and since multiple data providers perform data queries in a competitive manner, the speed of obtaining data query results can be improved.
[0166] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the system and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0167] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0168] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0169] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0170] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A data query method, characterized in that, it is applied to a transit processing end in a data storage system, and the data storage system further includes a data request end and multiple data providing ends; the data query method includes: receiving a data query request sent by the data request end, and determining a first identifier of a target data table in which the data to be queried is stored based on the data query request; querying, from a preset mapping dictionary, a target data providing end in which the target data table is stored based on the first identifier of the target data table; if it is queried that all the target data tables are stored in at least two target data providing ends, sending a data copy instruction and a data query instruction to each target data providing end; the data copy instruction is used to instruct any target data providing end to copy the target data table that it does not store to other target data providing ends; the data query instruction is used to instruct to perform a data query on all the target data tables; receiving a data query result sent by any target data providing end, and sending the data query result to the data request end.
2. The data query method according to claim 1, characterized in that, the data copy instruction carries an identifier modification rule for the target data table indicated to be copied, and a type modification rule for the data type of the data in the target data table indicated to be copied.
3. The data query method according to claim 1, characterized in that, after querying, from the preset mapping dictionary, the target data providing end in which the target data table is stored, and before receiving the data query result sent by any target data providing end, the data query method further includes: if it is queried that all the target data tables are stored in the same target data providing end, forwarding the data query request to the target data providing end.
4. The data query method according to claim 1, characterized in that, after receiving the data query result sent by any target data providing end, the data query method further includes: sending a copy cancellation instruction and a query cancellation instruction to other target data providing ends.
5. The data query method according to any one of claims 1 to 4, characterized in that, before receiving the data query request sent by the data request end, the data query request further includes generating the preset mapping dictionary according to the following steps: receiving a data table registration request sent by each data providing end; the data table registration request carries a second identifier of each data providing end and a first identifier of the data table stored by this data providing end; associatively storing the second identifier of each data providing end and the first identifier of the data table stored by this data providing end, and generating the preset mapping dictionary.
6. The data query method according to any one of claims 1 to 4, characterized in that, the data query method further includes: sending a request for obtaining the first identifier of the data table stored to each data providing end; after obtaining the first identifier of the data table stored by each data providing end, updating the preset mapping dictionary.
7. The data query method according to claim 1, characterized in that, Different data providers have different storage methods for data, and the storage methods include at least one of in-memory storage, disk storage, and database storage.
8. The data query method according to claim 1, wherein, determining a first identifier of a target data table in which the requested data is stored based on the data query request includes: parsing the data query request according to a preset parsing method to determine a first identifier of a target data table in which the requested query data is stored; wherein the preset parsing method includes a regular expression parsing method and an abstract syntax tree parsing method.
9. A data query method, wherein, applied to a first data provider in a data storage system, the data storage system further includes a data requester, a transit processing end, and a second data provider; the data query method includes: receiving a data copy instruction and a data query instruction sent by the transit processing end; the data copy instruction is used to indicate copying a first data table that it does not store to the second data provider; obtaining the first data table from the second data provider based on the data copy instruction; performing a data query on all the data tables indicated in the data query instruction to obtain a data query result; all the data tables include a stored second data table and the copied first data table; sending the data query result to the transit processing end.
10. The data query method according to claim 9, wherein, obtaining the first data table from the second data provider based on the data copy instruction includes: sending a data copy request to the second data provider; the data copy request carries a first identifier of the first data table to be copied; receiving the first data table sent by the second data provider.
11. The data query method according to claim 9, wherein, after obtaining the first data table from the second data provider and before performing a data query on all the data tables indicated in the data query instruction, the data query method further includes: modifying the first identifier of the copied first data table according to the identifier modification rule carried in the data copy instruction, and storing the first data table.
12. The data query method according to claim 9, wherein, after obtaining the first data table from the second data provider and before performing a data query on all the data tables indicated in the data query instruction, the data query method further includes: uniformly modifying the data types of the data in the first data table and the second data table to a preset type according to the type modification rule carried in the data copy instruction.
13. The data query method according to claim 9, wherein, during the process of obtaining the first data table from the second data provider, the data query method further includes: if a copy cancellation instruction sent by the transit processing end is received, stopping the operation of obtaining the first data table from the second data provider.
14. The data query method according to claim 9, wherein, in the process of performing data query on all data tables indicated in the data query instruction, the data query method further includes: if a query cancellation instruction sent by the transit processing end is received, stop the operation of performing data query on all data tables.
15. The data query method according to claim 9, wherein, the data query method further includes: sending a data table registration request to the transit processing end; wherein, the second identifier of the second data table is carried in the data table registration request.
16. A data query device, wherein, applied to a transit processing end in a data storage system, the data storage system further includes a data request end and a plurality of data providing ends; the data query device includes: a determination module, configured to receive a data query request sent by the data request end, and determine a first identifier of a target data table in which the requested query data is stored based on the data query request; a query module, configured to query a target data providing end storing the target data table from a preset mapping dictionary based on the first identifier of the target data table; a first sending module, configured to, if it is queried that all target data tables are stored in at least two target data providing ends, send a data copy instruction and a data query instruction to each target data providing end; the data copy instruction is used to instruct any target data providing end to copy the target data table that it does not store to other target data providing ends; the data query instruction is used to instruct to perform data query on all target data tables; a second sending module, configured to receive a data query result sent by any target data providing end, and send the data query result to the data request end.
17. A data query device, wherein, applied to a first data providing end in a data storage system, the data storage system further includes a data request end, a transit processing end, and a second data providing end; the data query device includes: a receiving module, configured to receive a data copy instruction and a data query instruction sent by the transit processing end; the data copy instruction is used to instruct to copy a first data table that it does not store to the second data providing end; an obtaining module, configured to obtain the first data table from the second data providing end based on the data copy instruction; a generating module, configured to perform data query on all data tables indicated in the data query instruction to obtain a data query result; all data tables include a stored second data table and the copied first data table; a third sending module, configured to send the data query result to the transit processing end.
18. An electronic device, wherein, including: A processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory through the bus, and when the machine-readable instructions are run by the processor, the steps of the data query method described in any one of claims 1 to 8 are executed, or the steps of the data query method described in any one of claims 9 to 15 are executed.
19. A computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the data query method described in any one of claims 1 to 8 are executed, or the steps of the data query method described in any one of claims 9 to 15 are executed.
Citation Information
Patent Citations
Multi-source data query method, device and equipment and computer readable storage medium
CN113722353A
Target data obtaining method and apparatus
US20210191934A1