Data partitioning method and device and related product
By extending the historical query request set, the query efficiency problem caused by the different new and old query requests in the prior art is solved, and more efficient data partitioning and querying is achieved.
Patent Information
- Application Number
- CN202311553401.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
When the prior art uses historical query requests to partition a data set, the new query request is different from the historical query request, resulting in slower query efficiency.
By obtaining the historical query request set and the data total set, multiple historical query request pairs are obtained based on the query time of the historical query request set, their similarity is calculated and the maximum similarity is used as the extended value, the historical query request set is expanded, and finally the data total set is partitioned according to the extended historical query request set.
Improve query efficiency, optimize data partitioning and reduce query cost by simulating the similarity between historical query requests and new query requests.
Smart Images

Figure CN120020754A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data partitioning technology, and in particular to a data partitioning method, device and related products. Background Technology
[0002] With the continuous development of science and technology, distributed systems are mainly used for the storage and query of data sets in applications. Distributed systems divide data sets into multiple data partitions to achieve the storage and query of data sets through the data partitions. In related technologies, most of them use historical query requests to partition data sets. Historical query requests are requests executed when querying data in a data set. However, since new query requests may be different from historical query requests, using the partition to query the data corresponding to the new query request in the data set will result in slower query efficiency. Therefore, how to improve query efficiency has become a technical problem that needs to be solved urgently in the current field. SUMMARY OF THE INVENTION
[0003] The present application embodiment provides a data partitioning method, device and related products, aiming to improve query efficiency.
[0004] The first aspect of the present application provides a data partitioning method, comprising:
[0005] Obtain a historical query request set and a total data set, wherein the total data set includes a data subset corresponding to the historical query request set and other query data subsets;
[0006] According to the query time of the historical query request set, a plurality of historical query request pairs in the historical query request set are obtained, wherein the query time of the first historical query request in the target historical query request pair is earlier than the query time of the second historical query request in the historical query request pair, and the target historical query request pair is any one of the plurality of historical query request pairs;
[0007] Calculate the multiple historical query request pairs respectively to obtain the similarities corresponding to the multiple historical query request pairs respectively, and use the maximum value of the similarities corresponding to the multiple historical query request pairs respectively as the extension value of the historical query request set;
[0008] Extending each historical query request in the historical query request set according to the extension value to obtain an extended historical query request set;
[0009] Perform a partition operation on the data set according to the expanded historical query request set to obtain data partitions corresponding to the data set.
[0010] The second aspect of the present application provides a data partitioning device, including:
[0011] A request set acquisition unit, configured to acquire a historical query request set and a total data set, where the total data set includes a data subset corresponding to the historical query request set and the remaining query data subsets;
[0012] A request pair acquisition unit, configured to obtain multiple historical query request pairs in the historical query request set according to the query time of the historical query request set, where the query time of the first historical query request in the target historical query request pair is earlier than the query time of the second historical query request in the historical query request pair, and the target historical query request pair is any one of the multiple historical query request pairs;
[0013] An extended value acquisition unit, configured to calculate each of the multiple historical query request pairs respectively to obtain the similarity corresponding to each of the multiple historical query request pairs, and use the maximum value among the similarities corresponding to the multiple historical query request pairs as the extended value of the historical query request set;
[0014] An extended request set acquisition unit, configured to extend each historical query request in the historical query request set according to the extended value to obtain an extended historical query request set;
[0015] A total data set partitioning unit, configured to perform a partitioning operation on the total data set according to the extended historical query request set to obtain a data partition corresponding to the total data set.
[0016] A third aspect of the present application provides a computer device, where the device includes a processor and a memory:
[0017] The memory is used to store a computer program and transmit the computer program to the processor;
[0018] The processor is configured to execute the steps of the data partitioning method provided in the first aspect according to the instructions in the computer program.
[0019] A fourth aspect of the present application provides a computer-readable storage medium, where the computer-readable storage medium is used to store a computer program, and when the computer program is executed by a computer device, the steps of the data partitioning method provided in the first aspect are implemented.
[0020] A fifth aspect of the present application provides a computer program product, including a computer program, and when the computer program is executed by a computer device, the steps of the data partitioning method provided in the first aspect are implemented.
[0021] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0022] In the technical solution of this application, first, a historical query request set and a total data set are obtained. Then, according to the query time of the historical query request set, multiple historical query request pairs in the historical query request set are obtained. It should be noted that the query time of the first historical query in the target historical query request pair is earlier than the query time of the second historical query in the historical query request pair. After that, calculations are performed on the multiple historical query request pairs respectively to obtain the similarities corresponding to the multiple historical query request pairs respectively, and the maximum value among the similarities corresponding to the multiple historical query request pairs respectively is used as the expansion value of the historical query request set. Finally, each historical query request in the historical query request set is expanded according to the expansion value to obtain an expanded historical query request set, so as to perform a partitioning operation on the total data set according to the expanded historical query request set to obtain the data partitions corresponding to the total data set. It can be seen that in this application, the query requests corresponding to different query times in the query request pair are used to simulate the historical query requests and new query requests, and the query requests in the query request set are expanded according to the maximum value of the similarities between the two in the multiple query request pairs. In this way, compared with the related art, the data partitions obtained by performing a partitioning operation on the total data set using the expanded query request set in this application can be used to query data, which can improve the query efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 FIG. is a schematic diagram of data partitioning provided in the related art;
[0024] Figure 2 FIG. is a scenario architecture diagram of a data partitioning method provided in an embodiment of this application;
[0025] Figure 3 FIG. is a flowchart of a data partitioning method provided in an embodiment of this application;
[0026] Figure 4 FIG. is a schematic diagram of the similarity of historical query request pairs of a data partitioning method provided in an embodiment of this application;
[0027] Figure 5 FIG. is a flowchart of another data partitioning method provided in an embodiment of this application;
[0028] Figure 6 FIG. is a schematic diagram of calculating distances of a data partitioning method provided in an embodiment of this application;
[0029] Figure 7 FIG. is a schematic diagram of query request expansion of a data partitioning method provided in an embodiment of this application;
[0030] Figure 8 FIG. is a schematic diagram of partition splitting of a data partitioning method provided in an embodiment of this application;
[0031] Figure 9Schematic diagram of regular rectangular partitioning for a data partitioning method provided by an embodiment of the present application;
[0032] Figure 10 Schematic diagram of regular rectangular partitioning for another data partitioning method provided by an embodiment of the present application;
[0033] Figure 11 Schematic diagram of irregular shape partitioning for a data partitioning method provided by an embodiment of the present application;
[0034] Figure 12 Schematic diagram of partitioning for a data partitioning method provided by an embodiment of the present application;
[0035] Figure 13 Schematic diagram of enlarged partitioning for a data partitioning method provided by an embodiment of the present application;
[0036] Figure 14 Schematic diagram of query process for a data partitioning method provided by an embodiment of the present application;
[0037] Figure 15 Schematic diagram of the structure of a data partitioning device provided by an embodiment of the present application;
[0038] Figure 16 Schematic diagram of the structure of a server in an embodiment of the present application;
[0039] Figure 17 Schematic diagram of the structure of a terminal device in an embodiment of the present application. Detailed implementation manners
[0040] Next, embodiments of the present application will be described with reference to the accompanying drawings.
[0041] First, several noun terms that may be involved in the following embodiments of the present application will be explained.
[0042] Historical query request: For a system, query requests that have been received.
[0043] New query request: For a system, query requests that will be received.
[0044] With the continuous development of science and technology, distributed systems are mainly used for the storage and query of data sets in applications. Distributed systems divide data sets into multiple data partitions to achieve the storage and query of data sets through the data partitions. In related technologies, most distributed systems use historical query requests to partition data sets. Historical query requests are requests executed when querying data in a data set. However, since new query requests may be different from historical query requests, using the partition to query the data corresponding to the new query request in the data set will result in slower query efficiency. Therefore, how to improve query efficiency has become a technical problem that needs to be solved urgently in the current field.
[0045] The following combination Figure 1 To illustrate data partitioning in related technologies. Figure 1 is a schematic diagram of data partitioning provided in the relevant technology, where Figure 1 (a) After partitioning the data set according to historical query request A, data partitions PA1, PA2 and PA3 are obtained. The query cost of data partition PA1 is 220MB, the query cost of data partition PA2 is 130MB, and the query cost of data partition PA3 is 150MB. At this time, if the data partitions are used to query the data corresponding to historical query request A in the data set, only data partition PA1 needs to be queried, and its query cost is 220MB; Figure 1 (b) After partitioning the data set according to the historical query request A, use the partition to query the data set with the new query request A 1 Corresponding data, at this time, the new query request A 1 Different from historical query request A, query new query request A 1 The corresponding data needs to be queried in data partitions PA1 and PA3, and the query cost is 370MB. As the query cost increases, the query efficiency becomes slower and slower.
[0046] In view of the above problems, a data partitioning method, device and related products are provided in the present application, aiming to improve query efficiency. In the technical solution provided in the present application, historical query requests and new query requests are simulated by query requests corresponding to different query times in historical query request pairs, and the similarity between the historical query request and the new query request is calculated. Furthermore, the maximum value of the similarities corresponding to multiple historical query request pairs is used as the extension value of the historical query request set, so as to expand the historical query requests in the historical query request set according to the extension value, and finally the data set is partitioned according to the expanded historical query request set to obtain the data partitions corresponding to the data set. In this way, querying data according to the data partitions in the present application can improve query efficiency.
[0047] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.
[0048] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0049] Furthermore, the data partitioning method provided in this application mainly involves big data. Big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time range. It is a vast, high-growth-rate, and diverse information asset that requires new processing models to have stronger decision-making power, insight discovery ability, and process optimization ability. With the advent of the cloud era, big data has attracted more and more attention. Big data requires special technologies to effectively process a large amount of data that can tolerate the elapsed time. Technologies applicable to big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems. In the embodiments of this application, big data technology is mainly used to process historical query request pairs in the historical query request set to obtain an extended historical query request set, and finally partition the total data set according to the extended historical query request set to obtain the data partition of the total data set, thereby solving the problem of slow query efficiency.
[0050] The execution subject of the data partitioning method provided by the embodiments of the present application can be a terminal device. For example, a historical query request set and a total data set are obtained on the terminal device. As an example, the terminal device may specifically include, but is not limited to, mobile phones, desktop computers, tablet computers, laptop computers, handheld computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The execution subject of the data partitioning method provided by the embodiments of the present application can also be a server, that is, a historical query request set and a total data set can be obtained on the server. The data partitioning method provided by the embodiments of the present application can also be executed collaboratively by the terminal device and the server. Therefore, the embodiments of the present application do not limit the implementation subject for executing the technical solution of the present application.
[0051] Figure 2 Exemplarily shows a scenario architecture diagram of a data partitioning method. The figure includes a server and various forms of terminal devices. Figure 2 The shown server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. Additionally, the server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.
[0052] See Figure 3 , which is a flowchart of a data partitioning method provided by the embodiments of the present application. As Figure 3 shown in the data partitioning method, the following steps are included:
[0053] S301: Obtain a historical query request set and a total data set.
[0054] In this step, the total data set includes a data subset corresponding to the historical query request set and the remaining query data subsets, where the historical query request set includes multiple historical query requests, and the data subset corresponding to the historical query request set includes data corresponding to the multiple historical query requests respectively. Further, the historical query request set can be understood as the request set executed when querying the data in the total data set that corresponds to the multiple historical query requests respectively. Correspondingly, the remaining query data subsets can be understood as the data subsets in the total data set that are not queried according to the historical query request set, that is, the data subsets not queried by the query requests.
[0055] S302: Obtain multiple pairs of historical query requests in the historical query request set according to the query time of the historical query request set.
[0056] In this step, the query time of the first historical query request in the target historical query request pair is earlier than that of the second historical query request in the historical query request pair, and the target historical query request pair is any one of multiple historical query request pairs. It can be understood that in this application, the historical query request set is divided into multiple historical query request pairs, and the query times in each historical query request pair are different. In this way, the query request with a relatively earlier query time in the historical query request pair can be simulated as a historical query request, and the query request with a relatively later query time in the historical query request pair can be simulated as a new query request, so as to expand the query request range of the historical query request according to the similarity between the historical query request and the new query request in the subsequent process.
[0057] S303: Calculate each of the multiple historical query request pairs respectively to obtain the similarities corresponding to the multiple historical query request pairs respectively, and use the maximum value among the similarities corresponding to the multiple historical query request pairs respectively as the expansion value of the historical query request set.
[0058] In this step, the similarity between the first historical query request and the second historical query request in the target historical query request pair can be calculated to obtain the similarity between the first historical query request and the second historical query request. Further, the similarity between the historical query requests in each historical query request pair can be calculated to obtain the similarity corresponding to each historical query request pair, so as to select the maximum value among the similarities corresponding to each historical query request pair as the expansion value of the historical query request set.
[0059] In this way, in this application, the historical query request with a relatively later query time in the historical query request pair is simulated as a new query request, and then the query request range of the historical query requests in the historical query request set is expanded according to the maximum similarity between the historical query request and the new query request, which can improve the query efficiency. Among them, the query dimension of the first historical query request and the query dimension of the second historical query request can be the same, and the query dimension of the first historical query request and the query dimension of the second historical query request can be different, but the query requests between the first historical query request and the second historical query request need to be partially overlapped.
[0060] As Figure 4 shown, Figure 4 is a schematic diagram of the similarity of the historical query request pair of a data partitioning method provided by an embodiment of this application. In Figure 4 (in Figure 4In the figure, solid circles are used to represent historical query requests. The set of pseudo-historical query requests includes historical query request pairs Q1 and Q2. Among them, historical query request pair Q1 includes historical query requests q1 and q2, and the query times of historical query requests q1 and q2 are different; historical query request pair Q2 includes historical query requests q3 and q4, and the query times of historical query requests q3 and q4 are different. As Figure 4 can be seen, the overlapping part of historical query requests q1 and q2 is smaller than the overlapping part of historical query requests q3 and q4, that is, the similarity of historical query request pair Q1 is less than the similarity of historical query request pair Q2. Therefore, the similarity of historical query request pair Q2 can be used as the expansion value of the set of historical query requests to expand the historical query requests in the set of historical query requests.
[0061] In an implementable embodiment, in this step, calculating multiple historical query request pairs respectively to obtain the similarities corresponding to the multiple historical query request pairs can specifically include steps S3031 - S3033. As Figure 5 shown, Figure 5 is a flowchart of another data partitioning method provided by an embodiment of the present application. Figure 5 It includes the following steps S3031 - S3033:
[0062] S3031: Perform graphic transformation on the first historical query request in the target historical query request pair to obtain the first regular rectangle corresponding to the first historical query request, and perform graphic transformation on the second historical query request in the target historical query request pair to obtain the second regular rectangle corresponding to the second historical query request.
[0063] It can be understood that in this step, the first historical query request in the target historical query request pair and the second historical query request in the target historical query request pair can be subjected to graphic transformation. Correspondingly, in the present application, the historical query requests in each historical query request pair can be subjected to graphic transformation so that the similarities between the historical query requests in each historical query request pair can be calculated in a concrete manner in the subsequent process, thereby obtaining the similarities between the pairwise historical query requests in each historical query request pair.
[0064] S3032: Calculate the first regular rectangle and the second regular rectangle to obtain the distance between the first regular rectangle and the second regular rectangle.
[0065] Next, in combination with Figure 6 introduce the process of calculating the distance of the regular rectangle in the present application. Figure 6 is a schematic diagram of calculating the distance of a data partitioning method provided by an embodiment of the present application. In Figure 6Among them, the distance between the first regular rectangle X1 corresponding to the first historical query request in the target historical query request pair and the second regular rectangle X2 corresponding to the second historical query request in the target historical query request pair is calculated to obtain the distance δ between the first regular rectangle X1 and the second regular rectangle X2. In this way, the calculation process of calculating the similarity of query requests is visualized in this application, so that the calculation result can be obtained more quickly.
[0066] S3033: Convert the data type of the distance between the first regular rectangle and the second regular rectangle to obtain the similarity between the first historical query request and the second historical query request.
[0067] In this step, the data type corresponding to the distance can be converted to the data type corresponding to the similarity. For example, if the distance between the first regular rectangle and the second regular rectangle is 1, after the data type conversion, the similarity between the first historical query request and the second historical query request is 99%. It can be further understood that the minimum value of the distances between multiple pairwise regular rectangles in this application is the maximum value of the similarities between multiple pairwise query requests.
[0068] In another implementable embodiment, after calculating the pairs of regular rectangles corresponding to multiple historical query request pairs in this application to obtain the distances corresponding to the multiple pairs of regular rectangles, the minimum distance value is selected from the multiple distances as the expansion value of the historical query request set, and then the similarity between the historical query requests in the historical query request set is verified through this expansion value.
[0069] Specifically, the historical query request set can be evenly divided into two historical query request subsets, namely the first historical query request subset Q H and the second historical query request subset Q F . The second historical query request subset Q F is simulated as a new query request subset, and then a query request is selected from the first historical query request subset Q H and the second historical query request subset Q F respectively to form a query request pair (q h , q f ), and it is judged whether the maximum distance value of (q h , q f ) is less than or equal to the expansion value of the historical query request set. If the maximum distance value of (q h , q f ) is less than or equal to the expansion value of the historical query request set, it means that the query requests in the query request pair are similar. The expanded historical query request set can be obtained according to the expansion value of the historical query request set in the subsequent process, and the data partition corresponding to the data total set obtained by using this expanded historical query request set is used to query data, and the query efficiency is relatively high; if (qh , q f ), if the maximum distance value is greater than the extended value of the historical query request set, then increase the numerical value of the extended value of the historical query request set until (q h , q f )'s maximum distance value is less than or equal to the extended value of the historical query request set.
[0070] Furthermore, in this application, formula (1) can be used to obtain the maximum distance value of (q h , q f ). The specific formula (1) is as follows:
[0071]
[0072] Among them, dist(q h , q f ) represents the maximum distance value between the query request pair (qh, qf), q h .l d represents the lower limit value l of q h in the d-th dimension d , q h .u d represents the upper limit value u of q h in the d-th dimension d , q f .l d represents the lower limit value l of q f in the d-th dimension d , q f .u d represents the upper limit value u of q f in the d-th dimension d , d max represents the maximum dimension range of the d-th dimension. It should be noted that the upper limit value and the lower limit value usually refer to the range limit of the query request under its corresponding query dimension, and this range limit can be determined based on business requirements, data characteristics, or the input of the query object.
[0073] S304: Expand each historical query request in the historical query request set according to the extended value to obtain an extended historical query request set.
[0074] Specifically, after obtaining the extended value of the historical query request set, each historical query request in the historical query request set can be separately subjected to graphic transformation to obtain a regular rectangle corresponding to each historical query request, so that the regular rectangle can be extended according to the extended value in the subsequent process. Then, the extended value is added to each rectangular side of the regular rectangle corresponding to the target historical query request to obtain a regular extended rectangle corresponding to the target historical query request. Correspondingly, according to this extension method, the regular rectangles corresponding to each historical query request are extended respectively to obtain regular extended rectangles corresponding to each historical query request. Finally, according to the regular extended rectangles corresponding to each historical query request, an extended historical query request set is obtained, where the target historical query request is any one of the multiple historical query requests. In this way, each historical query request in the historical query request set is extended concretely in this application to improve the efficiency of query request extension.
[0075] As Figure 7 shown, Figure 7 FIG. is a schematic diagram of query request extension for a data partitioning method provided by an embodiment of this application. Figure 7 It shows the process of extending the first regular rectangle X1 corresponding to the first historical query request in the target historical query pair to obtain a regular extended rectangle X1* corresponding to the first regular rectangle X1. It should be noted that Figure 7 the extension process in is only a partial example in this application. In practical applications, the regular rectangles corresponding to each historical query request in the historical query request set can also be extended. Specifically, the extended value δ is added to each side of the first regular rectangle X1 to obtain the regular extended rectangle X1*.
[0076] In an implementable embodiment, the formula (2) can also be used to extend each historical query request in the historical query request set. The formula (2) is specifically as follows:
[0077]
[0078] Wherein, represents the lower limit value l of the historical query request after extension in the d-th dimension of d , represents the upper limit value u of the historical query request after extension in the d-th dimension of d , represents the historical query request q i The lower limit value l in the x-axis direction of the d-th dimension is d increased by the extended value δ, represents the historical query request q i The lower limit value l in the y-axis direction of the d-th dimension isd An extended value δ is added. denotes a historical query request q i the upper limit value u in the x-axis direction of the d-th dimension d An extended value δ is added. denotes a historical query request q i the upper limit value u in the y-axis direction of the d-th dimension d An extended value δ is added.
[0079] Furthermore, formula (3) can be obtained from formula (2). Formula (3) is mainly used to reduce the query request range of each historical query request in the historical query request set according to the change of data analysis requirements, the optimization mechanism of automated query, and the input of query objects. Formula (3) is specifically as follows:
[0080]
[0081] wherein, denotes the lower limit value l of the reduced historical query request in the d-th dimension d , denotes the upper limit value u of the reduced historical query request in the d-th dimension d , denotes a historical query request q i ′ the lower limit value l in the x-axis direction of the d-th dimension d The extended value δ is reduced. denotes a historical query request q i ′ the lower limit value l in the y-axis direction of the d-th dimension d The extended value δ is reduced. denotes a historical query request q i ′ the upper limit value u in the x-axis direction of the d-th dimension d The extended value δ is reduced. denotes a historical query request q i ′ the upper limit value u in the y-axis direction of the d-th dimension d The extended value δ is reduced.
[0082] S305: Perform a partitioning operation on the total data set according to the extended historical query request set to obtain the data partition corresponding to the total data set.
[0083] Specifically, in this step, first, according to the expanded historical query request set, a data subset corresponding to the expanded historical query request set in the data total set can be obtained. It should be noted that the data subset corresponding to the expanded historical query request set in the data total set is different from the data subset corresponding to the historical query request set in the data total set, and the data subset corresponding to the expanded historical query request set includes the data subset corresponding to the historical query request set. It can be seen that through the historical query request expansion scheme of the present application, the expansion of the data subset corresponding to the historical query request set is further realized, thereby improving the data query efficiency.
[0084] Then, according to the data subset corresponding to the expanded historical query request set, the remaining query data subsets in the data total set are updated to obtain the updated remaining query data subsets. It should be noted that since the data subset corresponding to the expanded historical query request set includes the data subset corresponding to the historical query request set, the spatial range of the updated remaining query data subsets is smaller than the spatial range of the remaining query data subsets. Finally, the data subset corresponding to the expanded historical query request set is divided into the first data partition in the data total set, and the updated remaining query data subsets are divided into the second data partition in the data total set. It can be seen that through the embodiments of the present application, the spatial range of the remaining query data subsets in the data total set can be reduced, and the spatial range of the data subset corresponding to the expanded historical query request set can be increased, thus realizing the improvement of the query efficiency.
[0085] In an implementable specific example, a recursive algorithm can be used to partition the data total set, where the recursive algorithm includes an irregular splitting algorithm and a rectangular splitting algorithm. It can be understood that in the present application, the rectangular splitting algorithm can be used to divide the data subset corresponding to the expanded historical query request set into regular rectangular partitions in the data total set, and the irregular splitting algorithm can be used to divide the updated remaining query data subsets into irregularly shaped partitions in the data total set. In this way, the data total set is partitioned using regular rectangular partitions and irregularly shaped partitions to improve the query performance and reduce the query cost.
[0086] Furthermore, the present application can use formula (4) to obtain the query cost of the data partition in the data total set for the data subset corresponding to the expanded historical query request set Formula (4) is specifically as follows:
[0087]
[0089] Wherein, represents the number of the expanded historical query request sets, represents the query cost of the expanded historical query request sets in the data total set, denotes a historical query request in the expanded historical query request set, denotes the data partition layout of the total data set, p i denotes the i-th data partition in the data partition layout, p i .size denotes the partition space size of the i-th data partition, The value of can be rounded off.
[0090] It should be noted that after obtaining the first data partition by using the rectangle splitting algorithm to perform regular rectangle partitioning on the data subset corresponding to the expanded historical query request set in this application, the rectangle splitting algorithm can be further used to split the first data partition until the split partition layout meets or approaches the optimal partition layout. It can be understood that the query cost of the split data partition meets or approaches the query cost of the optimal data partition. The query cost of the optimal data partition can be determined according to the expanded historical query request set. For example, the query cost of each data partition of the data subset corresponding to the expanded historical query request set is 128MB.
[0091] Such as Figure 8 shown Figure 8 is a schematic diagram of partition splitting of a data partition method provided by an embodiment of this application. In Figure 8 (a), the proposed first data partition is P1, and then P1 is split to obtain data partition P1_A and data partition P1_B. Further, data partition P1_A and data partition P1_B can also be split until the query cost of the split data partition meets the query cost of the optimal partition layout, that is, splitting data partition P1_A into data partition P A and data partition P B , splitting data partition P1_B into data partition P C and data partition P D . It should be noted that in this application, a partition tree can be used for data partitioning. This partition tree can be a QD tree, which is not specifically limited here. Such as Figure 8 (b) shown, where the leaf nodes A, B, C, and D respectively correspond to the final data partitions P A , P B , P C , P D in the distributed storage system. It should also be noted that in this application, there is no physical storage for the partitions of non-leaf nodes, that is, the second data partition, and it can be maintained in the main node for convenient query routing.
[0092] In an achievable implementation, before the first data partition P1 is split in this application, the query cost of the first data partition P1 can also be determined. If the query cost of the first data partition P1 is greater than or equal to 2·b min , then the first data partition P1 is split using the rectangular splitting algorithm; if the query cost of the first data partition P1 is greater than or equal to α·b min , then the first data partition P1 is split using the irregular splitting algorithm and the rectangular splitting algorithm. Where b min is the minimum query cost of the data partition, and α is a constant greater than 1 and less than 2. In this way, this application uses the irregular splitting algorithm and the rectangular splitting algorithm to optimize the query cost, and combines the irregular splitting algorithm and the rectangular splitting algorithm to split the data partition with a smaller query cost to improve the splitting efficiency.
[0093] Furthermore, before obtaining the data subset in the data set corresponding to the extended historical query request set in this application, it can also be determined whether there is a partial overlap among the query requests in the extended historical query request set. It can be understood that determining whether there is a partial overlap situation is to determine whether there is an intersection situation.
[0094] Specifically, if there is no partial overlap among the query requests in the extended historical query request set, then according to the query requests without partial overlap in the extended historical query request set, the data subset in the data set corresponding to the query requests without partial overlap in the extended historical query request set is determined; or; if there is a partial overlap among the query requests in the extended historical query request set, then according to the query requests with partial overlap in the extended historical query request set, the data subset in the data set corresponding to the query requests with partial overlap in the extended historical query request set is determined. It can be seen that this application determines the data subset in the data set according to different situations of the query requests in the extended historical query request set. In this way, the data partition can be quickly constructed according to the data subset in the subsequent process.
[0095] Even further, the process of dividing the data subset corresponding to the extended historical query request set into the first data partition in the data set can be divided into the following two situations:
[0096] Situation 1: Directly perform a partitioning operation on the data subset corresponding to the query requests without partial overlap to obtain a regular rectangular partition corresponding to the data subset corresponding to the query requests without partial overlap, and use the regular rectangular partition corresponding to the data subset corresponding to the query requests without partial overlap as the first data partition in the data set.
[0097] As Figure 9 shown, Figure 9Schematic diagram of regular rectangle partitioning for a data partitioning method provided by an embodiment of the present application. In Figure 9 (a), if there is no partial overlap between the query request q 1 * in the pseudo-historical query request set and the query request q 2 * in the historical query request set, then partition the data subset corresponding to the query request q 1 * and the data subset corresponding to the query request q 2 *; as shown in Figure 9 (b), partition the data subset corresponding to the query request q 1 * into the data partition QP_1, and partition the data subset corresponding to the query request q 2 * into the data partition QP_2.
[0098] Case 2: Perform a partitioning operation on the data subsets corresponding to the query requests with partial overlap. Divide the data subsets corresponding to the query requests with partial overlap into regular rectangle partitions respectively, and then merge the regular rectangle partitions with partial overlap into the minimum bounding rectangle partition, so as to use the minimum bounding rectangle partition as the minimum regular rectangle partition corresponding to the data subset corresponding to the query request with partial overlap. Then, use the minimum regular rectangle partition corresponding to the data subset corresponding to the query request with partial overlap as the first data partition in the data set.
[0099] As shown in Figure 10 below, Figure 10 Schematic diagram of regular rectangle partitioning for another data partitioning method provided by an embodiment of the present application. In Figure 10 (a), if there is partial overlap between the query request q 3 * in the pseudo-historical query request set and the query request q 4 * in the historical query request set, then partition the data subset corresponding to the query request q 3 * and the data subset corresponding to the query request q 4 *; as shown in Figure 10 (b), use the minimum bounding rectangle partition corresponding to the data subset corresponding to the query request q 3 * and the data subset corresponding to the query request q 4 * as the data partition GP_1.
[0100] It should also be noted that in the present application, a partitioning operation can also be performed on the remaining query data subsets after update to obtain the non-regular shape partitions corresponding to the remaining query data subsets after update, and use the non-regular shape partitions corresponding to the remaining query data subsets after update as the second data partition.
[0101] As shown in Figure 11 below, Figure 11This is a schematic diagram of an irregular shape partition for a data partitioning method provided by an embodiment of the present application. In Figure 11 , data partition QP_1, data partition QP_2, and data partition GP_1 are the first data partitions, and data partition KP is the second data partition. It can be understood that data partition KP = data partition P0 - data partition QP_1 - data partition QP_2 - data partition GP_1. It can be seen that in the present application, the data subset corresponding to the expanded historical query request set is divided into regular rectangular partitions in the data set, and the remaining query data subsets after update are divided into irregular shape partitions in the data set. In this way, the data set is partitioned using regular rectangular partitions and irregular shape partitions to improve query performance and reduce query costs.
[0102] In some examples, after performing a partitioning operation on the data subset corresponding to the query requests with partial overlap to obtain the smallest regular rectangular partition corresponding to the data subset corresponding to the query requests with partial overlap, a partitioning operation can also be performed on the data subset corresponding to the query requests with partial overlap to obtain an irregular shape partition corresponding to the data subset corresponding to the query requests with partial overlap, and then it is determined whether the query cost of the smallest regular rectangular partition corresponding to the data subset corresponding to the query requests with partial overlap is less than the query cost of the irregular shape partition corresponding to the data subset corresponding to the query requests with partial overlap.
[0103] If the query cost of the smallest regular rectangular partition corresponding to the data subset corresponding to the query requests with partial overlap is less than the query cost of the irregular shape partition corresponding to the data subset corresponding to the query requests with partial overlap, then the smallest regular rectangular partition corresponding to the data subset corresponding to the query requests with partial overlap is used as the first data partition in the data set; if the query cost of the smallest regular rectangular partition corresponding to the data subset corresponding to the query requests with partial overlap is greater than or equal to the query cost of the irregular shape partition corresponding to the data subset corresponding to the query requests with partial overlap, then the irregular shape partition corresponding to the data subset corresponding to the query requests with partial overlap is used as the first data partition in the data set. In this way, on the basis of obtaining the first data partition by performing a partitioning operation on the data set in the present application, the query cost of the partition is further determined, and the data partition with a smaller query cost is selected as the first data partition, thereby improving the query cost.
[0104] In some other examples, in the case where the query cost of the minimum regular rectangle partition corresponding to the data subset corresponding to the partially overlapping query requests is greater than or equal to the query cost of the irregular shape partition corresponding to the data subset corresponding to the partially overlapping query requests, the data subset corresponding to the partially overlapping query requests is partitioned to obtain a regular rectangle partition corresponding to the data subset corresponding to the partially overlapping query requests, and it is determined whether the query cost of the irregular shape partition corresponding to the data subset corresponding to the partially overlapping query requests is small or the regular rectangle partition corresponding to the data subset corresponding to the partially overlapping query requests is small. Among them, the data partition corresponding to the smaller query cost is used as the first data partition. In this way, in this application, the data with the minimum query cost is selected as the first data partition through multiple comparisons to improve the query cost.
[0105] As Figure 12 shown, Figure 12 is a partition schematic diagram of a data partition method provided by an embodiment of this application. In Figure 12 (a), the query request q 5 * in the set of pseudo-historical query requests and the query request q 6 * in the set of historical query requests are partially overlapping, and the query cost of the total data set is 735MB. Among them, the query request q 5 * and the query request q 6 * are partitioned to obtain an irregular shape partition GP_2 (as shown in Figure 12 (b)), and the query cost of the query request q 5 * in the irregular shape partition GP_2 is 465MB; the query request q 5 * and the query request q 6 * are partitioned to obtain a regular rectangle partition QP_3, a regular rectangle partition QP_4, a regular rectangle partition QP_5, and a regular rectangle partition QP_6. The query cost of the query request q 5 * in the regular rectangle partitions QP_3, QP_4, QP_5, and QP_6 is 420MB (180MB + 240MB, as shown in Figure 12 (c)).
[0106] It can be seen that although the query request q 5 * and the query request q 6 * are partially overlapping, the overlapping range of their partial overlap is very small. Therefore, the query cost of the minimum regular rectangle partition constructed according to the query request q 5 * and the query request q 6 * will be very large. On this basis, in this application, by judging the query request q 5 * and the query request q 6*Is the query cost of the constructed regular rectangular partition smaller, or is the query cost of the constructed non-regular shaped partition smaller, to further determine the query request q 5 *and the query request q 6 *The constructed data partition improves the flexibility of the partitioning operation.
[0107] In an implementable embodiment, after the present application divides the remaining query data subsets after the update into the second data partition in the data set, it can also obtain an amplification factor and a minimum partition space range threshold, where the minimum partition space range threshold is 128MB, and the amplification factor is used to expand the partition space range of the first data partition and the second data partition. After that, it is judged whether the partition space range corresponding to the first data partition is less than the minimum partition space range threshold, or whether the partition space range corresponding to the second data partition is less than the minimum partition space range threshold. If the partition space range corresponding to the first data partition is less than the minimum partition space range threshold, the partition space range corresponding to the first data partition is expanded according to the amplification factor; or; if the partition space range corresponding to the second data partition is less than the minimum partition space range threshold, the partition space range corresponding to the second data partition is expanded according to the amplification factor, so as to meet the minimum partition space range threshold by expanding the partition space range of the data partition. It can be understood that the partition space range of the data partition is the query cost of the data partition.
[0108] Furthermore, in the present application, formulas (5), (6) and (7) can be used to obtain the partition space range corresponding to the expanded data partition. Formulas (5), (6) and (7) are specifically as follows:
[0109]
[0110]
[0111]
[0112] Among them, represents the center point vector of the data partition QP in the d-th dimension, represents the radius vector of the data partition QP in the d-th dimension, QP.l d represents the lower limit value l of the data partition QP in the d-th dimension d , QP.u d represents the upper limit value u of the data partition QP in the d-th dimension d , QP′ is the partition space range corresponding to the expanded data partition, f is the amplification factor, the amplification factor is preset, the amplification factor can be 1.5, and no specific limitation is made here.
[0113] Such as Figure 13 shownFigure 13 This is an enlarged partition schematic diagram of a data partitioning method provided by an embodiment of the present application. In Figure 13 , the data partition QP is expanded according to formulas (5), (6), and (7) to obtain an enlarged data partition QP'. It should be noted that the data partition QP can be a regular rectangular partition or an irregular shape partition. This is only an example in the present application and is not specifically limited herein. Figure 13 In the present application
[0114] In another implementable embodiment, the present application can also expand the data partition according to the historical query requests in the historical query request set. Specifically (the present application takes the data partition QP as an example for illustration), a smaller area QP_A is split from the data partition QP, and then the relative position of each historical query request in the data partition QP within the data partition QP is recorded. This relative position is obtained based on the center point vector and radius vector of the data partition QP. After that, the relative position of each historical query request in the area QP_A within the area QP_A can also be recorded, and the relative positions in the area QP_A are sorted in ascending order, and the historical query requests in QP_A are assigned to the data partition QP one by one according to this sorting until the partition space range of the data partition QP meets the minimum partition space range threshold. In this way, the relationship between the data partition size and the query efficiency is effectively balanced. By intelligently allocating the historical query requests, it is ensured that each data partition is neither too small nor can reflect the spatial distribution characteristics of the data.
[0115] Specifically, in the present application, formula (8) can be used to obtain the relative position. Formula (8) is specifically as follows:
[0116]
[0117] where F QP (x) represents the relative position of the historical query request x in the data partition QP, and x d represents that the dimension corresponding to the historical query request x is the d-th dimension. It should be noted that the dimension in the present application is different features or attributes of the data in the data set.
[0118] In the embodiments of the present application, after partitioning the total data set according to the expanded historical query request set to obtain the data partitions corresponding to the total data set, a target query request may be obtained, and then the target query request is used to query in the data partitions corresponding to the total data set to obtain the target query data corresponding to the target query request in the total data set. Among them, the target query request may be a historical query request, or the target query request may be a new query request. It can be understood that if the target query request is a historical query request, the first data partition is used to query the target query request; if the target query request is a new query request, the second data partition is used to query the target query request. In this way, according to different situations of the query request, different data partitions are used to query the query request, which improves the query efficiency and reduces the query cost.
[0119] Further, in combination with Figure 14 to illustrate the query process of the present application. Figure 14 FIG. is a schematic diagram of the query process of a data partitioning method provided by an embodiment of the present application. In Figure 14 it, the distributed storage system receives an SQL query request, and then the query rewriter determines whether the SQL query request is a historical query request or a new query request. After that, the query router determines the data partition corresponding to the SQL query request in the data partitions corresponding to the total data set, so as to query the SQL query request according to the data partition corresponding to the SQL query request.
[0120] In some other examples, the query rewriter may also determine the query range of the SQL query request to determine the data partition corresponding to the SQL query request according to the query range of the SQL query request. For example: the SQL query request of A>=10 AND B<=50 corresponds to the query range [10,∞)×(-∞,50], and the query router may decompose the query range into two non-overlapping query ranges [10,∞)×(-∞,∞) and (-∞,10)×(-∞,50] to determine the corresponding data partitions through the two non-overlapping query ranges.
[0121] It should also be noted that the data partitioning scheme in this application can be applied not only to distributed storage systems, but also to database management systems, log management systems, and big data partitioning systems, which improves the strong reliability for querying and storing data within the system. In summary, in the embodiments of this application, historical query requests are simulated as new query requests based on different query times, so as to expand the query range of historical query requests in the historical query request set through the maximum similarity between multiple query request pairs, and two partitioning schemes, namely regular rectangle partitioning and irregular shape partitioning, are also proposed to partition the total data set, thereby improving query efficiency, reducing query costs, and improving the robustness, reliability, and accuracy of system storage and query data.
[0122] Based on the data partitioning method provided in the foregoing embodiments, this application also correspondingly provides a data partitioning device. The data partitioning device provided in the embodiments of this application will be specifically introduced below.
[0123] See Figure 15 , which is a schematic structural diagram of a data partitioning device provided in the embodiments of this application. As Figure 15 shown, the data partitioning device specifically includes:
[0124] A request set acquisition unit 1501, configured to acquire a historical query request set and a total data set, where the total data set includes a data subset corresponding to the historical query request set and the remaining query data subsets;
[0125] A request pair acquisition unit 1502, configured to obtain multiple historical query request pairs in the historical query request set according to the query times of the historical query request set, where the query time of the first historical query request in the target historical query request pair is earlier than the query time of the second historical query request in the historical query request pair, and the target historical query request pair is any one of the multiple historical query request pairs;
[0126] An expansion value acquisition unit 1503, configured to calculate each of the multiple historical query request pairs to obtain the similarity corresponding to each of the multiple historical query request pairs, and use the maximum value among the similarities corresponding to the multiple historical query request pairs as the expansion value of the historical query request set;
[0127] An expanded request set acquisition unit 1504, configured to expand each historical query request in the historical query request set according to the expansion value to obtain an expanded historical query request set;
[0128] A total data set partitioning unit 1505, configured to perform a partitioning operation on the total data set according to the expanded historical query request set to obtain a data partition corresponding to the total data set.
[0129] Optionally, the extended value obtaining unit 1503 is specifically configured to:
[0130] Perform graphic transformation on the first historical query request in the target historical query request pair to obtain a first regular rectangle corresponding to the first historical query request, and perform graphic transformation on the second historical query request in the target historical query request pair to obtain a second regular rectangle corresponding to the second historical query request;
[0131] Calculate the first regular rectangle and the second regular rectangle to obtain the distance between the first regular rectangle and the second regular rectangle;
[0132] Perform data type conversion on the distance between the first regular rectangle and the second regular rectangle to obtain the similarity between the first historical query request and the second historical query request.
[0133] Optionally, the extended request set obtaining unit 1504 is specifically configured to:
[0134] Perform graphic transformation on each historical query request in the historical query request set to obtain a regular rectangle corresponding to each historical query request;
[0135] Increase the rectangular sides of the regular rectangle corresponding to each historical query request by the extended value to obtain a regular extended rectangle corresponding to each historical query request;
[0136] Obtain an extended historical query request set according to the regular extended rectangle corresponding to each historical query request.
[0137] Optionally, the data total set partitioning unit 1505 includes:
[0138] An extended data subset obtaining unit, configured to obtain a data subset in the data total set corresponding to the extended historical query request set according to the extended historical query request set;
[0139] An updated data subset obtaining unit, configured to update the remaining query data subsets according to the data subset corresponding to the extended historical query request set to obtain updated remaining query data subsets;
[0140] A data partitioning obtaining unit, configured to divide the data subset corresponding to the extended historical query request set into a first data partition in the data total set, and divide the updated remaining query data subsets into a second data partition in the data total set.
[0141] Optionally, the apparatus further includes:
[0142] A partial overlap determination unit, configured to determine whether there is a partial overlap among the query requests in the expanded historical query request set;
[0143] The expanded data subset obtaining unit is specifically configured to:
[0144] If there is no partial overlap among the query requests in the expanded historical query request set, determine, according to the query requests without partial overlap in the expanded historical query request set, the data subset in the total data set corresponding to the query requests without partial overlap in the expanded historical query request set; or,
[0145] If there is a partial overlap among the query requests in the expanded historical query request set, determine, according to the query requests with partial overlap in the expanded historical query request set, the data subset in the total data set corresponding to the query requests with partial overlap in the expanded historical query request set.
[0146] Optionally, the data partition obtaining unit includes:
[0147] A regular rectangle partition obtaining unit, configured to perform a partitioning operation on the data subset corresponding to the query requests without partial overlap to obtain a regular rectangle partition corresponding to the data subset corresponding to the query requests without partial overlap;
[0148] A regular rectangle partition conversion unit, configured to use the regular rectangle partition corresponding to the data subset corresponding to the query requests without partial overlap as the first data partition in the total data set; or,
[0149] A minimum regular rectangle partition obtaining unit, configured to perform a partitioning operation on the data subset corresponding to the query requests with partial overlap to obtain a minimum regular rectangle partition corresponding to the data subset corresponding to the query requests with partial overlap;
[0150] A minimum regular rectangle partition conversion unit, configured to use the minimum regular rectangle partition corresponding to the data subset corresponding to the query requests with partial overlap as the first data partition in the total data set.
[0151] Optionally, the apparatus further includes:
[0152] A first irregular shape partition obtaining unit, configured to perform a partitioning operation on the data subset corresponding to the query requests with partial overlap to obtain an irregular shape partition corresponding to the data subset corresponding to the query requests with partial overlap;
[0153] The minimum regular rectangle partition conversion unit is specifically configured to:
[0154] If the query cost of the minimum regular rectangular partition corresponding to the data subset corresponding to the partially overlapping query requests is less than the query cost of the irregular-shaped partition corresponding to the data subset corresponding to the partially overlapping query requests, then the minimum regular rectangular partition corresponding to the data subset corresponding to the partially overlapping query requests is taken as the first data partition in the data set.
[0155] Optionally, the data partition obtaining unit includes:
[0156] A second irregular-shaped partition obtaining unit, configured to perform a partitioning operation on the remaining query data subsets after the update to obtain an irregular-shaped partition corresponding to the remaining query data subsets after the update;
[0157] A second data partition obtaining unit, configured to take the irregular-shaped partition corresponding to the remaining query data subsets after the update as the second data partition.
[0158] Optionally, the apparatus further includes:
[0159] A magnification factor obtaining unit, configured to obtain a magnification factor and a minimum partition space range threshold, where the magnification factor is used to expand the partition space ranges of the first data partition and the second data partition;
[0160] A first partition expansion unit, configured to, if the partition space range corresponding to the first data partition is less than the minimum partition space range threshold, expand the partition space range corresponding to the first data partition according to the magnification factor; or,
[0161] A second partition expansion unit, configured to, if the partition space range corresponding to the second data partition is less than the minimum partition space range threshold, expand the partition space range corresponding to the second data partition according to the magnification factor.
[0162] Optionally, the apparatus further includes:
[0163] A target query request obtaining unit, configured to obtain a target query request;
[0164] A target query data obtaining unit, configured to query in the data partitions corresponding to the data set according to the target query request to obtain target query data in the data set corresponding to the target query request.
[0165] An embodiment of the present application provides a computer device, and the computer device may be a server. Figure 16FIG. 0 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 900 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 922 (e.g., one or more processors) and a memory 932, and one or more storage media 930 (e.g., one or more mass storage devices) for storing application programs 942 or data 944. Among them, the memory 932 and the storage media 930 may be transient storage or persistent storage. The programs stored in the storage media 930 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 922 may be configured to communicate with the storage media 930 and execute a series of instruction operations in the storage media 930 on the server 900.
[0166] The server 900 may further include one or more power supplies 926, one or more wired or wireless network interfaces 950, one or more input / output interfaces 958, and / or one or more operating systems 941, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0167] Among them, the CPU 922 is used to execute the following steps:
[0168] Obtain a historical query request set and a total data set, where the total data set includes a data subset corresponding to the historical query request set and the remaining query data subsets;
[0169] According to the query time of the historical query request set, obtain a plurality of historical query request pairs in the historical query request set, where the query time of the first historical query request in the target historical query request pair is earlier than the query time of the second historical query request in the historical query request pair, and the target historical query request pair is any one of the plurality of historical query request pairs;
[0170] Calculate each of the plurality of historical query request pairs respectively to obtain the similarities corresponding to the plurality of historical query request pairs respectively, and use the maximum value among the similarities corresponding to the plurality of historical query request pairs respectively as the expansion value of the historical query request set;
[0171] Expand each historical query request in the historical query request set according to the expansion value to obtain an expanded historical query request set;
[0172] Partition the total data set according to the expanded historical query request set to obtain the data partitions corresponding to the total data set.
[0173] An embodiment of the present application further provides another computer device, which may be a terminal device. As Figure 17 shown, for ease of explanation, only the parts related to the embodiments of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiments of the present application. Taking the terminal device as a mobile phone as an example:
[0174] Figure 17 Shown is a block diagram of a part of the structure of the mobile phone provided by the embodiments of the present application. Refer to Figure 17 , the mobile phone includes: a radio frequency (full English name: Radio Frequency, English abbreviation: RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a wireless fidelity (full English name: wirelessfidelity, English abbreviation: WiFi) module 1070, a processor 1080, and a power supply 1090 and other components. Those skilled in the art can understand that Figure 17 the mobile phone structure shown in
[0175] does not limit the mobile phone, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 17 The following specifically introduces each component of the mobile phone:
[0176] The RF circuit 1010 can be used for receiving and transmitting information or signals during communication. Specifically, after receiving the downlink information from the base station, it is sent to the processor 1080 for processing. Additionally, the uplink data is sent to the base station. Generally, the RF circuit 1010 includes, but is not limited to, antennas, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (full English name: Low Noise Amplifier, English abbreviation: LNA), a duplexer, etc. Moreover, the RF circuit 1010 can also communicate with the network and other devices via wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (full English name: Global System of Mobile communication, English abbreviation: GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (full English name: Code Division Multiple Access, English abbreviation: CDMA), Wideband Code Division Multiple Access (full English name: Wideband Code Division Multiple Access, English abbreviation: WCDMA), Long Term Evolution (full English name: Long Term Evolution, English abbreviation: LTE), email, Short Messaging Service (full English name: Short Messaging Service, SMS), etc.
[0177] The memory 1020 can be used to store software programs and modules. The processor 1080 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1020. The memory 1020 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store the data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 1020 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0178] The input unit 1030 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 1030 may include a touch panel 1031 and other input devices 1032. The touch panel 1031, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel 1031), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 1031 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 1080, and can receive and execute commands sent by the processor 1080. In addition, the touch panel 1031 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1031, the input unit 1030 may further include other input devices 1032. Specifically, the other input devices 1032 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0179] The display unit 1040 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 1040 may include a display panel 1041. Optionally, the display panel 1041 can be configured in forms such as a liquid crystal display (English full name: Liquid Crystal Display, English abbreviation: LCD), an organic light-emitting diode (English full name: Organic Light-Emitting Diode, English abbreviation: OLED), etc. Further, the touch panel 1031 can cover the display panel 1041. When the touch panel 1031 detects a touch operation on or near it, it transmits it to the processor 1080 to determine the type of touch event. Subsequently, the processor 1080 provides corresponding visual output on the display panel 1041 according to the type of touch event. Although in Figure 17 the touch panel 1031 and the display panel 1041 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 1031 and the display panel 1041 can be integrated to realize the input and output functions of the mobile phone.
[0180] The mobile phone may further include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1041 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1041 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the mobile phone can also be configured with, they will not be elaborated here.
[0181] The audio circuit 1060, the speaker 1061, and the microphone 1062 can provide an audio interface between the user and the mobile phone. The audio circuit 1060 can transmit the electrical signal converted from the received audio data to the speaker 1061, and the speaker 1061 converts it into a sound signal for output; on the other hand, the microphone 1062 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1060 and then converted into audio data. After the audio data is output to the processor 1080 for processing, it is sent through the RF circuit 1010 to, for example, another mobile phone, or the audio data is output to the memory 1020 for further processing.
[0182] WiFi belongs to short-range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 1070, which provides users with wireless broadband Internet access. Although Figure 17 the WiFi module 1070 is shown, it can be understood that it does not belong to an essential component of the mobile phone and can be completely omitted within the scope of not changing the essence of the invention according to needs.
[0183] The processor 1080 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1020, and calling data stored in the memory 1020, it executes various functions of the mobile phone and processes data, thereby collecting overall data and information of the mobile phone. Optionally, the processor 1080 may include one or more processing units; preferably, the processor 1080 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1080 either.
[0184] The mobile phone further includes a power source 1090 (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the processor 1080 through a power management system, so as to implement functions such as charging management, discharging management, and power consumption management through the power management system.
[0185] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated herein.
[0186] In the embodiment of the present application, the processor 1080 included in the mobile phone further has the following functions:
[0187] Obtain a historical query request set and a data total set, where the data total set includes a data subset corresponding to the historical query request set and the remaining query data subsets;
[0188] According to the query time of the historical query request set, obtain multiple historical query request pairs in the historical query request set, where the query time of the first historical query request in the target historical query request pair is earlier than the query time of the second historical query request in the historical query request pair, and the target historical query request pair is any one of the multiple historical query request pairs;
[0189] Calculate each of the multiple historical query request pairs respectively to obtain the similarity corresponding to each of the multiple historical query request pairs, and use the maximum value among the similarities corresponding to each of the multiple historical query request pairs as the expansion value of the historical query request set;
[0190] Expand each historical query request in the historical query request set according to the expansion value to obtain an expanded historical query request set;
[0191] Perform a partitioning operation on the data total set according to the expanded historical query request set to obtain a data partition corresponding to the data total set.
[0192] The embodiment of the present application further provides a computer-readable storage medium for storing a computer program. When the computer program runs on a computer device, the computer device is caused to execute any one of the implementation manners of a data partitioning method described in the foregoing various embodiments.
[0193] The embodiment of the present application further provides a computer program product including a computer program. When it runs on a computer device, the computer device is caused to execute any one of the implementation manners of a data partitioning method described in the foregoing various embodiments.
[0194] Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes of the above-described system and device can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein.
[0195] In several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the system is only a logical function division. In actual implementation, there may be other division methods. For example, multiple systems can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0196] The systems described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0197] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0198] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (English full name: Read-Only Memory, English abbreviation: ROM), random access memories (English full name: Random Access Memory, English abbreviation: RAM), magnetic disks or optical discs that can store computer programs.
[0199] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.
[0200] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data partitioning method, characterized in that: include: Acquire a historical query request set and a total data set, wherein the total data set includes a data subset corresponding to the historical query request set and other query data subsets; According to the query time of the historical query request set, a plurality of historical query request pairs in the historical query request set are obtained, wherein the query time of a first historical query request in a target historical query request pair is earlier than the query time of a second historical query request in the historical query request pair, and the target historical query request pair is any one of the plurality of historical query request pairs; Calculating the multiple historical query request pairs respectively to obtain the similarities respectively corresponding to the multiple historical query request pairs, and taking the maximum value of the similarities respectively corresponding to the multiple historical query request pairs as the extension value of the historical query request set; Expand each historical query request in the historical query request set according to the expansion value to obtain an expanded historical query request set; The data set is partitioned according to the expanded historical query request set to obtain data partitions corresponding to the data set.
2. The method according to claim 1, characterized in that The calculating the plurality of historical query request pairs respectively to obtain the similarities respectively corresponding to the plurality of historical query request pairs includes: Performing a graphic conversion on a first historical query request in the target historical query request pair to obtain a first regular rectangle corresponding to the first historical query request, and performing a graphic conversion on a second historical query request in the target historical query request pair to obtain a second regular rectangle corresponding to the second historical query request; Calculating the first regular rectangle and the second regular rectangle to obtain a distance between the first regular rectangle and the second regular rectangle; The distance between the first regular rectangle and the second regular rectangle is converted into a data type to obtain the similarity between the first historical query request and the second historical query request.
3. The method according to claim 2, characterized in that The step of respectively extending each historical query request in the historical query request set according to the extension value to obtain an extended historical query request set includes: Performing graphic conversion on each historical query request in the historical query request set to obtain a regular rectangle corresponding to each historical query request; Adding the extension value to the rectangular sides of the regular rectangle corresponding to each of the historical query requests, respectively, to obtain the regular extended rectangle corresponding to each of the historical query requests; According to the rule expansion rectangle corresponding to each historical query request, an expanded historical query request set is obtained.
4. The method according to claim 3, characterized in that The performing a partitioning operation on the total data set according to the expanded historical query request set to obtain data partitions corresponding to the total data set includes: According to the expanded historical query request set, obtaining a data subset in the total data set corresponding to the expanded historical query request set; According to the data subset corresponding to the expanded historical query request set, the remaining query data subsets are updated to obtain updated remaining query data subsets; The data subset corresponding to the expanded historical query request set is divided into a first data partition in the total data set, and the remaining query data subset after the update is divided into a second data partition in the total data set.
5. The method according to claim 4, characterized in that Before obtaining the data subset corresponding to the expanded historical query request set in the total data set according to the expanded historical query request set, the method further includes: Determining whether there is partial overlap in the query requests in the expanded historical query request set; The step of obtaining, according to the expanded historical query request set, a data subset in the total data set corresponding to the expanded historical query request set comprises: If there is no partial overlap between the query requests in the expanded historical query request set, then determining, based on the query requests in the expanded historical query request set that do not have partial overlap, a data subset in the total data set corresponding to the query requests in the expanded historical query request set that do not have partial overlap; or, If there are partial overlaps in the query requests in the expanded historical query request set, a data subset corresponding to the query requests partially overlapping in the expanded historical query request set is determined in the total data set according to the query requests partially overlapping in the expanded historical query request set.
6. The method according to claim 5, characterized in that The step of dividing the data subset corresponding to the expanded historical query request set into a first data partition in the total data set includes: Performing a partitioning operation on the data subsets corresponding to the query requests without partial overlap to obtain regular rectangular partitions corresponding to the data subsets corresponding to the query requests without partial overlap; Using the regular rectangular partition corresponding to the data subset corresponding to the query request without partial overlap as the first data partition in the total data set; or Performing a partition operation on the data subsets corresponding to the partially overlapping query requests to obtain minimum regular rectangular partitions corresponding to the data subsets corresponding to the partially overlapping query requests; The smallest regular rectangular partition corresponding to the data subset corresponding to the query request with partial overlap is used as the first data partition in the total data set.
7. The method according to claim 6, characterized in that After performing a partitioning operation on the data subsets corresponding to the partially overlapping query requests to obtain the minimum regular rectangular partitions corresponding to the data subsets corresponding to the partially overlapping query requests, the method further includes: Performing a partition operation on the data subsets corresponding to the partially overlapping query requests to obtain irregular shape partitions corresponding to the data subsets corresponding to the partially overlapping query requests; The step of taking the smallest regular rectangular partition corresponding to the data subset corresponding to the query request with partial overlap as the first data partition in the total data set includes: If the query cost of the minimum regular rectangular partition corresponding to the data subset corresponding to the query request with partial overlap is less than the query cost of the irregular shape partition corresponding to the data subset corresponding to the query request with partial overlap, the minimum regular rectangular partition corresponding to the data subset corresponding to the query request with partial overlap is used as the first data partition in the total data set.
8. The method according to claim 4, characterized in that The step of dividing the updated remaining query data subset into a second data partition in the total data set comprises: Performing a partition operation on the remaining query data subsets after the update to obtain irregular shape partitions corresponding to the remaining query data subsets after the update; The irregular shape partition corresponding to the remaining query data subset after the update is used as the second data partition.
9. The method according to claim 4, characterized in that After dividing the updated remaining query data subset into the second data partition in the total data set, the method further includes: Obtaining an amplification factor and a minimum partition space range threshold, wherein the amplification factor is used to expand the partition space ranges of the first data partition and the second data partition; If the partition space range corresponding to the first data partition is smaller than the minimum partition space range threshold, the partition space range corresponding to the first data partition is enlarged according to the amplification factor; or If the partition space range corresponding to the second data partition is smaller than the minimum partition space range threshold, the partition space range corresponding to the second data partition is enlarged according to the magnification factor.
10. The method according to claim 1, characterized in that After performing a partition operation on the data set according to the expanded historical query request set to obtain data partitions corresponding to the data set, the method further includes: Get the target query request; According to the target query request, a query is performed in the data partition corresponding to the data set to obtain the target query data corresponding to the target query request in the data set.
11. A data partitioning device, characterized in that: include: A request set acquisition unit, used to acquire a historical query request set and a total data set, wherein the total data set includes a data subset corresponding to the historical query request set and other query data subsets; a request pair obtaining unit, configured to obtain a plurality of historical query request pairs in the historical query request set according to the query time of the historical query request set, wherein the query time of the first historical query request in the target historical query request pair is earlier than the query time of the second historical query request in the historical query request pair, and the target historical query request pair is any one of the plurality of historical query request pairs; an extension value obtaining unit, configured to calculate the plurality of historical query request pairs respectively, obtain the similarities respectively corresponding to the plurality of historical query request pairs, and use the maximum value of the similarities respectively corresponding to the plurality of historical query request pairs as the extension value of the historical query request set; an extended request set obtaining unit, configured to respectively extend each historical query request in the historical query request set according to the extended value, to obtain an extended historical query request set; The data set partitioning unit is used to perform a partitioning operation on the data set according to the expanded historical query request set to obtain data partitions corresponding to the data set.
12. A computer device, characterized in that: The device comprises a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the steps of the data partitioning method according to any one of claims 1 to 10 according to the instructions in the computer program.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a computer device, the steps of the data partitioning method according to any one of claims 1 to 10 are implemented.
14. A computer program product, characterized in that The invention comprises a computer program, which implements the steps of the data partitioning method according to any one of claims 1 to 10 when executed by a computer device.