A data index construction, data query method, device, medium and product
By assigning offsets and sorting to unsorted data, constructing position vectors and learned indexes, and using machine learning to fit the cumulative distribution function, the problem that learned indexes cannot be applied to unsorted data queries is solved, achieving efficient querying of unsorted data and reducing space costs.
Patent Information
- Application Number
- CN202511262436.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing learning indexing methods are limited in their application to unordered data query scenarios, cannot effectively accelerate the query process of non-primary key data, and have high space costs.
By assigning offsets to the original unsorted data and sorting it, a position vector and a learned index are constructed. A machine learning algorithm is used to fit the cumulative distribution function to generate a data index. By combining the position matching vector and hash value to filter invalid positions, fast location of unsorted data can be achieved.
It improves the efficiency of querying unordered data, reduces the space cost of data indexing, and enables efficient querying of unsorted data.
Smart Images

Figure CN120743913B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine learning technology, and in particular to a data index construction, data query method, apparatus, medium and product. Background Technology
[0002] Databases use primary keys to identify each record and typically also use them to organize and sort data for fast access; this data access method is called a primary key index. If a user needs to query based on a non-primary key field, directly using the primary key index is not efficient enough. A secondary index is usually used instead. Secondary indexes are data structures or data organization methods designed to accelerate queries on non-primary key columns. Traditional secondary indexes include prefix indexes, bitmap indexes, and dictionary indexes.
[0003] With the development of artificial intelligence in the database field, database researchers are using machine learning techniques to learn the distribution characteristics or patterns of sorted data, and replacing traditional index searches with search methods based on data distribution fitting functions. This reduces the space cost of indexes and improves query performance. This indexing method that uses machine learning to replace traditional indexes to accelerate data queries is called a learned index.
[0004] However, existing learning indexing methods are typically based on primary key indexes because their underlying data is already sorted, making them more suitable for machine learning to capture. Learning indexes have a strong dependency on data order and cannot index or accelerate queries of non-primary key (unordered) data, thus limiting the application expansion of learning indexes in scenarios involving unordered data queries. Summary of the Invention
[0005] The purpose of this invention is to provide a data index construction, data query method, apparatus, medium, and product that can apply machine learning algorithms to unsorted data, construct data indexes, accelerate the query process of unsorted data, and reduce the space cost of data indexes.
[0006] To achieve the above objectives, embodiments of the present invention provide a method for constructing a data index, comprising:
[0007] Assign an offset to each data element at each original position in the original data; wherein the original data is unsorted data;
[0008] Sort each data element in the original data to form sorted data;
[0009] The offset of the data element at each sorting position in the sorted data is stored according to the sorting position to form a position vector; wherein, the position vector is used to associate the sorting position of the data element in the sorted data with the offset of the data element;
[0010] adopting a machine learning algorithm to build a learning index for the sorted data; wherein the learning index is used to predict the data element and the sorted position of the data element in the sorted data;
[0011] According to the learning index and the position vector, a data index of the original data is generated.
[0012] As an improvement of the above scheme, the method of adopting a machine learning algorithm to build a learning index for the sorted data comprises:
[0013] According to the data element and the sorted position in the sorted data, a cumulative distribution function is built;
[0014] Adopting a machine learning algorithm to fit the cumulative distribution function to learn the relevance of the data element and the sorted position in the sorted data, and build a learning index.
[0015] As an improvement of the above scheme, the method of adopting a machine learning algorithm to fit the cumulative distribution function to learn the relevance of the data element and the sorted position in the sorted data, and build a learning index comprises:
[0016] Adopting a spline interpolation method to divide the sorted data into a plurality of segmented data;
[0017] Adopting a machine learning algorithm to fit the cumulative distribution function to learn the relevance of the data element and the sorted position in each of the segmented data, and build a learning index.
[0018] As an improvement of the above scheme, after the method of sorting each data element in the original data to form sorted data, the method further comprises:
[0019] Calculating the hash value of the data element at each sorted position in the sorted data;
[0020] Storing the hash value of the data element at each sorted position according to the sorted position to form a position matching vector; wherein the position matching vector is used to associate the sorted position of the data element in the sorted data with the hash value of the data element;
[0021] Then, the method of generating a data index of the original data according to the learning index and the position vector comprises:
[0022] According to the learning index, the position vector and the position matching vector, a data index of the original data is generated.
[0023] As an improvement of the above scheme, the method further comprises:
[0024] Discarding the sorting data.
[0025] The embodiment of the present application further provides a data query method, comprising:
[0026] The original position of the target data element is located by using a preset data index, wherein the data index is obtained by using the data index construction method in any one of the above embodiments, and the data index comprises a position vector and a learning index.
[0027] As an improvement of the above scheme, the original position of the target data element is located by using a preset data index, and the locating comprises:
[0028] The target sorting position of the target data element is determined by using the learning index, wherein the learning index is used to predict the sorting position of a data element in the sorting data.
[0029] The target offset associated with the target sorting position is found according to the position vector, wherein the position vector is used to associate the sorting position of a data element in the sorting data with the offset of the data element, and the offset is used to identify the original position of the data element in the original data.
[0030] The original position of the target data element is located according to the target offset.
[0031] As an improvement of the above scheme, the learning index is a bounded-error learning index fitted according to a cumulative distribution function of the data elements and the sorting positions in the sorting data.
[0032] The target sorting position of the target data element is determined by using the learning index, and the determining comprises:
[0033] The sorting position of the target data element corresponding to an existing error is predicted according to the learning index, and the sorting position is recorded as a predicted sorting position.
[0034] The sorting position range is determined according to the predicted sorting position and a preset error value, wherein the sorting position range comprises a plurality of sorting positions.
[0035] The target sorting position is determined in the sorting position range.
[0036] As an improvement of the above scheme, the data index further comprises a position matching vector, and the position matching vector is formed by storing the hash value of the data element at each sorting position in the sorting data according to the sorting position.
[0037] determining a target sorting position in the sorting position range, comprising:
[0038] According to the position matching vector, a hash value associated with each sorting position in the sorting position range is found;
[0039] The hash value of the target data element is calculated;
[0040] The hash value of the target data element is matched with the hash value associated with each sorting position in the sorting position range to determine the target sorting position.
[0041] Embodiments of the present application also provide a data index construction device, comprising:
[0042] An offset allocation module is configured to allocate an offset to a data element at each original position in original data; wherein the original data is unsorted data;
[0043] A sorting data generation module is configured to sort each data element in the original data to form sorting data;
[0044] A position vector construction module is configured to store the offset of a data element at each sorting position in the sorting data according to the sorting position to form a position vector; wherein the position vector is used to associate the sorting position of the data element in the sorting data with the offset of the data element;
[0045] A learning index construction module is configured to use a machine learning algorithm to construct a learning index for the sorting data; wherein the learning index is used to predict the data element and the sorting position of the data element in the sorting data;
[0046] A data index generation module is configured to obtain a data index of the original data according to the learning index and the position vector.
[0047] Embodiments of the present application also provide a data query device, comprising:
[0048] A target data element index module is configured to use a preset data index to locate an original position of a target data element to be found; wherein the data index is constructed according to the data index construction method of any one of the above.
[0049] Embodiments of the present application also provide a terminal device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the data index construction method of any one of the above, or the data query method of any one of the above.
[0050] The embodiment of the present application also provides a computer readable storage medium, which comprises a stored computer program, wherein the computer program controls a device where the computer readable storage medium is located to perform the data index construction method according to any one of the preceding embodiments or the data query method according to any one of the preceding embodiments when the computer program is running.
[0051] The embodiment of the present application also provides a computer program product, which comprises a computer program or computer instructions, and the computer program or the computer instructions realize the data index construction method according to any one of the preceding embodiments or the data query method according to any one of the preceding embodiments when the computer program or the computer instructions are executed by a processor.
[0052] Compared with the prior art, the data index construction method, the data query method, the device, the medium and the product disclosed by the present application utilize the function of mapping the original position of non-sequenced original data by using a position vector, combine the data position prediction capability of a learning index, provide an efficient data index scheme for non-sequenced data, and solve the problem that the existing machine learning algorithm cannot be applied to non-sequenced data query. The learning index is extended and applied to the scene of non-sequenced data query, compared with the traditional two-level index mode, can accelerate the query process of non-sequenced data, improve the data query performance, and reduce the space cost of data index. BRIEF DESCRIPTION OF DRAWINGS
[0053] Figure 1 is a flowchart of a data index construction method provided by the embodiment of the present application;
[0054] Figure 2 is a flowchart of a preferred data index construction method in the embodiment of the present application;
[0055] Figure 3 is a flowchart of a data query method provided by the embodiment of the present application;
[0056] Figure 4 is a flowchart of a preferred data query method in the embodiment of the present application;
[0057] Figure 5 is an example diagram of data index construction and data query in a specific implementation scene in the embodiment of the present application;
[0058] Figure 6 is a structural schematic diagram of a data index construction device provided by the embodiment of the present application;
[0059] Figure 7 is a structural schematic diagram of a data query device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0060] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the protection scope of the present application.
[0061] In the description of the present application, it should be understood that the terms "center", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the purpose of facilitating the description of the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation to the present application.
[0062] The terms "first", "second" are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0063] In the description of the present application, it should be noted that, unless otherwise specified and limited, the terms "mounting", "connecting", "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integral connection; it can be mechanical connection, or electrical connection; it can be direct connection, or indirect connection through intermediate medium, or internal communication of two elements. For a person of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0064] Referring to Figure 1 is a flowchart of a data index construction method provided by an embodiment of the present application. The present embodiment provides a data index construction method, which includes steps S11 to S15:
[0065] S11, assigning an offset to each data element at a raw position in raw data; wherein the raw data is unsorted data;
[0066] S12, sorting each data element in the raw data to form sorted data;
[0067] S13, store the offset of the data element at each sorting position in the sorting data according to the sorting position to form a position vector, wherein the position vector is used to associate the sorting position of the data element in the sorting data with the offset of the data element;
[0068] S14, construct a learning index for the sorting data by using a machine learning algorithm, wherein the learning index is used to predict the data element and the sorting position of the data element in the sorting data;
[0069] S15, generate a data index of the original data according to the learning index and the position vector.
[0070] It should be noted that in actual application, the query demand for non-ordered data is large, and the population information table in the OLAP system is taken as an example, as shown in Table 1, the population information table takes region, birth date and gender as a composite sorting field, and identity card number, age and height as index fields. The data is first sorted according to the region, and in the case of the same region field, the birth date is sorted. If the region and the birth date are the same, the data is sorted according to the gender field. Although the population information table is a composite sorting table according to three fields, only the region field is actually ordered. Machine learning can only learn the data distribution rule of the region field, construct a learning index, and use a function to fit and calculate the position of a certain value. For the query of other fields, including birth date, gender, identity card number and the like, it is a query of non-primary key data, that is, a query of unordered data, and the user usually has a query demand for the identity card number or age and other non-primary key fields.
[0071] Table 1
[0072]
[0073] The embodiment of the application is used to solve the problem that the existing learning index cannot index non-ordered data, and proposes a data indexing method based on machine learning for non-ordered data, applies some existing machine learning models to non-ordered columns, replaces the traditional secondary index, and accelerates the analysis and query process of non-primary key columns.
[0074] In the embodiment of the present application, the basic structure of the data index of the original data includes a position vector and a learning index, the position vector is the ordering information of the unsorted data after ordering, any data structure capable of using subscript to index data can represent the position vector, and the simplest data structure is, for example, an array. Specifically, the i-th element of the position vector is the index of the i-th position data in the original data. In other words, if the user wants to find the smallest data in the original data, the value j of the subscript 0 of the position vector is obtained, and the data of the subscript j of the original data is the smallest value. It is the presence of the position vector that the present application can construct a learning index on the ordered data, and then map the prediction result of the learning index to the underlying unsorted original data.
[0075] The process of constructing the data index is to read the local data block in the data file to obtain the original data, and the original data includes a plurality of data elements. The original data refers to an unsorted data sequence, for example, the population information table described above, in which the data is organized by taking the region as the primary key, and the birth date data column, the age data column, and the ID number data column are all unsorted data.
[0076] An offset is assigned to each element in the original data, which is used to identify the original position of each data element in the data. For example, taking the age field as the original data, the original data includes data elements such as 44, 20, 67 in order, and each data element is assigned an offset of 0, 1, 2, respectively. Offset 0 indicates that the data element 44 is in the original position in the original data, and offset 1 indicates that the value 20 is in the original position in the original data.
[0077] Each data element in the original data is sorted according to a predetermined sorting standard to form sorted data. For example, the original data of the age field can be sorted according to the numerical value, the original data of the region section can be sorted according to the initial letter of the pinyin or the number of strokes, etc. Correspondingly, the offset of each data element in the original data is also reordered with the data element to form a position vector. The position vector maps the sorted data to the unsorted original data, and associates the ordering position of the data element in the sorted data with the offset of the data element, for example, the position vector value corresponding to the data 20 is 1, indicating that the original position order of the value 20 in the original data is 1, i.e., the second position in the original data.
[0078] Further, a machine learning algorithm is used to learn the data distribution characteristics of the sorted data, and a learning index is constructed for the sorted data, which describes the calculation method of the ordering position of a given target data element, i.e., the learning index associates the data element with the ordering position of the data element in the sorted data.
[0079] It can be understood that the sorted data is data that has been sorted, and a learning index can be constructed by using an existing machine learning algorithm to learn data distribution characteristics, and the learning index can be in the form of an RMI, a spline interpolation, or a linear regression, and specific construction means can be referred to in related technologies, and are not specifically limited herein.
[0080] According to the learning index and the position vector, a data index of the original data can be obtained. In an application process, a target sorting position of a target data element in the sorted data can be found according to the learning index, and an offset of the target data element can be determined according to the position vector, and then an original position of the target data element in the original data can be found, so that fast positioning of the target data element is realized.
[0081] By using the technical means of the embodiment of the present application, the original position of the non-sorted original data is mapped by using the position vector, and the data position prediction capability of the learning index is combined, so that an efficient data index scheme is provided for unsorted data, and the problem that existing machine learning algorithms cannot be applied to unordered data query is solved. The learning index is extended and applied to the unordered data query scene, compared with the traditional two-level index mode, the query process of the non-sorted data can be accelerated, the data query performance is improved, and the space cost of the data index is reduced.
[0082] As a preferred embodiment, the embodiment of the present application is further implemented on the basis of the above-mentioned embodiment, and reference is made to Figure 2 FIG. 14 is a flowchart of a preferred data index construction method in the embodiment of the present application, and step S14, that is, the step of constructing a learning index for the sorted data by using a machine learning algorithm, includes steps S141 to S142.
[0083] S141, constructing a cumulative distribution function according to data elements and sorting positions in the sorted data;
[0084] S142, fitting the cumulative distribution function by using a machine learning algorithm to learn the relevance of the data elements in the sorted data and the sorting positions, and constructing a learning index.
[0085] In the embodiment of the present application, a cumulative distribution function CDF is constructed based on the sorted data, the function describes the relationship between each data element and the sorting position of the data element in the sorted data, a machine learning algorithm is used to fit the cumulative distribution function to learn the relevance of each data element and the corresponding sorting position, and a prediction model for predicting the sorting position of a target data element is generated as the learning index.
[0086] In an alternative embodiment, the cumulative distribution function is fitted by a machine learning algorithm to learn the association between data elements in the sorting data and the sorting positions, to obtain a learning index, comprising:
[0087] The sorting data is divided into a plurality of segmented data by using a spline interpolation method;
[0088] The cumulative distribution function is fitted by a machine learning algorithm to learn the association between data elements in each of the segmented data and the sorting positions, to obtain a learning index with bounded error.
[0089] Specifically, after constructing the cumulative distribution function CDF, the sorting data is segmented by using a spline interpolation based on the mapping relationship of CDF, and a corresponding prediction model is constructed for each segmented data to obtain the learning index.
[0090] As an example, the prediction model is a linear prediction model, and the formula is:
[0091]
[0092] wherein, is the sorting position of the data element x, is the sorting position of the first data element in the segmented data, is the first data element of the segmented function, is the slope of the segmented data, is the target data element to be queried.
[0093] Suppose the sorting data is divided into two segmented data, and the corresponding prediction models are constructed as and The learning index contains the prediction models of the above two segmented data, and in the subsequent application process, the corresponding prediction model can be selected by the target data element to be queried to predict the corresponding sorting position.
[0094] Alternatively, the learning index is a learning index with bounded error. That is, the sorting position of the target data element predicted by the learning index is a position with error. When querying a data element, the database system uses the learning index to predict the corresponding position with error pos of the data element, and then performs binary search in the sorted range [pos-min_err, pos+max_err] to complete the search data.
[0095] By means of the technical means of the embodiment of the present application, the mapping relationship between the sorting data and the sorting position is described by constructing a cumulative distribution function, and a machine learning algorithm is used to fit the cumulative distribution function to learn the data distribution law of the sorting data, thereby constructing a learning index, which is beneficial to realize the fast searching of the sorting position of the data element, and further improve the efficiency of searching the original position of the data element by using the sorting position and the position vector, and realize the extended application of the machine learning algorithm on unordered data.
[0096] As a preferred embodiment, the embodiment of the present application is further implemented on the basis of the above-mentioned embodiment, and after the step S12, i.e. the step of sorting each data element in the original data to form sorting data, the method further comprises steps S16 and S17:
[0097] S16, calculating the hash value of the data element at each sorting position in the sorting data;
[0098] S17, storing the hash value of the data element at each sorting position according to the sorting position to form a position matching vector; wherein the position matching vector is used to associate the sorting position of the data element in the sorting data with the hash value of the data element; and the position matching vector is used for accuracy verification of the sorting position.
[0099] Then, the step S15 specifically comprises:
[0100] According to the learning index, the position vector and the position matching vector, a data index of the original data is generated.
[0101] In the embodiment of the present application, the basic structure of the data index comprises a position vector, a position matching vector and a learning index.
[0102] After each data element in the original data is sorted according to the preset sorting standard to form sorting data, the hash value of each data element is calculated, and the hash value of each data element is also sorted according to the sorting position to form a position matching vector.
[0103] Optionally, a simple decimal to hexadecimal conversion of the first two bits can be used as a hash rule to calculate the hash value of the data element, and MD5, SHA and the like can also be used as a hash rule, which is not limited here.
[0104] The position matching vector is used to solve the query delay problem caused by invalid disk reading. Since the learning index is a learning index with bounded error, the ranking position of the target data element to be queried calculated by the learning index may be erroneous, and the position matching vector can quickly filter out the invalid ranking position. Since the hash value is unique, the same data hash value is consistent, and different data hash values are different with high probability. The hash value in the position matching vector can directly associate the data element with the ranking position, providing a basis for subsequent quick filtering of invalid ranking positions, avoiding the need to directly read the original data to verify whether it matches during subsequent queries, thereby reducing invalid disk reading and query delay.
[0105] When the hash value is matched, the hash value of the value to be queried is calculated and compared with the hash value in the position matching vector corresponding to the predicted ranking position range. If the hash value of a certain position is inconsistent with the hash value of the value to be queried, the ranking data corresponding to the position is not the target data, and there is no need to read the position vector value corresponding to the position, i.e., there is no need to locate the original data by the position vector and read it, and the position can be directly skipped. Disk reading is a high-time-consuming operation in database query. By filtering out invalid positions in advance through hash matching, the number of disk readings can be greatly reduced, and only the original data corresponding to the position where hash matching is successful needs to be read, thereby significantly reducing query delay.
[0106] As a preferred embodiment, the method further comprises a step S18:
[0107] S18, discarding the ranking data.
[0108] In the embodiment of the application, the ranking data is no longer needed due to the presence of the position vector, and the ranking data will be discarded subsequently, thereby saving data storage space.
[0109] By using the technical means of the embodiment of the application, the hash value of the ranking data is calculated to construct the position matching vector. In the data query application process, consistency matching can be performed between the hash value of the target data element and the hash value associated with the predicted ranking position range, to realize screening of the target ranking position, filter out invalid ranking positions, reduce invalid disk reading, and find the target data element through only one valid disk reading, thereby greatly improving the query efficiency of non-ranking fields.
[0110] The embodiment of the application also provides a data query method, comprising the following steps:
[0111] A preset data index is used to locate the original position of the target data element to be found. The data index is constructed according to the construction method of the data index of any one of the above embodiments, and the data index comprises a position vector and a learning index.
[0112] Specifically, referring to Figure 3 is a flowchart of a data query method provided by an embodiment of the present application, which adopts a preset data index to locate the original position of a target data element to be searched, and includes steps S21 to S23.
[0113] S21, a learning index is adopted to determine a target ranking position of the target data element; wherein the learning index is used to predict the ranking position of a data element in ranking data and the ranking position of the data element in the ranking data;
[0114] S22, a target offset associated with the target ranking position is searched according to the position vector; wherein the position vector is used to associate the ranking position of a data element in ranking data and the offset of the data element, and the offset is used to identify the original position of the data element in original data;
[0115] S23, the original position of the target data element is located according to the target offset.
[0116] In the embodiment of the present application, a target data element to be searched is determined, and a data index corresponding to original data in which the target data element is located is determined. The data index is constructed according to the construction method of the data index in any one of the above embodiments, and the data index at least includes a position vector and a learning index.
[0117] The learning index is used to predict the ranking position of the target data element in the ranking data, and the target ranking position is determined. Then, the offset associated with the target ranking position is searched according to the position vector as the target offset, and the original position of the target data element in the original data is searched according to the target offset.
[0118] By using the technical means of the embodiment of the present application, the function of mapping the original position of non-ordered original data by the position vector is combined with the data position prediction ability of the learning index, and an efficient data index scheme is provided for unordered data, which solves the problem that the existing machine learning algorithm cannot be applied to unordered data query. Compared with the traditional two-level index method, the learning index is extended and applied to the unordered data query scene, which can accelerate the query process of non-ordered data, improve the data query performance, and reduce the space cost of data index.
[0119] Referring to Figure 4 is a flowchart of a data query method preferred in the embodiment of the present application, and in one embodiment, the learning index is a bounded-error learning index fitted according to the cumulative distribution function constructed according to the data element and the ranking position in the ranking data.
[0120] The step S21, i.e. determining the target sorting position of the target data element by using the learning index, comprises steps S211-S213:
[0121] S211, predicting a sorting position with an error corresponding to the target data element according to the learning index, denoted as a predicted sorting position;
[0122] S212, determining a sorting position range according to the predicted sorting position and a preset error value; wherein the sorting position range comprises a plurality of sorting positions;
[0123] S213, determining the target sorting position in the sorting position range.
[0124] Specifically, according to the construction process of the data index, when the learning index is a learning index with a bounded error, the predicted sorting position of the target data element predicted according to the learning index is a position with an error, the sorting data range containing the target sorting position is further determined according to the predicted sorting position, and the target sorting position is further determined in the sorting data range, thereby improving the prediction accuracy of the sorting position.
[0125] As an example, the predicted sorting position pos is predicted by the learning index, the error value err is preset, i.e. the range of the predicted sorting position ± err contains the real position, and the sorting data range [pos-err, pos+err] is obtained.
[0126] As a preferred embodiment, the data index further comprises a position matching vector, which is formed by storing the hash value of the data element at each sorting position in the sorting data according to the sorting position.
[0127] Then, the step S213, i.e. determining the target sorting position in the sorting position range, comprises:
[0128] According to the position matching vector, searching for the hash value associated with each sorting position in the sorting position range;
[0129] Calculating the hash value of the target data element;
[0130] Matching the hash value of the target data element with the hash value associated with each sorting position in the sorting position range to determine the target sorting position.
[0131] The embodiment of the present application quickly filters invalid sorting positions by using the position matching vector to determine the accurate target sorting position, thereby avoiding reading irrelevant original data.
[0132] Specifically, a hash value of the target data element is calculated, and a consistency match is performed on the hash values associated with each ranking position in the ranking position range in the position matching vector, so that a ranking position corresponding to a consistent hash value is obtained as the target ranking position.
[0133] By using the technical means of the embodiment of the application, the target ranking position is screened by performing a consistency match on the hash value of the target data element and the hash value associated with the predicted ranking position range, invalid ranking positions are filtered out, invalid disk reading is reduced, only one effective disk reading is required to find the target data element, and the query efficiency of the non-ranking field is greatly improved.
[0134] Referring to Figure 5 is an example diagram of data index construction and query data in a specific implementation scenario in the embodiment of the application, and the data index construction and query process of the embodiment of the application are explained and described by using a specific embodiment.
[0135] Data index construction process:
[0136] The non-ordered original data obtained is: 44, 2, 67, 89, 34, 5, 12, 55, 22, 7, 63, 98, 14, 27, and 88.
[0137] A unique offset is allocated to each data element in the original data, as shown in Table 2:
[0138] Table 2
[0139]
[0140] The data elements in the original data are sorted in ascending order to obtain sorted data, and correspondingly, the offsets are also sorted together with the data elements of the original data to form a position vector. A corresponding hash value is calculated for each data element to obtain a position matching vector. As shown in Table 3:
[0141] Table 3
[0142]
[0143] A CDF is constructed based on the sorted data to describe the mapping relationship between the data elements of the sorted data and the ranking positions in the sorted data. Based on the CDF mapping relationship, the sorted data is segmented by using spline interpolation, and a linear prediction model is trained for each segment, and the formula is: .
[0144] As an example, the sorted data is divided into two spline segments, which are segment data 1 and segment data 2, wherein the segment data 1 is the data corresponding to the ranking positions 0-6, and the segment data 2 is the data corresponding to the ranking positions 7-14.
[0145] Calculate the model parameters of each segment data:
[0146] Segment data 1: , = 2, the slope k is the ratio of the position change amount and the data change amount, that is, the slope , and the model formula is obtained as: .
[0147] Segment data 2: , = 34, the slope , and the model formula is obtained as: .
[0148] It can be understood that, in order to avoid the prediction position deviation, an error range is set, for example, the maximum error is set to 2, that is, the range of the predicted ranking position ± 2 contains the true position.
[0149] Thus, the learning index contains a prediction model of two segment data, and the corresponding model can be selected to predict the ranking position range through the target data element to be searched. Thus, the data index includes the position vector, the position matching vector and the learning index.
[0150] Data query process:
[0151] Suppose the target data element to be queried is 27, the corresponding predicted ranking position is predicted by using the learning index, and the prediction model of the above segment data 1 is used to calculate , that is, the predicted ranking position is 6. Since the learning index is a bounded error learning index, the predicted ranking position 6 obtained by prediction may not be the accurate target ranking position.
[0152] According to the predicted ranking position 6 and the preset error value 2, it is determined that the ranking position range is [4, 8], including 4, 5, 6, 7, 8, a total of 5 ranking positions, and the target ranking position needs to be searched in the range of ranking positions 4~8.
[0153] The hash value of the target data element is calculated as 14, and the hash values associated with the ranking positions 4~8 in the position matching vector are 5F, 6A, 14, E5 and 08 respectively. The hash value of the target data element is calculated as 14, and the matching is successful at the 6th position, that is, the target ranking position is 6. It can be understood that if the hash matching is unsuccessful, it means that there is no need to read the value of the position vector corresponding to this position, thereby reducing the number of times of reading the original data from the disk.
[0154] In the position vector, the position vector value corresponding to the target ranking position 6 is obtained, and it is queried that the value on the 6th position is 13, which is the target offset, representing that the target data element to be searched is located at the 13th position in the original data. According to the target offset, the original position of the target data element is quickly located in the original data, and other field data associated with the target data element can be quickly queried.
[0155] Referring to Figure 6 , which is a structural schematic diagram of a data index construction device provided by an embodiment of the present application. The embodiment of the present application provides a data index construction device 10, which comprises:
[0156] An offset allocation module 11 is configured to allocate an offset to each data element at an original position in original data; wherein the original data is unsorted data;
[0157] An ordered data generation module 12 is configured to sort each data element in the original data to form ordered data;
[0158] A position vector construction module 13 is configured to store the offset of each data element at an ordered position in the ordered data according to the ordered position to form a position vector; wherein the position vector is configured to associate the ordered position of the data element in the ordered data with the offset of the data element;
[0159] A learning index construction module 14 is configured to construct a learning index for the ordered data by using a machine learning algorithm; wherein the learning index is configured to associate the data element with the ordered position of the data element in the ordered data;
[0160] A data index generation module 15 is configured to obtain a data index of the original data according to the learning index and the position vector.
[0161] Preferably, the device 10 further comprises:
[0162] A position matching vector construction module is configured to calculate the hash value of each data element at an ordered position in the ordered data; and store the hash value of each data element at the ordered position according to the ordered position to form a position matching vector; wherein the position matching vector is configured to associate the ordered position of the data element in the ordered data with the hash value of the data element; and the position matching vector is configured to verify the accuracy of the ordered position.
[0163] Therefore, the data index generation module 15 is specifically configured to:
[0164] According to the learning index, the position vector and the position matching vector, a data index of the original data is generated.
[0165] It should be noted that the data index construction device provided by the embodiment of the present application is used to execute all process steps of the data index construction method of the above-mentioned embodiment, and the working principles and beneficial effects of the two are one-to-one correspondence, thus not being repeated.
[0166] Referring to Figure 7 is a structural schematic diagram of a data query device provided by the embodiment of the present application, and the embodiment of the present application further provides a data query device 20, which comprises:
[0167] A target data element index module is configured to locate a raw position of a target data element to be searched by using a preset data index, wherein the data index is constructed according to the data index construction method of any one of the above-mentioned embodiments, and the data index comprises a position vector and a learning index.
[0168] The target data element index module specifically comprises:
[0169] A target sorting position determination module 21 is configured to determine a target sorting position of the target data element by using the learning index, wherein the learning index is used to predict a sorting position of a data element in sorting data and the data element.
[0170] A target offset amount searching module 22 is configured to search a target offset amount associated with the target sorting position according to the position vector, wherein the position vector is used to associate a sorting position of a data element in sorting data and an offset amount of the data element, and the offset amount is used to identify a raw position of the data element in original data.
[0171] A raw position locating module 23 is configured to locate the raw position of the target data element according to the target offset amount.
[0172] It should be noted that the data query device provided by the embodiment of the present application is used to execute all process steps of the data query method of the above-mentioned embodiment, and the working principles and beneficial effects of the two are one-to-one correspondence, thus not being repeated.
[0173] The embodiment of the present application further provides a terminal device, which comprises a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the data index construction method of any one of the above-mentioned embodiments or the data query method of any one of the above-mentioned embodiments when executing the computer program.
[0174] The embodiment of the present application further provides a computer readable storage medium, which comprises a stored computer program, wherein the computer program controls a device where the computer readable storage medium is located to perform the data index construction method according to any one of the above embodiments or the data query method according to any one of the above embodiments when the computer program is running.
[0175] The embodiment of the present application further provides a computer program product, which comprises a computer program or computer instructions, and the computer program or the computer instructions are executed by a processor to realize the data index construction method according to any one of the above embodiments or the data query method according to any one of the above embodiments.
[0176] Those skilled in the art can understand that all or part of the processes of the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the program can be stored in a computer readable storage medium. When the program is executed, the program can include the processes of the above-mentioned embodiments. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or a random access memory (RAM) and the like.
[0177] The above is the preferred embodiment of the present application. It should be noted that those skilled in the art can make several improvements and refinements without departing from the principles of the present application, and these improvements and refinements are also considered to be within the protection scope of the present application.
Claims
1. A method of constructing a data index, characterized by, The method comprises: allocating an offset to a data element at each original position in original data; wherein the original data is unsorted data; sorting each data element in the original data to form sorted data; storing the offset of the data element at each sorted position in the sorted data according to the sorted position to form a position vector; wherein the position vector is used to associate the sorted position of the data element in the sorted data with the offset of the data element; using a machine learning algorithm to build a learning index for the sorted data; wherein the learning index is used to predict the data element and the sorted position of the data element in the sorted data; generating a data index of the original data according to the learning index and the position vector; after the step of sorting each data element in the original data to form sorted data, the method further comprises: calculating a hash value of the data element at each sorted position in the sorted data; storing the hash value of the data element at each sorted position according to the sorted position to form a position matching vector; wherein the position matching vector is used to associate the sorted position of the data element in the sorted data with the hash value of the data element; then the step of generating a data index of the original data according to the learning index and the position vector is specifically: generating a data index of the original data according to the learning index, the position vector and the position matching vector.
2. The method of claim 1, wherein, The step of using a machine learning algorithm to build a learning index for the sorted data comprises: building a cumulative distribution function according to the data elements and sorted positions in the sorted data; using a machine learning algorithm to fit the cumulative distribution function to learn the association between the data elements and the sorted positions in the sorted data and build a learning index.
3. The method of claim 2, wherein, The step of using a machine learning algorithm to fit the cumulative distribution function to learn the association between the data elements and the sorted positions in the sorted data and build a learning index comprises: using a spline interpolation method to divide the sorted data into a plurality of segmented data; using a machine learning algorithm to fit the cumulative distribution function to learn the association between the data elements and the sorted positions in each of the segmented data and build a learning index.
4. The method of claim 1, wherein, The method further comprises: discarding the sorted data.
5. A data query method, characterized by, The method comprises: using a preset data index to locate the original position of a target data element to be searched; wherein the data index is built according to the method for building a data index according to any one of claims 1 to 4, and the data index comprises a position vector and a learning index.
6. The data query method of claim 5, wherein, The step of using a preset data index to locate the original position of a target data element to be searched comprises: using the learning index to determine a target sorted position of the target data element; wherein the learning index is used to predict a data element and a sorted position of the data element in sorted data; According to the position vector, a target offset corresponding to the target ranking position is searched; wherein the position vector is used to associate a ranking position of a data element in ranking data and an offset of the data element, and the offset is used to identify an original position of the data element in original data; According to the target offset, an original position of the target data element is located.
7. The data query method of claim 6, wherein, The learning index is a learning index with bounded error obtained by fitting a cumulative distribution function according to data elements and ranking positions in the ranking data; The learning index is used to determine the target ranking position of the target data element, including: According to the learning index, a ranking position with an existing error corresponding to the target data element is predicted, which is recorded as a predicted ranking position; According to the predicted ranking position and a preset error value, a ranking position range is determined; wherein the ranking position range includes a plurality of ranking positions; The target ranking position is determined in the ranking position range.
8. The data query method of claim 7, wherein, The data index further includes a position matching vector, which is formed by storing hash values of data elements at each ranking position in the ranking data according to the ranking position; The target ranking position is determined in the ranking position range, including: According to the position matching vector, a hash value corresponding to each ranking position in the ranking position range is searched; The hash value of the target data element is calculated; The hash value of the target data element is matched with the hash value corresponding to each ranking position in the ranking position range to determine the target ranking position.
9. A data index construction apparatus, characterized in that, Including: An offset allocation module is configured to allocate an offset to a data element at each original position in original data; wherein the original data is non-ranking data; A ranking data generation module is configured to sort each data element in the original data to form ranking data; A position vector construction module is configured to store offsets of data elements at each ranking position in the ranking data according to the ranking position to form a position vector; wherein the position vector is used to associate a ranking position of the data element in the ranking data and the offset of the data element; A learning index construction module is configured to use a machine learning algorithm to construct a learning index for the ranking data; wherein the learning index is used to predict the data element and the ranking position of the data element in the ranking data; A data index generation module is configured to obtain a data index of the original data according to the learning index and the position vector; The device further includes: A position matching vector construction module is configured to calculate hash values of data elements at each ranking position in the ranking data; and store the hash values of the data elements at each ranking position according to the ranking position to form a position matching vector; wherein the position matching vector is used to associate a ranking position of the data element in the ranking data and the hash value of the data element; The data index generation module is specifically configured to: According to the learning index, the position vector and the position matching vector, a data index of the original data is generated.
10. A data query apparatus, characterized by comprising: The application comprises: The target data element index module is configured to locate the original position of the target data element to be searched by using a preset data index, wherein the data index is constructed according to the data index construction method in any one of claims 1 to 4, and the data index comprises a position vector and a learning index.
11. A terminal device, characterized by comprising: The computer program is stored in the memory and configured to be executed by the processor, and the processor implements the data index construction method in any one of claims 1 to 4 or the data query method in any one of claims 5 to 8 when the computer program is executed.
12. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a stored computer program, wherein the computer readable storage medium controls a device in which the computer readable storage medium is located to execute the data index construction method in any one of claims 1 to 4 or the data query method in any one of claims 5 to 8 when the computer program is executed.
13. A computer program product, characterised in that, The computer program product comprises a computer program or computer instructions, and the computer program or the computer instructions implement the data index construction method in any one of claims 1 to 4 or the data query method in any one of claims 5 to 8 when executed by a processor.
Citation Information
Patent Citations
Data processing method and device and device for data processing
CN112668015A
Seismic data query method and device, electronic equipment and medium
CN119025481A