Spatial data index construction method based on hierarchical clustering and Hilbert dimension reduction

By applying hierarchical clustering and Hilbert dimensionality reduction technology in agricultural multidimensional data, multidimensional data is converted into one-dimensional data and combined with multi-segment linear models for searching, the problem of slow retrieval of multidimensional data is solved and efficient data retrieval is achieved.

CN119917501APending Publication Date: 2025-05-02HARBIN AEROSPACE STAR DATA SYST TECH CO LTD

Patent Information

Application Number
CN202411982296.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The multi-condition complex retrieval speed of agricultural multi-dimensional data is slow, and the existing index structure has the problem of index space expansion in high-dimensional space, and the complex algorithm of learning indexes leads to excessive memory consumption, reducing index efficiency.

Method used

The spatial data index construction method based on hierarchical clustering and Hilbert dimensionality reduction is adopted. The multidimensional data is divided into spatially similar clusters through the hierarchical clustering algorithm, and the Hilbert spatial fill curve is established in the cluster to obtain one-dimensional data and Hilbert encoding, and searched with a multi-segment linear model.

Benefits of technology

It effectively reduces the memory usage of data storage, retains the original location of data, and significantly improves the retrieval speed of multi-dimensional data, which is more efficient than traditional indexes and learning indexes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917501A_ABST
    Figure CN119917501A_ABST
Patent Text Reader

Abstract

The invention discloses a spatial data index construction method based on hierarchical clustering and Hilbert dimensionality reduction, belongs to the technical field of spatial indexes, and aims at solving the problem that the retrieval speed of mass agricultural multi-dimensional data is low, overcoming the limitation that an existing database spatial data index is low in construction speed and high in memory occupation and achieving rapid data retrieval. The method comprises the following steps: collecting agricultural multi-dimensional data and carrying out attribute binding, constructing clusters by utilizing a hierarchical clustering algorithm and setting threshold truncation points to segment high-dimensional spatial data into clusters in similar spatial distribution, filling different clusters through a Hilbert curve to map disordered high-dimensional data into ordered one-dimensional data to generate cluster type one-dimensional data, data point codes are calculated by utilizing Hilbert codes and are sequenced, then a multi-segment linear model is trained through positions, codes and errors, and finally corresponding data points are determined and retrieved according to query conditions. The method has wide applicability and can be applied to various mass data with multiple data attributes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of spatial indexing, and in particular relates to a method for constructing a spatial data index based on hierarchical clustering and Hilbert dimension reduction. Background Art

[0002] With the continuous development of smart agriculture, the level of agricultural informatization is getting higher and higher. The demand for multi-condition and complex retrieval after the binding of agricultural multi-dimensional data with agricultural business is gradually increasing in the agricultural system, and the amount of agricultural data covering a large area can reach millions. When performing multi-condition global retrieval or classified retrieval of these data, the response speed has far exceeded the system response time that users can wait for, which greatly affects the user experience. Therefore, it is urgent to solve the problem of slow retrieval of agricultural multi-dimensional data.

[0003] At present, in order to optimize the speed of agricultural data retrieval, the method of segmenting data and sub-table query is often used. However, in the early data processing stage, the service access speed is slowed down due to the problem of data memory occupation. Traditional index structures such as R-tree and its variants perform well in two-dimensional space, but when the data dimension increases, the problem of index space expansion may occur. The present invention aims to combine agricultural data retrieval with database indexing technology.

[0004] The current learning index can improve data retrieval rate and reduce memory usage by combining machine learning algorithms, but the usage scenarios do not take into account the multi-dimensional data space distribution, and complex machine learning algorithms may lead to excessive memory consumption when calculating the index, reducing indexing efficiency. Summary of the invention

[0005] The present invention aims to solve the problem of slow retrieval speed of multi-dimensional agricultural data with complex conditions. The following solutions are provided:

[0006] A method for constructing a spatial data index based on hierarchical clustering and Hilbert dimension reduction, the method comprising:

[0007] S1. Collect agricultural multidimensional data with an organizational structure and complete the binding of attributes and business of the agricultural multidimensional data;

[0008] S2. constructing a hierarchical clustering algorithm according to the organizational structure of the agricultural multidimensional data, setting a threshold cutoff point to divide the agricultural multidimensional data into spatially similar clusters, wherein the number of spatially similar clusters depends on the threshold cutoff point;

[0009] S3, establish Hilbert space filling curves in different clusters to obtain cluster one-dimensional data and Hilbert coding of data points;

[0010] S4, constructing a multi-segment linear model according to the cluster one-dimensional data, the Hilbert coding of the data points and the error e training;

[0011] S5. According to the input search condition, combined with the multi-segment linear model, find data points that meet the search condition.

[0012] Furthermore, in S1, the agricultural multi-dimensional data includes: the company to which it belongs, the department to which it belongs, the right holder and the name of the plot of land.

[0013] Furthermore, S2 specifically includes:

[0014] S21, performing one-hot encoding processing on the agricultural multidimensional data, splicing the processed data together to obtain binary data with data spatial distribution information, and considering the number of elements of the agricultural multidimensional data as m clusters;

[0015] S22, the binary data is formed into m*n dimensional spatial data, the m*n dimensional spatial data is:

[0016] (a i ,b i ,...,n i ), (a j ,b j ,...,n j ),...,(a m ,b m ,...,n m )

[0017] Wherein, i=1, 2, ..., m; j=1, 2, ..., m; a, b, c...n are the 1st to nth bits of the binary system;

[0018] S23, processing the m*n dimensional spatial data by a hierarchical clustering algorithm, and calculating the spatial distance D between the m clusters ij , the spatial distance D ij pass:

[0019]

[0020] get;

[0021] Find the minimum spatial distance, merge the two clusters with the closest spatial distance into a new cluster G1 using the shortest clustering method, and repeat the merging process for the remaining m-1 clusters until the number of remaining clusters is reduced to 1, and stop clustering. The leaf nodes represent the data points, the internal nodes represent the merged clusters, and the length of the edges reflects the similarity or distance between clusters, thus forming a tree diagram;

[0022] S24, setting a threshold cutoff point on the dendrogram to obtain spatially similar clusters.

[0023] Furthermore, S3 specifically includes:

[0024] S31, establish Hilbert space filling curves in different clusters, establish a unique Hilbert coding mapping relationship between the n-dimensional space and the one-dimensional space, and each data point in the n-dimensional space has a corresponding H value to indicate the position of the data point on the curve;

[0025] The H value is determined by:

[0026]

[0027] H{m1,m2,...,m i}={H(m1),{H(m2),...,{H(m i )}

[0028] m∈A n , H(m)∈A

[0029] Get, where A n represents n-dimensional space, A represents one-dimensional space, m i is a data point in n-dimensional space, i is the index of the data point, H(m i ) is the data point m i The corresponding H value;

[0030] S32, converting the H value obtained in each stage into a binary code through an iterative algorithm, and finally concatenating and converting it into a decimal value, thereby obtaining the Hilbert code of the data point.

[0031] Furthermore, S4 specifically includes:

[0032] S41, inputting the Hilbert code of the data point into the cumulative distribution function to obtain the position of the data point, the position of the data point is obtained by:

[0033] y=G(x)*M

[0034] Obtain, where x is the Hilbert code of the data point, y is the location of the data point, M is the total number in the one-dimensional space, and G(x) is the cumulative distribution function;

[0035] S42, fitting the cumulative distribution function according to the piecewise linear regression model, starting from the starting data point f, establishing a linear model with a slope k between the starting data point f and the next data point, and continuously calculating toward the subsequent data points, if the ordinate of a certain data point plus the error e cannot satisfy the linear model condition, then taking the data point as the new starting point f1, repeatedly establishing the linear model;

[0036] S43, according to the input starting point (a, b), based on the i-th training sample, determine the maximum slope and minimum slope of each data point, until the minimum slope of the next data point is less than the minimum slope of the previous data point or the maximum slope of the next point is greater than the maximum slope of the previous point, the first linear model is established, and (x i+1 ,y i+1 ) is a new starting point position, and the steps are continued to generate the multi-segment linear model;

[0037] The slope maximum is given by:

[0038]

[0039] Get, where x i is the horizontal coordinate of the i-th training sample, y i is the ordinate of the i-th training sample, k h is the maximum slope;

[0040] The slope minimum is obtained by:

[0041]

[0042] Get, where k l is the minimum slope value.

[0043] Further, in S5, the search conditions include: data point search and data range search.

[0044] Furthermore, S5 specifically includes:

[0045] S51, when the search condition is the data point search, all attribute fields are input to search for a data point that meets the search condition; first, the data point that meets the search condition is one-hot encoded, and then the corresponding cluster is found through the hierarchical clustering algorithm, and the Hilbert space filling curve mapping is performed in the cluster to map the search condition to the ordered one-dimensional data, and the trained linear model is found according to the Hilbert encoding to obtain the data point that meets the search condition;

[0046] S52. When the search condition is the data range search, input part of the attribute field to retrieve multiple data points that meet the search condition; first, perform one-hot encoding on the input search condition, and fill all 0s and all 1s on the search condition that has not been input to form a hyper-rectangular surface, and map it through the Hilbert space filling curve to ensure that when the hyper-rectangle and the curve have an intersection, the data points involved in the intersection are the data points that meet the search condition, and then execute the steps in S51 on the data points that meet the search condition, and finally obtain all the data points that meet the search condition.

[0047] Beneficial effects:

[0048] The present invention avoids the server data pressure caused by segmentation processing of agricultural data of huge magnitude or the table retrieval method, adopts a hierarchical clustering algorithm to perform cluster analysis on agricultural multidimensional data, effectively conforms to the complex hierarchical structure of the data, introduces the Hilbert filling curve, and through the iterative algorithm of the curve, can keep the relatively close position of the adjacent data in the multidimensional space when mapping to the one-dimensional space, which not only reduces the memory of data storage, but also retains the original position of the data. Compared with the traditional fully connected neural network model, the multi-segment linear model adopted only needs to incrementally build the model for the data. When a certain data point cannot satisfy the previous model within the error range, a new linear model is built with this data point, so that the training cost is small and less time is needed to fit the data. Compared with the traditional spatial geographic data index R-tree and the current learning index, the present invention can reduce the memory usage and greatly improve the retrieval speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is a flowchart of a method for constructing a spatial data index based on hierarchical clustering and Hilbert dimension reduction;

[0050] Figure 2 It is a schematic diagram of the effect of hierarchical clustering algorithm;

[0051] Figure 3 This is a schematic diagram of the tree diagram effect;

[0052] Figure 4 It is a schematic diagram of the effect of multi-segment linear model;

[0053] Figure 5 This is a schematic diagram of the data range retrieval effect. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.

[0055] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected specific embodiments of the present invention. Based on the specific embodiments of the present invention, all other specific embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0056] In order to further understand the content, features and effects of the present invention, the following specific implementation methods are given as examples.

[0057] Example 1: Combination Figure 1 This embodiment describes a method for constructing a spatial data index based on hierarchical clustering and Hilbert dimension reduction, including the following steps:

[0058] A method for constructing a spatial data index based on hierarchical clustering and Hilbert dimension reduction, the method comprising:

[0059] S1. Collect agricultural multidimensional data with an organizational structure and complete the binding of attributes and business of the agricultural multidimensional data;

[0060] S2. constructing a hierarchical clustering algorithm according to the organizational structure of the agricultural multidimensional data, setting a threshold cutoff point to divide the agricultural multidimensional data into spatially similar clusters, wherein the number of spatially similar clusters depends on the threshold cutoff point;

[0061] S3, establish Hilbert space filling curves in different clusters to obtain cluster one-dimensional data and Hilbert coding of data points;

[0062] S4, constructing a multi-segment linear model according to the cluster one-dimensional data, the Hilbert coding of the data points and the error e training;

[0063] S5. According to the input search condition, combined with the multi-segment linear model, find data points that meet the search condition.

[0064] Specifically, compared with traditional data retrieval methods, the present invention solves the problem of complex organizational structure of agricultural data attributes. On the basis of retaining the data hierarchy, it combines dimensionality reduction technology to construct ordered one-dimensional data, and accurately predicts the data storage location through multi-segment linear model training results, which significantly reduces the memory size occupied by data indexes and greatly improves the data retrieval rate.

[0065] In S1, the agricultural multi-dimensional data includes: the company to which it belongs, the department to which it belongs, the right holder and the name of the plot of land.

[0066] S2 specifically includes:

[0067] S21, performing one-hot encoding processing on the agricultural multidimensional data, splicing the processed data together to obtain binary data with data spatial distribution information, and considering the number of elements of the agricultural multidimensional data as m clusters;

[0068] S22, the binary data is formed into m*n dimensional spatial data, the m*n dimensional spatial data is:

[0069] (a i ,b i ,...,n i ), (a j ,b j ,...,n j ),...,(a m ,b m ,...,n m )

[0070] Wherein, i=1, 2, ..., m; j=1, 2, ..., m; a, b, c...n are the 1st to nth bits of the binary system;

[0071] S23, processing the m*n dimensional spatial data by a hierarchical clustering algorithm, and calculating the spatial distance D between the m clusters ij , the spatial distance D ij pass:

[0072]

[0073] get;

[0074] Find the minimum spatial distance, merge the two clusters with the closest spatial distance into a new cluster G1 using the shortest clustering method, and repeat the merging process for the remaining m-1 clusters until the number of remaining clusters is reduced to 1, and stop clustering. The leaf nodes represent the data points, the internal nodes represent the merged clusters, and the length of the edges reflects the similarity or distance between clusters, thus forming a tree diagram;

[0075] S24, setting a threshold cutoff point on the dendrogram to obtain spatially similar clusters.

[0076] Specifically, in S21, the company, department, owner, and plot name are processed by one-hot encoding, and the processed data are spliced ​​together to obtain binary data with data spatial distribution information;

[0077] In S22, after one-hot encoding processing: Company A 001, Company B 010, Company C 100, Department D 01, Department E 10, Zhang San: 001, Li Si: 010, Wang Wu: 100, Plot 1: 01, Plot 2: 10; then the one-hot encoding corresponding to the data of Company A, Department E, Zhang San, and Plot 2 is {0011000110}.

[0078] like Figure 2 As shown, given a two-dimensional data set, according to step S23, find the minimum spatial distance of 9 data points, merge data points 4 and 5 into G1, replace the above two data points with G1 to find the minimum spatial distance of 8 data points, merge points 6 and 7 into G2, replace the above two data points with G2 to find the minimum spatial distance of 7 data points, merge points 8 and 9 into G3, replace the above two data points with G3 to find the minimum spatial distance of 6 data points, merge data points G2 and G3 into G4, and so on. The final tree diagram is as follows: Figure 3 As shown in the figure, the distance of the clustering process will form a trend of gradually decreasing and then gradually increasing, because similar data points will be merged first, and the similarity of the remaining data will become smaller and smaller, indicating that these data points do not belong to the same cluster. According to the change of the curve, the mutation value is found, and the threshold cutoff point is set to 0.85, and the 9 data are clustered into 4 spatially similar clusters.

[0079] S3 specifically includes:

[0080] S31, establish Hilbert space filling curves in different clusters, establish a unique Hilbert coding mapping relationship between the n-dimensional space and the one-dimensional space, and each data point in the n-dimensional space has a corresponding H value to indicate the position of the data point on the curve;

[0081] The H value is determined by:

[0082]

[0083] H{m1,m2,...,m i}={H(m1),{H(m2),...,{H(m i )}

[0084] m∈A n , H(m)∈A

[0085] Get, where A n represents n-dimensional space, A represents one-dimensional space, m i is a data point in n-dimensional space, i is the index of the data point, H(m i ) is the data point m i The corresponding H value;

[0086] S32, converting the H value obtained in each stage into a binary code through an iterative algorithm, and finally concatenating and converting it into a decimal value, thereby obtaining the Hilbert code of the data point.

[0087] Furthermore, S4 specifically includes:

[0088] S41, inputting the Hilbert code of the data point into the cumulative distribution function to obtain the position of the data point, the position of the data point is obtained by:

[0089] y=G(x)*M

[0090] Obtain, where x is the Hilbert code of the data point, y is the location of the data point, M is the total number in the one-dimensional space, and G(x) is the cumulative distribution function;

[0091] S42, fitting the cumulative distribution function according to the piecewise linear regression model, starting from the starting data point f, establishing a linear model with a slope k between the starting data point f and the next data point, and continuously calculating toward the subsequent data points, if the ordinate of a certain data point plus the error e cannot satisfy the linear model condition, then taking the data point as the new starting point f1, repeatedly establishing the linear model;

[0092] S43, according to the input starting point (a, b), based on the i-th training sample, determine the maximum slope and minimum slope of each data point, until the minimum slope of the next data point is less than the minimum slope of the previous data point or the maximum slope of the next point is greater than the maximum slope of the previous point, the first linear model is established, and (x i+1 ,y i+1 ) is a new starting point position, and the steps are continued to generate the multi-segment linear model;

[0093] The slope maximum is given by:

[0094]

[0095] Get, where x i is the horizontal coordinate of the i-th training sample, y i is the ordinate of the i-th training sample, k h is the maximum slope;

[0096] The slope minimum is obtained by:

[0097]

[0098] Get, where k l is the minimum slope value.

[0099] Specifically, the minimum slope value of the next data point is less than the minimum slope value of the previous data point, which is expressed as:

[0100]

[0101] The maximum value of the slope of the next point is greater than the maximum value of the slope of the previous point, which is expressed as:

[0102]

[0103] The Hilbert codes of the multidimensional data obtained according to the Hilbert space filling curve are arranged in order: 1, 3, 5, 7, 10, 16, 17, 21, 24, corresponding to the positions in the one-dimensional data: 0, 1, 2, 3, 4, 5, 6, 7, 8. Set the error e to 1, and follow the steps of S43, starting from (1, 0) and calculating until (16, 5) to calculate k l is 4 / 15, which is smaller than the k of the previous point (10, 4) l =1 / 3, which does not meet the requirements for linear model construction. Taking (16, 5) as a new starting point, we continue to repeat the above steps. The resulting multi-segment linear model is as follows: Figure 4 shown.

[0104] In S5, the search conditions include: data point search and data range search.

[0105] S5 specifically includes:

[0106] S51, when the search condition is the data point search, all attribute fields are input to search for a data point that meets the search condition; first, the data point that meets the search condition is one-hot encoded, and then the corresponding cluster is found through the hierarchical clustering algorithm, and the Hilbert space filling curve mapping is performed in the cluster to map the search condition to the ordered one-dimensional data, and the trained linear model is found according to the Hilbert encoding to obtain the data point that meets the search condition;

[0107] S52. When the search condition is the data range search, input part of the attribute field to retrieve multiple data points that meet the search condition; first, perform one-hot encoding on the input search condition, and fill all 0s and all 1s on the search condition that has not been input to form a hyper-rectangular surface, and map it through the Hilbert space filling curve to ensure that when the hyper-rectangle and the curve have an intersection, the data points involved in the intersection are the data points that meet the search condition, and then execute the steps in S51 on the data points that meet the search condition, and finally obtain all the data points that meet the search condition.

[0108] Specifically, Figure 5As shown, the points where the hyperrectangular surface of the data range retrieval intersects with the Hilbert space filling curve are 10, 11, 28, 29, 30, 31, 32, 33, 34, and 35. These points are the points that meet the retrieval conditions. The steps in S51 are executed for the data points involved to obtain all the data points that meet the retrieval conditions.

Claims

1. A method for constructing a spatial data index based on hierarchical clustering and Hilbert dimension reduction, characterized in that: The steps include: S1. Collect agricultural multidimensional data with an organizational structure and complete the binding of attributes and business of the agricultural multidimensional data; S2. constructing a hierarchical clustering algorithm according to the organizational structure of the agricultural multidimensional data, setting a threshold cutoff point to divide the agricultural multidimensional data into spatially similar clusters, wherein the number of spatially similar clusters depends on the threshold cutoff point; S3, establish Hilbert space filling curves in different clusters to obtain cluster one-dimensional data and Hilbert coding of data points; S4, constructing a multi-segment linear model according to the cluster one-dimensional data, the Hilbert coding of the data points and the error e training; S5. According to the input search condition, combined with the multi-segment linear model, find data points that meet the search condition.

2. The method for constructing a spatial data index based on hierarchical clustering and Hilbert dimension reduction according to claim 1, characterized in that: In S1, the agricultural multi-dimensional data includes: the company to which it belongs, the department to which it belongs, the right holder and the name of the plot of land.

3. The method for constructing a spatial data index based on hierarchical clustering and Hilbert dimension reduction according to claim 1, characterized in that: S2 specifically includes: S21, performing one-hot encoding processing on the agricultural multidimensional data, splicing the processed data together to obtain binary data with data spatial distribution information, and considering the number of elements of the agricultural multidimensional data as m clusters; S22, the binary data is formed into m*n dimensional spatial data, the m*n dimensional spatial data is: (a i ,b i ,...,n i ),(a j ,b j ,...,n j ),...,(a m ,b m ,...,n m ) Wherein, i=1, 2, ..., m; j=1, 2, ..., m; a, b, c...n are the 1st to nth bits of the binary system; S23, processing the m*n dimensional spatial data by a hierarchical clustering algorithm, and calculating the spatial distance D between the m clusters ij , the spatial distance D ij pass: get; Find the minimum spatial distance, merge the two clusters with the closest spatial distance into a new cluster G1 using the shortest clustering method, and repeat the merging process for the remaining m-1 clusters until the number of remaining clusters is reduced to 1, and stop clustering. The leaf nodes represent the data points, the internal nodes represent the merged clusters, and the length of the edges reflects the similarity or distance between clusters, thus forming a tree diagram; S24, setting a threshold cutoff point on the dendrogram to obtain spatially similar clusters.

4. The method for constructing a spatial data index based on hierarchical clustering and Hilbert dimension reduction according to claim 1, characterized in that: S3 specifically includes: S31, establish Hilbert space filling curves in different clusters, establish a unique Hilbert coding mapping relationship between the n-dimensional space and the one-dimensional space, and each data point in the n-dimensional space has a corresponding H value to indicate the position of the data point on the curve; The H value is determined by: H{m1,m2,...,m i }={H(m1),{H(m2),...,{H(m i )} m∈A n ,H(m)∈A Get, where A n represents n-dimensional space, A represents one-dimensional space, m i is a data point in n-dimensional space, i is the index of the data point, H(m i ) is the data point m i The corresponding H value; S32, converting the H value obtained in each stage into a binary code through an iterative algorithm, and finally concatenating and converting it into a decimal value, thereby obtaining the Hilbert code of the data point.

5. The method for constructing a spatial data index based on hierarchical clustering and Hilbert dimension reduction according to claim 1, characterized in that: S4 specifically includes: S41, inputting the Hi1bert code of the data point into the cumulative distribution function to obtain the position of the data point, the position of the data point is obtained by: y=G(x)*M Obtain, where x is the Hilbert code of the data point, y is the location of the data point, M is the total number in the one-dimensional space, and G(x) is the cumulative distribution function; S42, fitting the cumulative distribution function according to the piecewise linear regression model, starting from the starting data point f, establishing a linear model with a slope k between the starting data point f and the next data point, and continuously calculating toward the subsequent data points, if the ordinate of a certain data point plus the error e cannot satisfy the linear model condition, then taking the data point as the new starting point f1, repeatedly establishing the linear model; S43, according to the input starting point (a, b), based on the i-th training sample, determine the maximum slope and minimum slope of each data point, until the minimum slope of the next data point is less than the minimum slope of the previous data point or the maximum slope of the next point is greater than the maximum slope of the previous point, the first linear model is established, and (x i+1 ,y i+1 ) is a new starting point, and the steps are continued to generate the multi-segment linear model; The slope maximum is given by: Get, where x i is the horizontal coordinate of the i-th training sample, y i is the ordinate of the i-th training sample, k h is the maximum slope; The slope minimum is obtained by: Get, where k l is the minimum slope value.

6. The method for constructing a spatial data index based on hierarchical clustering and Hilbert dimension reduction according to claim 1, characterized in that: In S5, the search conditions include: data point search and data range search.

7. The method for constructing a spatial data index based on hierarchical clustering and Hilbert dimension reduction according to claim 6, characterized in that: S5 specifically includes: S51, when the search condition is the data point search, all attribute fields are input to search for a data point that meets the search condition; first, the data point that meets the search condition is one-hot encoded, and then the corresponding cluster is found through the hierarchical clustering algorithm, and the Hilbert space filling curve mapping is performed in the cluster to map the search condition to the ordered one-dimensional data, and the trained linear model is found according to the Hilbert encoding to obtain the data point that meets the search condition; S52. When the search condition is the data range search, input part of the attribute field to retrieve multiple data points that meet the search condition; first, perform one-hot encoding on the input search condition, and fill all 0s and all 1s on the search condition that has not been input to form a hyper-rectangular surface, and map it through the Hilbert space filling curve to ensure that when the hyper-rectangle and the curve have an intersection, the data points involved in the intersection are the data points that meet the search condition, and then execute the steps in S51 on the data points that meet the search condition, and finally obtain all the data points that meet the search condition.

Citation Information

Patent Citations

  • Parallel high-speed railway survey data retrieval method based on grid indexes

    CN110297952A

  • Multi-modal model optimization retrieval training method and storage medium

    CN118094216A

  • Updatable spatial learning index method and device based on partition and dimension reduction

    CN118820537A

Cited By

  • Multi-dimensional data clustering dimensionality reduction collaborative learning index construction method for agricultural machinery track

    CN121350038A