Modeling method for solving distance from sample point to simplex topological space in high-dimensional space
By gradually reducing the dimensionality through orthogonal projection in high-dimensional space and utilizing recursive convex quadratic programming, the efficiency and accuracy issues of distance calculation from sample points to simplexes in high-dimensional space are solved, and fast and accurate distance calculation is achieved, which is suitable for processing large-scale data sets.
Patent Information
- Application Number
- CN202510806413.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-30
AI Technical Summary
In high-dimensional space, existing technologies find it difficult to efficiently calculate the distance from sample points to simplex topological space. The calculation complexity is high, the efficiency is low, and it is difficult to process large-scale data sets.
The sample points are projected into the subspace where the simplex is located by the method of stepwise orthogonal projection dimensionality reduction. The distance from the sample points to the simplex is solved by a recursive strict convex quadratic programming problem, which reduces the computational complexity and improves the computational efficiency.
It realizes the rapid and accurate calculation of the distance from the sample point to the simplex in high-dimensional space, improves the computational efficiency and accuracy, and is suitable for processing large-scale data sets, especially in the fields of machine learning and pattern recognition.
Smart Images

Figure CN120724167A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and specifically relates to a modeling method for solving the distance from a sample point to a simplex topological space in a high-dimensional space, which can be applied to technologies such as deep learning, topological space distance solution, and computer vision. Background Art
[0002] With the development of big data and machine learning technologies, the processing and analysis of high-dimensional data has become a significant research area. Efficiently calculating the distance from a sample point to a specific topological object (such as a simplex) in high-dimensional space is a fundamental problem in many applications. For example, in fields such as pattern recognition, computer vision, signal processing, and bioinformatics, calculating the distance from a sample point to a specific set helps understand data distribution characteristics and make classification decisions.
[0003] The simplex is the simplest convex set in multidimensional space and plays an important role in topology. Mathematically, a simplex can be viewed as a geometric solid consisting of a set of vertices. For example, a line segment is a one-dimensional simplex, a triangle is a two-dimensional simplex, and a tetrahedron is a three-dimensional simplex. The concept of the simplex also applies to higher-dimensional spaces and has broad applications in solving optimization problems and cluster analysis.
[0004] However, directly calculating the distance from a sample point to a topological space of simplexes in high-dimensional space is often a complex and computationally intensive task. Traditional distance calculation methods often fail to meet the speed and accuracy requirements of practical applications. Efficient algorithm design becomes particularly critical when processing large datasets.
[0005] Currently, various methods have been proposed and applied to different scenarios for calculating the distance between a sample point and a simplex in high-dimensional space. These methods can be roughly divided into two categories: methods based on analytic geometry and methods based on numerical optimization.
[0006] (1) Methods based on analytic geometry
[0007] This type of method mainly relies on geometric principles and solves the distance from a point to a simplex in an analytical way.
[0008] Figure 1 The principle of calculating the distance from a sample point to a simplex based on geometric methods is demonstrated. For example, for a two-dimensional simplex (i.e., a triangle), the distance from the point to each side of the triangle can be calculated and the minimum value is selected as the final result. For a point O outside the line AB, a perpendicular line is drawn through O to the line AB, with its foot at O', as shown in the following example: Figure 1 Connect OA and find the angle between OA and AB, and then find AO'. Adding AO' to point A gives the coordinates of O', which is Then OO' is the distance from point O to line segment AB.
[0009] However, this approach faces challenges in high-dimensional spaces because the complexity of the analytical expression rises dramatically with the increase of dimension, and even closed-form solutions are difficult to obtain.
[0010] (2) Methods based on numerical optimization
[0011] Given the limitations of analytic geometry methods in high-dimensional situations, another approach is to transform the problem into a numerical optimization problem. Specifically, one can define an objective function that represents the squared distance from a point to a simplex and then use gradient descent, quasi-Newton methods, or other numerical optimization algorithms to find a solution that minimizes this objective function. Although this approach is theoretically applicable to spaces of arbitrary dimensions, in practice, the numerical optimization process can be inefficient due to the risk of local optima or slow convergence.
[0012] In summary, existing methods for calculating the distance from sample points to simplexes in high-dimensional spaces exhibit limitations on multiple levels. These limitations primarily manifest in low computational efficiency, insufficient modeling robustness, and difficulty coping with large-scale datasets. Furthermore, computational complexity increases exponentially with increasing dimensionality. This poses challenges for existing modeling methods for calculating the distance from sample points to simplex topological spaces in high-dimensional spaces, including high computational effort, low parallelism, and high memory usage. Summary of the Invention
[0013] Faced with these challenges, there is an urgent need for a new modeling method that can not only ensure the necessary accuracy requirements, but also significantly improve computing performance and have good scalability to meet the needs of ever-changing application scenarios.
[0014] Therefore, the present invention proposes a new modeling method for solving the distance between a sample point and a simplex topological space in a high-dimensional space, which is used to efficiently calculate the minimum distance between a sample point and a high-dimensional simplex. Through this modeling method, the present invention can not only effectively reduce the computational complexity, but also improve the computational efficiency and stability, thereby providing a more advanced, efficient and reliable solution for solving the distance between a point and a simplex in a high-dimensional space. It aims to overcome the shortcomings of the existing technology and can quickly and accurately calculate the distance between a sample point and a simplex in a high-dimensional space, which is of great significance for promoting research and technological development in related fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 The principle of calculating the distance from a sample point to a simplex based on geometric methods is demonstrated.
[0016] Figure 2 Shows a flowchart of a method for solving the distance from a sample point in a high-dimensional space to a simplex according to the present invention.
[0017] Figure 3 Shows a hierarchical clustering dendrogram obtained according to the distance calculation method of the present invention.
[0018] Figure 4 Shows a hierarchical clustering dendrogram obtained according to the Euclidean distance calculation method. Detailed implementation mode
[0019] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0020] 1. Problem description
[0021] In a high-dimensional space a d-dimensional simplex (d < n) is the convex hull formed by d + 1 affinely independent points constituting it.
[0022]
[0023] The problem is to given a point in a high-dimensional space and a simplex in a high-dimensional space find their Euclidean distance
[0024]
[0025] 2. Problem analysis
[0026] The main difficulty of this problem lies in that the simplex σ d has a range, and we don't know which point in σ d has the minimum distance to the point p. Moreover, the situation in a high-dimensional space may not conform to the geometric intuition in the daily three-dimensional space.
[0027] The idea of the modeling method for solving the topological space distance from a sample point to a simplex in a high-dimensional space of the present invention is to gradually perform orthogonal projection onto a lower-dimensional space to achieve dimensionality reduction. First, project the point p orthogonally down to the subspace d where the simplex σ is located. In this way, the points to be considered can be represented by the basis vectors of the simplex σ d and all subsequent steps only need to be discussed in the subspace X. A coefficient vector can uniquely determine a point in X.
[0028] Then recursively perform orthogonal projection onto a lower-dimensional simplex, and the square of the distance from the current point to the space where the lower-dimensional simplex is located can be obtained each time. If the obtained coefficient vector satisfies belonging to the simplex σ dIf the condition is met, the square of this distance is returned; otherwise, it means that the distance obtained is not the distance from the point to the low-dimensional simplex, and further dimensionality reduction is needed. The square of this distance is added to the distance from the newly obtained orthogonal projection point to the lower-dimensional simplex. This distance is obtained in the same way, so it is recursive.
[0029] 3. Method for solving the distance from sample points to simplex topological space in high-dimensional space
[0030] Figure 2 The flowchart of the method for solving the distance from a sample point to a simplex in a high-dimensional space according to the present invention is shown.
[0031] Step 1 Initialization:
[0032] Take one of the vertices a0 of the simplex as The origin in the real space, n is the The number of dimensions. is a basis vector, a i is the i-th point in the sample space, denoted as: i=1,2,...,dd represents the dimension of the simplex, Representing the real number space of n times d, construct the matrix A, which is written as: A is a simplex σ d The matrix composed of all basis vectors of the space X0, A multiplied by a coefficient vector can represent the simplex σ d Any point in the space spanned by . We use is a vector, p is a point, recorded as: I understand. Both and p can represent points.
[0033] For a point q inside the simplex, calculate the matrix μi, then we have then Where D represents the condition that the coefficient vector of the point in the simplex should satisfy, is a point inside the simplex The corresponding coefficient vector, c i Represents the coefficient vector The i-th component of .
[0034] Suppose we need to solve to the simplex σ d The distance between the sample point p0 and q * is a simplex σ d The point closest to point p0. It's q * The corresponding vector, we use in subsequent calculations To represent the outliers in the simplex The nearest point, The coefficient vector of μi represents The i-th component of .
[0035] Step 2: Reduce the dimension to the space spanned by the original simplex and calculate the distance from the sample point to the space spanned by the original simplex:
[0036] Let X0 represent the original simplex σ d The spanned space calculates the distance from the sample point p0 to the space X0 spanned by the original simplex.
[0037] represents the a1 vector, A real number representing d dimensions, where d is the dimension of the simplex, and the orthogonal projection point p of the sample point p0 on the subspace X0 spanned by the simplex X0 The corresponding coefficient vector Expressed as in for The i-th component of
[0038] Need to solve To solve the orthogonal projection point p X0 The corresponding parameter vector, represents the perpendicular vector (orthogonal projection point) from p0 to the simplex space X0; A is the construction matrix. This is a least squares problem, and using the least squares method to solve it, we can get in is a d-dimensional real number space. Then we can calculate Represents sample points The square of the distance to the space X0 spanned by the original simplex.
[0039] In step 2, the orthogonal projection point p of the sample point p0 on the subspace X0 spanned by the simplex is used. X0 dist 2 Assign initial value In subsequent iterative calculations, dist can be gradually updated 2 . Among them, dist 2 Indicates the original sample point p0 to be calculated to the original simplex σ d The square of the distance. is the sample point The orthogonal projection point to the space X0 spanned by the original simplex.
[0040] Step 3: Determine whether the conditions are met:
[0041] If the current coefficient vector but Figure 2 The displayed calculation process is terminated and dist is returned 2 Otherwise, let p X0 is the current sample point (denoted as p), σ d For the current simplex, go to step 4.
[0042] Step 4: Recursive dimensionality reduction:
[0043] If the current simplex dimension is k, in order to solve the current sample point To the orthogonal projection point on the subspace spanned by the lower dimensional simplex, construct a convex quadratic programming problem with equality constraints, and then the current sample point can be obtained The distance to the space spanned by the k+1 k-1 dimensional sub-simplices contained in the current simplex (denoted as Here m is used to distinguish different k-1 dimensional simplexes and the spaces they span. For each space spanned by a k-1 dimensional simplex, an optimization problem is constructed and solved.
[0044] The constraints for each optimization problem are Here we use m to distinguish the subscripts of all possible choices, B m It is obtained by selecting a k-rank matrix from a matrix B whose number of rows is the simplex dimension d plus 1 and whose number of columns is the simplex dimension. The first simplex dimension row of B is the identity matrix of the simplex dimension, and the last row is a vector of all 1s. m It is a submatrix consisting of k components selected from a vector C of simplex dimension d plus 1 (the selection of Cm and Bm should be corresponding). C is the simplex dimension. The first components are all 0, and the last component is 1. The resulting simplex dimension + 1 equation corresponds to the inequality constraints in set D becoming equality constraints. i=1.2…d,where c i Some of them are 0, that is, c j =0,1≤j≤d. represents the coefficient vector The subscript m corresponds to the selected equality constraint, which determines the space spanned by the mth k-1-dimensional subsimplex contained in the current simplex.
[0045] In the following formula represents the representation of the orthogonal projection points in the subspace spanned by the lower dimensional simplex, where The coefficient vector representing the orthogonal projection point. Solving the following strict convex quadratic programming problem with equality constraints can yield the optimal coefficient vector in the subspace spanned by the lower-dimensional simplex. is a vector The transpose of A T is the transpose of matrix A, is the coefficient vector of the current point p.
[0046]
[0047] According to optimization theory, for strictly convex quadratic programming problems, the optimal solution is to satisfy the KKT condition (Karush-Kuhn-Tucker condition). Therefore, for each optimization problem, we only need to solve this set of linear equations.
[0048]
[0049] in is a placeholder vector, the unknown quantity is We solve the linear equations to get The d components of the previous simplex dimension are the coefficient vectors of the orthogonal projection points we expect to obtain in the subspace spanned by the current lower-dimensional simplex. Then we can calculate the square of the distance from the current point p to the space spanned by the mth k-1 dimensional simplex
[0050] Step 4 is executed recursively, see also step 5. Let the dimension of the current simplex of the input of step 4 be k, and the current point be p. The dimension of the current simplex when step 4 was executed last time (j-1 time) was k+1. When step 4 was executed for dimensionality reduction and projection, the projection point px_k was obtained; when entering the j-1th execution step, the current point p = px_k. Then, for the jth execution of step 4, represents the parameter vector of p; The parameter vector representing the orthogonal projection point of the current point p onto the space spanned by the k-1 dimensional simplex.
[0051] Step 5: Check the coefficient vector after dimensionality reduction
[0052] remember is the square of the distance from the point to the space spanned by the k-1 dimensional simplex, and dist 2 (k-1) is the square of the distance from the point to the k-1 dimensional simplex.
[0053] If we solve the above linear equations (*) we get The corresponding point is inside the k-1 dimensional simplex The current point The square of the distance to the space spanned by the mth k-1-dimensional subsimplex contained in the current simplex It is the current point The square of the distance to the mth k-1-dimensional subsimplex contained in the current simplex
[0054] The current point The minimum distance to all k-1 dimensional simplexes is If for each of the multiple k-1 dimensional simplexes in step 4, the solution obtained by solving the above equations (*) is All satisfied Then from the multiple Get Thus the recursion of the current layer can be ended, and the current point is obtained at the current layer of the recursion Distance to the k-dimensional simplex And recursively return to the previous layer. In the previous layer, the current point of the recursive current layer is the orthogonal projection point p of the current point in the previous layer relative to the k-dimensional simplex X , so the square of the distance between the current point in the previous layer and the k-dimensional simplex is in It is the orthogonal projection point p of the current point in the previous layer to the current point in the previous layer relative to the k-dimensional simplex X , that is, the square of the distance from the current point in the previous layer to the space spanned by the relative k-dimensional simplex.
[0055] according to You can recursively get point p X to the d-dimensional simplex σ d The square of the distance dist 2 (d). Where l represents the k-dimensional simplex, which is the l-th simplex among the multiple k-dimensional simplexes obtained by dimensionality reduction from the k+1-dimensional simplex. It is called the first estimation part, which represents the square of the distance from the current point of the upper recursion to the space spanned by the l-th k-dimensional simplex. The second recursive part represents the current point of the lower recursion (i.e., the orthogonal projection point of the current point of the upper recursion relative to the l-th k-dimensional simplex) to each k-1-dimensional sub-simplex contained in the l-th k-dimensional simplex (note that it is not the original simplex σ d The minimum value of the square of the distance between all k-1 dimensional sub-simplexes contained in it. Represents the square of the distance from the current point of the lower recursion to the mth k-1-dimensional subsimplex contained in the lth k-dimensional simplex.
[0056] Last return where ||p0-p X0 || 2 Equivalent to That is, the square of the distance between the original sample point p0 and the space spanned by the d-dimensional simplex. Represents the projection point pX0 To the point in the simplex closest to the original sample point The square of the distance, there is only one d-dimensional simplex, so m has only one value.
[0057] return Figure 2 Step 5, if we solve the above linear equations (*) in step 4, we get The corresponding point is not inside the k-1 dimensional simplex Then recursively execute step 4 until you get The second recursive part In the recursive execution of step 4, the mth k-1 dimensional simplex is used as the new current simplex, and the orthogonal projection point p of the current point p obtained in step 4 to the subspace spanned by the mth k-1 dimensional simplex is set as m,k-1 ( The corresponding point) is taken as the new current point p. For all k-2 dimensional sub-simplexes of the m-th k-1 dimensional simplex, we get all Then go to step 5 to determine whether there is If the situation occurs, it will decide whether to enter step 4 and continue recursion. S is used to distinguish the different values of m corresponding to the recursion.
[0058] Notice Corresponding represents a point in the space spanned by a zero-dimensional simplex, is the coefficient vector of the zero-dimensional simplex itself (i.e., a point), so there must be The recursion will definitely stop.
[0059] Finally return dist 2 ←dist 2 +dist 2 (d) is enough, where dist 2 (d) is point p X0 to the d-dimensional simplex σ d The square of the distance.
[0060] In relation to the mth k-1 dimensional simplex, When recursively executing step 4, the constraints should be Add a condition that was not previously present on the basis of Among them B s and C s Satisfy the conditions in step 4, and continue recursively. It is easy to see that we can add at most as many equality constraints as the simplex dimension d. These are the vertices of the simplex, and their coefficient vectors must be in the simplex and meet the condition of belonging to the set D, so the recursion will definitely stop.
[0061] 4. Performance Evaluation
[0062] In order to verify the effectiveness of the modeling method of the present invention for solving the distance from sample points to simplex topological space in high-dimensional space, we used a gastric cancer pathology section dataset to construct a simplex topological space, and used the method proposed in the present invention to model the distance from sample points to simplex topological space in high-dimensional space. Hierarchical clustering was used for performance evaluation.
[0063] In our evaluation, we construct sample points using pathological slide data, and regard each cluster obtained by clustering the sample points as a simplex in a high-dimensional space. We then use the modeling method proposed in this invention to solve the distance from sample points to simplex topological space in high-dimensional space to calculate the distance from the sample points to these simplices.
[0064] Figure 3 The hierarchical clustering dendrogram obtained by the distance calculation method of the present invention is shown.
[0065] The horizontal axis of the Hierarchical Clustering Dendrogram represents the sample points (Sample Index). 0 on the horizontal axis represents a sample point classified as negative, and 1 on the horizontal axis represents a sample point classified as positive. The vertical axis represents the distance value obtained using the proposed method for modeling the distance between sample points in high-dimensional space and simplex topological space. The colored lines in the figure represent the distance between the two closest sample points in different clusters. Figure 3 In the figure, the horizontal axis of each sample point in the area where the horizontal axis is 0 is 0, and there is no 1, while the horizontal axis of each sample point in the area where the horizontal axis is 1 is 1, and there is no 0. This shows that after clustering, sample points of different categories are not clustered into the same cluster, that is, there are no misclassified samples. Figure 3 The meaning of the expression is that by using the distance calculation method from the sample point to the simplex representing the cluster provided by the present invention, the result obtained by hierarchical clustering is correct, thereby proving that the modeling method for solving the distance from the sample point to the simplex topological space in the high-dimensional space of the present invention is effective and can correctly separate different samples with an accuracy rate of 100%.
[0066] Figure 4 The hierarchical clustering dendrogram obtained by the Euclidean distance calculation method is shown.
[0067] Figure 4 For comparison. Figure 4 , for Figure 3 The sample points constructed from the same pathological slide data are hierarchically clustered. During the clustering process, the Euclidean distance is used to calculate the distance from the sample point in the high-dimensional space to the simplex topological space.
[0068] Figure 4There are errors in the clustering results obtained. Two positive samples (horizontally 1) are mistakenly classified into the negative sample (horizontally 0) cluster, and one negative sample (horizontally 0) is mistakenly classified into the positive sample (horizontally 1) cluster.
[0069] Conclusion: Comparison Figure 3 and Figure 4 The results prove the effectiveness of the modeling method proposed in this invention for solving the distance from sample points to simplex topological space in high-dimensional space.
[0070] Beneficial effects brought by the technical solution of the present invention
[0071] The modeling method proposed in this paper for solving the distance between sample points and simplex topological space in high-dimensional space has the following significant beneficial effects:
[0072] 1. Improved Computational Efficiency: Compared to traditional methods, this invention utilizes an advanced modeling approach for calculating the distance between sample points and simplex topological space in high-dimensional space, reducing unnecessary computational steps and significantly improving the efficiency of calculating the distance between sample points and simplexes in high-dimensional space. This makes it possible to process large-scale datasets, particularly in applications such as machine learning and pattern recognition that require frequent similarity measurements. This aspect can be applied to the identification of gastric cancer pathology slides in medical imaging.
[0073] 2. Enhanced accuracy of modeling methods for solving distances between sample points and simplex topological spaces in high-dimensional space: By deeply analyzing the geometric structure of simplexes, this invention can more accurately measure the distance between sample points and simplexes, thereby improving the overall accuracy of modeling methods for solving distances between sample points and simplex topological spaces in high-dimensional space. This is particularly important for tasks that rely on accurate distance metrics, such as classification and clustering.
[0074] 3. Expanded application areas: Since the present invention can process high-dimensional data while maintaining high efficiency and accuracy, it is not only applicable to traditional computer science fields, but can also be applied to problems involving high-dimensional data analysis in multiple fields such as bioinformatics and financial analysis.
[0075] 4. Promote interdisciplinary research: The combination of topology and machine learning has brought new research directions to the field of data science. The technical solution of this invention is not only applicable to data mining and pattern recognition, but also can promote the cross-integration of fields such as topology, geometry, and machine learning.
Claims
1. A method for calculating the distance between a sample point and a cluster, where a sample point is represented by a weighted sum of basis vectors, and a cluster is a collection of multiple sample points of the same category. The sample points are sample points in a medical imaging dataset, a bioinformatics dataset, a financial dataset, an image dataset, a user feature dataset, a text dataset, or a gene sequence dataset. The method comprises: Step 1, initialization: Using the d-dimensional simplex σ d represents a cluster where the simplex σ d are the d+1 affine-independent points of the cluster {a0,a1,...,a d+1 }Convex hull Take a0 as The origin of the cluster is denoted as p0. use represents the vector of p0, A is a simplex σ d The matrix of all basis vector combinations of the space spanned, A multiplied by a coefficient vector can represent the simplex σ d Any point in the spanned space, is a basis vector, a i is the i-th sample point, there is where i = 1, 2, ..., d; The sample point p0 and the cluster σ d The distance is expressed as where q is the simplex σ d The point in represents the vector of point q Expressed as where μ i is the basis vector The corresponding coefficients, the coefficient vector corresponding to the point q in the simplex is μ i Represents the coefficient vector The i-th component of where D represents the i-th component of the simplex σ d The coefficient vector of the points inside should satisfy the conditions, is a point inside the simplex The corresponding coefficient vector, c i Represents the coefficient vector The i-th component of . Let q * is a simplex σ d The point closest to point p0 is Among them, dist 2 Representative sample points With simplex σ d The square of the distance between the clusters represented by X0 is the original simplex σ d The spanned affine space, p X0 is the sample point Orthogonal projection point on X0; Step 2: Calculate sample points to simplex σ d The distance of the spanned space X0: Let X0 represent the original simplex σ d Zhang Cheng's affine space, Get the orthogonal projection point of the sample point p0 on space X0 The corresponding coefficient vector Expressed as in for The i-th component of Solution To solve the orthogonal projection point p X0 The corresponding coefficient vector, where For sample points arrive vector of dist 2 Assign initial value where dist 2 Represents the sample point p0 to the simplex σ d The square of the distance; Step 3: like The obtained dist 2 It is the sample point p0 to the simplex σ d The distance is returned directly to dist 2 It can be ended; Otherwise, the simplex σ d As the current simplex, point p X0 As the current point p, and execute step 4; Step 4: Let the dimension of the current simplex be k. Next, focus on the k+1 k-1 dimensional simplices contained in the current simplex, where k is an integer greater than or equal to 1. Calculate the distance from the current point p to the subspace spanned by each k-1 dimensional simplex As the distance from the current point p to the k-1 dimensional simplex Estimate of , where the coefficient vector corresponding to the current point p is m is used to distinguish k+1 possible k-1 dimensional simplexes, which can be obtained and It is the coefficient vector of the orthogonal projection point of the current point p in the subspace spanned by the m-th k-1 dimensional simplex; Step 5: Among them, according to Continuously recursively obtain Finally, you can get dist 2 (d), as the initial sample point p0 to the original simplex σ d The square of the distance and return the result; the first term in the above formula is called is the first estimated score, which represents the square of the distance from the current point to the space spanned by the l-th k-dimensional simplex; the second term is the second recursive part, which represents the minimum value of the square of the distance between the current point and each k-1-dimensional subsimplex contained in the required l-th k-dimensional simplex; Case 1: For the case where the orthogonal projection point obtained in step 4 is inside the simplex, the coefficient vector corresponding to the orthogonal projection point is The income It is the sample point p to the mth k-1 dimensional simplex σ k-1 distance Case 2: For the case where the orthogonal projection point obtained in step 4 is outside the simplex, the coefficient vector corresponding to the orthogonal projection point is Then recursively execute step 4 until the second recursive part is obtained Thus we get Among them, the mth k-1 dimensional simplex is used as the new current simplex, and the orthogonal projection point p of the subspace spanned by the current point p in step 4 to the mth k-1 dimensional simplex is m,k-1 As the new current point p.
2. The method according to claim 1, wherein In case 2 of step 5, recursively execute step 4 to reduce the dimension of the m-th k-1 dimensional simplex to obtain multiple k-2 dimensional simplexes, and obtain all Then, go to step 5 to determine whether situation 2 still occurs, and then decide whether to go to step 4 to continue recursion, where s represents the sequence number of the k-2 dimensional simplex.
3. The method according to claim 2, wherein: In step 5, if the dimension of the current simplex is 0, Corresponding represents a point in the space spanned by a zero-dimensional simplex, is the coefficient vector of the zero-dimensional simplex itself, so there must be The recursion stops and dist is finally returned 2 ←dist 2 +dist 2 (d), where dist 2 (d) is point p X0 to the d-dimensional simplex σ d The square of the distance.
4. The method according to claim 3, wherein: In step 4, by solving the following strictly convex quadratic programming problem, a unique optimal solution can be obtained, that is, the orthogonal projection point p from the current point p to the space spanned by the m-th k-1-dimensional simplex m,k-1 The corresponding coefficient vector Then we can calculate the square of the distance from the current point p to the space spanned by the mth k-1 dimensional simplex in Represents the parameter vector of the current point p; The parameter vector representing the orthogonal projection point of the current point p onto the space spanned by the k-1 dimensional simplex; in is the coefficient vector The transpose of A T is the transpose of matrix A, is the coefficient vector of the current point p, B m It is a matrix of rank k selected from a matrix B whose number of rows is the simplex dimension d plus 1 and whose number of columns is the simplex dimension d. The first simplex dimension row of B is the identity matrix of the simplex dimension, and the last row is a vector of all 1s. C m It is a submatrix consisting of k components selected from a vector C with a simplex dimension d plus 1. C is the first simplex dimension whose components are all 0 and the last component is 1. Represents the number of constraints that the space spanned by the mth k-1-dimensional subsimplex of the current simplex needs to satisfy compared to the space X0.
5. The method according to claim 4, wherein According to convex optimization theory, we get the following The convex quadratic programming problem is solved by solving the following system of equations, and we get in, is a placeholder vector, It is the unknown quantity to be solved in the system of equations, and its first simplex dimension component is the coefficient vector of the orthogonal projection point of the current point p in the subspace spanned by the mth k-1 dimensional subsimplex of the current simplex 6. The method according to any one of claims 1 to 5, wherein: In step 2, the least squares method is used to solve Get sample points Orthogonal projection point p on the X0 subspace X0 The corresponding parameter vector 7. A method for performing hierarchical clustering on a data set, wherein the data set includes a plurality of sample points, the method comprising: S1: Set each sample point as a cluster; S2: Calculate the distance between all cluster pairs; S3: Find the two closest cluster pairs and merge them into a new cluster; Repeat steps 2-3 until the number of clusters reaches the threshold; in In step S2, the minimum value of the distance between any sample point in one cluster and the other cluster in the cluster pair is used as the distance between the cluster pair, and the distance between the sample point and the cluster is calculated according to the method of any one of claims 1-6.
8. An information processing device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.