Clustering method and system based on complex entropy and complex information level construction method
Patent Information
- Application Number
- CN202510294545.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
Artificial intelligence lacks the ability to think in a detailed way when solving complex problems, and lacks standards to quantitatively describe the complexity of problems, resulting in low accuracy of answers.
A clustering method based on complex entropy is adopted to generate clustering results of tree-like structures by initializing the category center of mass, calculating similarity, merging categories and calculating complex entropy values, and independently constructing a hierarchical solution strategy for complex problems.
It improves the accuracy of answers to artificial intelligence when solving complex problems, enhances the basis and adaptability of problem splitting, and can flexibly deal with problems of different complexity levels.
Smart Images

Figure CN120217016A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a clustering method based on complex entropy, a system, and a complex information hierarchy construction method. Background Art
[0002] For the solution of complex problems, a hierarchical solution strategy can be adopted, that is, the complex problem is gradually decomposed into multiple sub-problems to form a tree-like solution path, and then solved one by one. When facing complex problems, artificial intelligence, by means of proposing larger-scale, stronger computing power, and more complex models, and by stacking model parameters and expanding the scale of training data, although breakthroughs have been made in performance, it relies on large computing power for calculation, resulting in the lack of refined thinking ability of artificial intelligence in solving complex problems.
[0003] To improve the ability of artificial intelligence to solve complex problems, an AI hierarchical solution mechanism is proposed, such as Chain of Thought (CoT) design, workflow design, multi-level image label calibration, etc. These mechanisms are based on manual guidance and split problems into refined problems, which helps to improve the ability of AI to solve complex problems.
[0004] However, AI lacks a standard for quantitatively describing the complexity of problems, lacks a basis for splitting complex problems, and cannot autonomously generate a complex problem hierarchical solution strategy similar to that of humans, resulting in low accuracy in answering complex problems. Summary of the Invention
[0005] The present application provides a clustering method based on complex entropy, a system, and a complex information hierarchy construction method to solve the problem of low accuracy in answering complex problems.
[0006] In a first aspect, the present application provides a clustering method based on complex entropy, including:
[0007] Initialize the features of the elements in the first category, and calculate the category centroid based on the initialized elements;
[0008] Calculate the similarity based on the first category centroid to obtain a first merged category;
[0009] If the number of categories in the first merged category is greater than or equal to a first preset value, calculate a first entropy value through complex entropy, and the first entropy value is obtained by calculating the similarity difference;
[0010] Generate a clustering result based on the first entropy value, and the clustering result includes at least a first clustering result and a second clustering result, and the clustering result is a tree structure.
[0011] In some feasible embodiments, initializing the features of the elements in the first category and calculating the category centroid based on the initialized elements includes:
[0012] Obtain the vector representation of the elements in the first category;
[0013] Initialize the features of the elements in the first category, where the initialization is the element embedding vector representation;
[0014] Obtain the number of elements in the first category and the sample feature vector;
[0015] Calculate the first category centroid based on the number of elements and the sample feature vector.
[0016] In some feasible embodiments, calculating the similarity based on the first category centroid to obtain the first merged category includes:
[0017] Obtain the second category and calculate the second category centroid, where the similarity between the second category centroid and the first category centroid is greater than the similarity threshold;
[0018] Merge the first category and the second category to obtain a pair of the first category;
[0019] Based on the pair of the first category, obtain the third category and calculate the third category centroid, where the similarity between the third category centroid and the pair of the first category is greater than the similarity threshold;
[0020] Merge the pair of the first category and the third category to obtain the first merged category.
[0021] In some feasible embodiments, the first preset value is three;
[0022] If the number of categories in the first merged category is greater than or equal to the first preset value, calculate the first entropy value through complex entropy, including:
[0023] If the number of categories in the first merged category is greater than or equal to three, obtain the fourth category and calculate the fourth category centroid;
[0024] Calculate the first entropy value based on the first category centroid, the second category centroid, the third category centroid, and the fourth category centroid through the following formula:
[0025] H = -∑ 1≤i<j≤n ∑ 1≤k<1≤n,(i,k)<(k,l) log(1 - |S ij - S kl |);
[0026] Where i is the first category, j is the second category, k is the third category, l is the fourth category, S ijis the similarity between the centroid of the first category and the centroid of the second category, S kl is the similarity between the centroid of the third category and the centroid of the fourth category; |S ij -S kl | is the absolute difference in the distance between the first category pair and the second category pair, where the second category pair is the third category and the fourth category.
[0027] In some feasible embodiments, generating the clustering result based on the first entropy value includes:
[0028] If the first entropy value is less than the entropy value threshold, mark the first merged category as the retained category to generate the first clustering result;
[0029] If the first entropy value is greater than or equal to the entropy value threshold, mark the first category pair as the first merged category to generate the first-round clustering result;
[0030] Calculate the centroid of the first-round clustering result, and calculate the similarity based on the centroid to generate the second clustering result.
[0031] In some feasible embodiments, calculating the similarity based on the centroid to generate the second clustering result includes:
[0032] Calculate the similarity based on the centroid to obtain the second merged category;
[0033] If the second merged category is obtained, calculate the second entropy value;
[0034] If the second merged category is not obtained, or if the second entropy value is greater than or equal to the entropy value threshold, generate the second clustering result.
[0035] In some feasible embodiments, after generating the clustering result based on the first entropy value, it further includes:
[0036] Traverse the clustering result to obtain the assigned labels, where the assigned labels include the first label and the second label, the first label is the label of the first clustering result, and the second label is the label of the second clustering result;
[0037] Match the initial labels based on the assigned labels and update the tracking information. The initial labels are the labels generated by the statement list and are the indexes of the statements, and the tracking information is used to record the initial labels of the first category;
[0038] Generate the hierarchical clustering result, where the hierarchical clustering result includes the clustering result and the tracking information.
[0039] In a second aspect, the present application provides a clustering system based on complex entropy, including:
[0040] A centroid calculation module, configured to initialize the features of elements in the first category and calculate the category centroid based on the initialized elements;
[0041] A merging module, configured to calculate the similarity based on the first category centroid to obtain a first merged category;
[0042] If the number of categories in the first merged category is greater than or equal to a first preset value, a complex entropy calculation module is configured to calculate a first entropy value through complex entropy, and the first entropy value is obtained through similarity difference calculation;
[0043] A clustering result generation module, configured to generate a clustering result based on the first entropy value, where the clustering result at least includes a first clustering result and a second clustering result, and the clustering result is a tree structure.
[0044] In a third aspect, the present application provides a complex information hierarchy construction method, including:
[0045] Obtain a first answer text, where the first answer text is an answer text generated by a language model;
[0046] Segment the first answer text to generate a list of sentences, where the list of sentences includes multiple sentences;
[0047] Based on the list of sentences, generate a clustering result through a hierarchical clustering method based on complex entropy.
[0048] In some feasible embodiments, segmenting the first answer text to generate a list of sentences includes:
[0049] Preprocess the first answer text to obtain a preprocessed first answer text, where the preprocessing includes removing irregular symbols in the first answer text through regularization and performing screening and extraction on specific symbols;
[0050] Segment the preprocessed first answer text to generate a list of sentences.
[0051] As can be seen from the above technical solutions, the present application provides a clustering method, a system, and a complex information hierarchy construction method based on complex entropy. The clustering method includes: initializing the features of the elements in the first category, calculating the category centroid based on the initialized elements, and then calculating the similarity based on the first category centroid to obtain the first merged category. If the number of categories in the first merged category is greater than or equal to the first preset value, calculate the first entropy value through complex entropy, where the first entropy value is obtained by calculating the similarity difference, and generate a clustering result based on the first entropy value. The clustering result includes at least the first clustering result and the second clustering result, and the clustering result is a tree structure. The method calculates and generates a tree-shaped clustering result according to complex entropy, and aggregates similar problems to solve the problem of low accuracy in answering complex questions. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0053] Figure 1 It is a flowchart of the clustering method based on complex entropy provided by the embodiment of the present application;
[0054] Figure 2 It is a schematic diagram of initializing element features provided by the embodiment of the present application;
[0055] Figure 3 It is a schematic diagram of similarity calculation provided by the embodiment of the present application;
[0056] Figure 4 It is a schematic diagram of the cluster merging process provided by the embodiment of the present application;
[0057] Figure 5 It is a flowchart of the first entropy value comparison provided by the embodiment of the present application;
[0058] Figure 6 It is a flowchart of the complex information hierarchy construction method provided by the embodiment of the present application;
[0059] Figure 7 It is a schematic diagram of the first clustering result provided by the embodiment of the present application;
[0060] Figure 8 It is a schematic diagram of the second clustering result provided by the embodiment of the present application;
[0061] Figure 9 It is a schematic diagram of the input sentence group provided by the embodiment of the present application;
[0062] Figure 10 provided by the embodiment of the present applicationFigure 8 Schematic diagram of the clustering result of the input sentence group;
[0063] Figure 11 The first schematic diagram of the clustering result after expansion provided by the embodiment of the present application;
[0064] Figure 12 The second schematic diagram of the clustering result after expansion provided by the embodiment of the present application. Detailed implementation manners
[0065] The embodiments will be described in detail below, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application detailed in the claims.
[0066] In the process of solving complex problems, by excluding some factors with relatively low relevance and focusing on solving a small range of problems, the definition of complexity should involve whether it is possible to exclude factors with relatively low relevance. If it is not possible to exclude them, the complexity is low. If some factors with low relevance can be excluded, it is considered that the complexity of the task can be further reduced.
[0067] Therefore, determining whether a problem is complex depends on whether it is possible to exclude factors with relatively low relevance in the task, thereby focusing on a smaller problem, and then combining the conclusion of the focused task with the excluded factors to further solve the problem. Whether the complexity of the problem can be reduced does not depend on the number of elements, but on whether there are problems that can be focused on among the elements and whether there are factors that can be excluded.
[0068] The essence of the hierarchical solution strategy adopted when dealing with complex tasks is to construct structured sub-problems. The construction process of each sub-problem needs to focus on the specific components related to the original problem in it, while shielding off factors with relatively weak relevance, thereby reducing the dimension of the input of the sub-problem and reducing the difficulty of the sub-problem.
[0069] The process of constructing sub-problems mainly has two characteristics. First, in the process of constructing sub-problems, attention is paid to the mutual relevance between factors, rather than the individual feature distributions of each element. Second, the next-level factors of the same sub-problem often have similar relevance, rather than simply classifying factors below a certain threshold into one category.
[0070] In the second point, relevance is a concept of relative scale, and its nature changes with different observation perspectives. Therefore, relevance should not be measured solely by a fixed threshold to avoid the limitations of generalization. Secondly, if there are significantly relatively weak correlations among a set of factors, they should be excluded because these factors have less impact on other factors and will reduce the necessity of attention. On the contrary, priority should be given to focusing on those factors with stronger correlations.
[0071] That is to say, when determining the number of factors for a task, the more inconsistent the correlations among the factors are, the more complex the problem decomposition process will become.
[0072] If we start constructing sub-problems from the factor with the strongest correlation, a tree-like group of problems can be obtained. To sum up, when determining the number of factors for a task, the more inconsistent the correlations among the factors are, the more complex the problem decomposition process will become.
[0073] Moreover, when AI answers complex questions, it lacks a standard to quantitatively describe the complexity of the questions, lacks a basis for decomposing complex questions, and cannot independently generate a hierarchical problem-solving strategy similar to that of humans, resulting in low accuracy in answering complex questions.
[0074] To solve the above problems, the present application provides a clustering method based on complex entropy, which is applied to answering complex questions. In this method, each factor is regarded as a separate cluster, that is, a category, and the categories with the highest similarity are gradually merged until the complex entropy threshold condition is met or no further merging is possible. By setting the threshold of complex entropy, the complexity of each sub-problem can be limited, which not only provides a basis for clustering decision-making but also has strong adaptability, enabling it to flexibly handle questions of different complexity levels.
[0075] As Figure 1 shown, the method includes:
[0076] S100: Initialize the features of the elements in the first category and calculate the category centroid based on the initialized elements.
[0077] The first category is the category set for the initial clustering. For example, the first category is a set of sentence vectors, and the sentence vectors are obtained after processing the sentences generated by a large language model in a natural language generation task. Each sentence vector represents the feature representation of a sentence.
[0078] To initialize the features of the elements in the first category, in some embodiments, first obtain the vector representation of the elements in the first category, and then obtain the vector representation of each sentence through an embedding method. Pack the embedding vectors of each sentence into dictionary features, where the key is the category label and the value is the corresponding vector. Embedding is the result of converting a sentence into a continuous vector space through a specific embedding method, enabling each element to be represented in a low-dimensional vector space to complete the initialization of the factor features.
[0079] As Figure 2 shown, exemplarily, multiple texts can be used to obtain feature vectors through an embedding operation and then calculate the centroid.
[0080] For the elements of multiple samples, the centroid can be obtained by calculating the mean of the vector representations after embedding. The centroid of the elements of multiple samples is calculated by the following formula:
[0081]
[0082] where, |C i | is the number of samples in the first category C i , and x is the sample feature vector.
[0083] The centroid represents the central position of the first category in the vector space, is used to calculate the similarity between categories, and further judge the similarity degree between categories, providing a basis for clustering and merging.
[0084] S200: Calculate the similarity based on the centroid of the first category to obtain the first merged category.
[0085] Similarity is used to measure the similarity degree between two categories (represented by the centroid), and can be calculated in various ways. For example, cosine similarity, Euclidean distance, etc. In this embodiment, taking cosine similarity as an example, the representation of distance is illustrated.
[0086] To calculate the similarity, first obtain the second category. The first category is C i , and the second category is C j . Calculate the similarity S ij based on the following formula. The greater the similarity, the more similar the two categories are:
[0087]
[0088] where, ‖C i ‖ is the feature vector of the first category, and ‖C j ‖ is the feature vector of the second category.
[0089] For merging categories, in some embodiments, a second category is obtained and the centroid of the second category is calculated, where the similarity between the centroid of the second category and the centroid of the first category is greater than the similarity threshold, that is to say, the similarity between the first category and the second category is the greatest among all current categories.
[0090] Then the first category and the second category are merged to obtain a first category pair. Based on the first category pair, a third category C k is obtained, and the centroid of the third category is calculated. The similarity between the centroid of the third category and the first category pair is greater than the similarity threshold. Similarly to the process of merging with the first category pair, that is to say, the similarity between the third category and the first category pair is the greatest among all current categories. The first category pair and the third category are merged to obtain a first merged category.
[0091] Among them, the centroid of the first merged category can be calculated by the weighted average of the centroid of the first category, the centroid of the second category, and the centroid of the third category.
[0092] As Figure 3 shown, exemplarily, Feature Vector 1 and Feature Vector 2 are measured, or Feature Vector 1 and Feature Vector 3 are measured, etc.
[0093] As Figure 4 shown, exemplarily, the clustering rounds include Round0, Round1, Round2, Round3, and RoundN. Among them, in Round0, each feature vector forms a separate cluster, that is, Cluster 1 to Cluster N. At this time, there are N independent clusters. In Round1, the clusters with high similarity are merged. Cluster 1, Cluster 2, and Cluster 3 are merged into Cluster 123, Cluster 4, Cluster 6, and Cluster 7 are merged into Cluster 467, and the clusters that do not participate in this merger, such as Cluster 5, remain unchanged. Round2 continues the operation of the previous round and further merges the clusters with high similarity. For example, Cluster 123 and Cluster 5 are merged into Cluster 1235. Round3 follows the same pattern and merges more clusters with high similarity. For example, Cluster 1235 and Cluster 467 are merged into Cluster 1234567. RoundN represents the subsequent rounds of clustering, and the operation of merging clusters with high similarity continues until the clustering termination condition is met.
[0094] S300: If the number of elements in the first merged category is greater than or equal to the first preset value, calculate the first entropy value through complex entropy.
[0095] In the field of artificial intelligence, all factors will be embedded into the vector space to obtain feature vectors, and the correlation between factors can be calculated. Among them, the content that factors can be embedded in can be regarded as factors, which can be a picture, an image category in image classification, or a word, a sentence in NLP. The distance between factors is used to represent the correlation between factors, and the distribution density of them in the feature space is reflected by comparing the feature distances of each pair of categories.
[0096] By defining S ij and S kl to represent the feature distances between class i and class j, and between class k and class l respectively. When the difference between these two distance values S ij and S kl is relatively large, it indicates that the relevant difference between these two classes is also large. Therefore, we calculate |S ij -S kl | to measure the correlation difference between different class pairs, and use the measure of relation matching to replace the expression of distance.
[0097] Based on the above analysis, through the metric of complex entropy, the distance differences between factor pairs are aggregated to quantify the overall complexity of the task. In the calculation of complex entropy, not only the differences of single class pairs are considered, but the distance differences between all class pairs are accumulated.
[0098] Let D = {C1, C2,..., C n} be a task set containing n factors. Define the distance S ij between any two factors as the distance metric value between class C i and C j , satisfying the following conditions:
[0099] 1. S ij ≥0, and S ij = S ji (symmetry);
[0100] 2. S ii = 0, (reflexivity).
[0101] Let |S ij -S kl | represent the absolute difference in distances between two pairs of classes. The complex entropy H is calculated by the following formula:
[0102] H = -∑ 1≤i<j≤n ∑ 1≤k<1≤n,(i,k)<(k,l) log(1 - |S ij -S kl |);
[0103] where the constraint (i, j) < (k, l) indicates the enumeration order to avoid duplicate calculations. The distance here is a generalized normalized distance, which can be the normalized Euclidean distance, the cosine similarity distance, or other distance metrics.
[0104] Complex entropy is a second-order analysis of the feature distribution. By quantifying the relationships between features, especially by considering the similarity or correlation between features, complex entropy evaluates the complexity of the problem. Compared with the first-order analysis that only focuses on the distribution characteristics of a single feature, such as the mean and variance, complex entropy emphasizes the dependence and interaction between features, considers the covariation relationship between features, and the definition of complex entropy accumulates the multi-level interactions between features, thereby capturing the structural complexity.
[0105] According to the definition of complex entropy, zero entropy means that when the correlation between all factors is highly consistent, there is no need to construct sub-problems for the problem. The feature difference S of each factor pair is small and close to zero, resulting in the complex entropy H value tending to 0. The consistency of the correlation between all factors indicates that it is difficult to clearly exclude certain elements.
[0106] The limit entropy means that the consistency of the correlation of factor pairs is weak, and it is necessary to construct sub-problems. If there are three factors i, j, and k in a task, if S ij = 0 (that is, the two features are exactly the same), S ik = 1 (that is, the two features i and k are in an orthogonal relationship), then the complex entropy H tends to infinity. In this case, the change of k will not have any impact on i and j, and i and j can be grouped into a sub-problem to be solved, and k can be regarded as an independent problem to be solved.
[0107] As the complex entropy increases, the necessity of constructing sub-problems also increases. When the complex entropy value gradually climbs, it means that the complexity of the problem has increased significantly. At this time, the problem can be split into smaller manageable parts. Once the sub-problems are constructed and incorporated into the original problem as a whole, after recalculating the complex entropy, it will be found that its value has decreased.
[0108] Complex entropy provides a quantitative basis for the process of constructing sub-problems for complex tasks. Through the quantitative analysis of the relationships between features, complex entropy can reveal the mutual correlation of each factor in the task and its importance in the overall problem. When the complex entropy value increases, it indicates that the complexity of the problem has increased, and the task needs to be split into smaller sub-problems for easy processing. At this time, the numerical change of complex entropy provides a quantifiable guidance, which can help determine which factors can be grouped into the same sub-problem and which factors should be processed separately. Through this quantitative mechanism, it allows AI to autonomously formulate a problem hierarchical decomposition strategy and optimize the solution. It not only enhances the systematicness of problem-solving, but also provides empirical support for complex analysis and decision-making, and can more accurately identify and construct appropriate sub-problems when facing complex tasks, thereby improving work efficiency and the reliability of the solution results.
[0109] Among them, the first preset value is three. When the number of merged categories reaches three or more, enough category pairs can be formed for comparing the similarity differences, so as to accurately calculate the complex entropy. After the merging operation is completed, in some embodiments, if the number of categories in the first merged category is greater than or equal to three, obtain the fourth category and calculate the centroid of the fourth category.
[0110] Based on the centroid of the first category, the centroid of the second category, the centroid of the third category, and the centroid of the fourth category, calculate the first entropy value through the following formula. The first entropy value is obtained by calculating the similarity difference;
[0111] H = -∑ 1≤i<j≤n ∑ 1≤k<1≤n,(i,k)<(k,l) log(1 - |S ij - S kl |);
[0112] Among them, i is the first category, j is the second category, k is the third category, l is the fourth category, S ij is the similarity between the centroid of the first category and the centroid of the second category, S kl is the similarity between the centroid of the third category and the centroid of the fourth category; |S ij - S kl | is the absolute difference of the distance between the first category pair and the second category pair. The second category pair is the third category and the fourth category.
[0113] S400: Generate a clustering result based on the first entropy value.
[0114] As Figure 5 shown, after obtaining the first entropy value, compare the first entropy value with the entropy value threshold through a preset entropy value threshold. In some embodiments, if the first entropy value is less than the entropy value threshold, mark the first merged category as the retained category to generate the first clustering result; if the first entropy value is greater than or equal to the entropy value threshold, mark the first category pair as the first merged category and generate the first-round clustering result; calculate the centroid of the first-round clustering categories and generate the second clustering result based on the centroid to calculate the similarity.
[0115] When the first merged category is marked as the retained category, it represents retaining this merging operation. When the first category pair is marked as the first merged category, that is to say, the differences between these three merged categories are relatively large and the similarities are relatively low. This merging is unreasonable and this merging operation is not retained. Restore the category formed by the merger of the first category and the second category, cancel the merger with the third category, and at the same time terminate the clustering loop process.
[0116] Among them, the first clustering result consists of one or more reserved categories. The first-round clustering result is an intermediate result obtained by canceling the merger with the third category when the first merged category is unreasonable, providing a basis for generating the second clustering result. The second clustering result is a clustering result generated by further performing the clustering process on the basis of the first-round clustering result. Similarly to the first clustering result, it can consist of one or more reserved categories.
[0117] It can reduce the probability of wrongly merging categories with large differences, making the similarity of elements within each classification in the clustering result meet the expectation. Then, based on the merged categories, restart the next-round clustering calculation using the updated centroids, and continue to find suitable categories for merging until the termination condition is met.
[0118] In some embodiments, calculate the similarity based on the centroid to obtain the second merged category. If the second merged category is obtained, calculate the second entropy value. If the second merged category is not obtained, or if the second entropy value is greater than or equal to the entropy threshold, generate the second clustering result.
[0119] The second merged category is a new category obtained by calculating the similarity based on the centroid after calculating the centroid of the first-round clustering result and merging the categories that meet the similarity condition. Then calculate the entropy value through the complex entropy formula, that is, the second entropy value, which is to re-execute the steps from S300 to S400 until the termination condition is met. Among them, the termination condition is that the second entropy value is greater than or equal to the entropy value, or there are only two categories left and they cannot be merged into the second merged category.
[0120] In the process of calculating the centroid based on the first-round clustering result and merging categories according to the centroid, if no category combination that meets the merging condition, that is, the similarity is greater than the preset threshold, can be found and the second merged category cannot be formed, this means that, from the perspective of similarity, there are no obviously mergeable categories in the current clustering state, indicating that the clustering structure of the data is relatively stable at the current stage, and continuing to merge may not optimize the clustering effect and may even damage the existing reasonable clustering result.
[0121] When the second merged category is obtained, if the second entropy value is greater than or equal to the entropy threshold, it indicates that the similarity of elements within the second merged category is low and this merger is unreasonable. This means that although the merge operation has been performed, the merged categories have large differences and do not achieve the expected clustering effect. That is to say, in the clustering process, when the complex entropy after merging cannot be less than the entropy threshold, it indicates that the current clustering structure has reached a relatively stable state, and continuing the merge operation may not further optimize the clustering effect and may even make the clustering result worse.
[0122] During the clustering process, as the merging operation is continuously executed, the number of categories gradually decreases. When only two categories remain, since at least three categories are required for merging, at this time, no new merging operation can be performed and the clustering process terminates.
[0123] Among them, the clustering result includes at least a first clustering result and a second clustering result, and the clustering result is a tree structure. Each node in the tree structure represents a category, showing the hierarchical clustering relationship of the data, which can reflect the similarity between different sentence vectors and the clustering process.
[0124] After generating the clustering result based on the first entropy value, in some embodiments, traverse the clustering result to obtain the assigned labels, then match the initial labels based on the assigned labels to update the tracking information; generate a hierarchical clustering result, and the hierarchical clustering result includes the clustering result and the tracking information.
[0125] Among them, the assigned labels include a first label and a second label. The first label is the label of the first clustering result, and the second label is the label of the second clustering result. The initial label is the label generated through the statement list and is the index of the statement, which is used to track the original statement during the clustering process.
[0126] The tracking information is used to record the initial labels of the first category. After each clustering operation is completed, according to the merged or divided categories, update the recorded initial label information therein. Use the assigned labels to match to find the corresponding clustering result, and then match the elements in the clustering result with the initial labels to determine the original statement labels included in the clustering result, and then update the tracking information, which can establish the connection between the clustering result and the original statement, provide data support for analyzing the clustering effect, and help evaluate the rationality of the clustering.
[0127] By establishing the connection between the clustering result and the original statement, it is possible to know which original statements are included in the clustering result, and then judge whether the clustering is reasonable and whether the statements with similar semantics or features are clustered together.
[0128] After integrating the clustering result and the updated tracking information, a hierarchical clustering result is formed. The hierarchical clustering result is used to display the hierarchical structure and the corresponding relationship between the clustering and the original data. It can be understood that for the clustering process of multiple iterations, the tracking information is updated and the hierarchical clustering result is generated after each round of clustering. The hierarchical clustering result can also be presented in a tree diagram or other graphical way to show the hierarchical structure of the clustering and the tracking information, making the clustering result more intuitive.
[0129] The clustering method based on complex entropy provided in this embodiment, where the complex entropy provides a quantitative basis for the process of constructing sub-problems for complex tasks. Through the quantitative analysis of the relationships between features, the complex entropy can reveal the mutual correlations of various factors in the task and their importance in the overall problem. When the complex entropy value increases, it indicates that the complexity of the problem increases, and the task needs to be split into smaller sub-problems for easier processing. At this time, the numerical change of the complex entropy provides quantifiable guidance to help determine which factors can be grouped into the same sub-problem and which factors should be processed separately.
[0130] Through this quantization mechanism, it allows the AI to autonomously formulate a problem hierarchical decomposition strategy and optimize the entire solution. This not only enhances the systematicness of problem-solving but also provides empirical support for complex analysis and decision-making, enabling more accurate identification and construction of appropriate sub-problems when facing complex tasks, thereby improving work efficiency and the reliability of the solution results.
[0131] Based on the above clustering method based on complex entropy, some embodiments of the present application also provide a clustering system based on complex entropy, including:
[0132] A centroid calculation module, configured to initialize the features of the elements in the first category and calculate the category centroid based on the initialized elements;
[0133] A merging module, configured to calculate the similarity based on the first category centroid to obtain a first merged category;
[0134] If the number of elements in the first merged category is greater than or equal to the first preset value, the complex entropy calculation module is configured to calculate a first entropy value through complex entropy, and the first entropy value is obtained by calculating the similarity difference;
[0135] A clustering result generation module, configured to generate a clustering result based on the first entropy value, and the clustering result includes at least a first clustering result and a second clustering result, and the clustering result is in a tree structure.
[0136] For the effect of the above system embodiment during operation, reference can be made to the effect of the above method embodiment, which will not be elaborated here.
[0137] Based on the above clustering method based on complex entropy, some embodiments of the present application also provide a complex information hierarchy construction method, as Figure 6 shown, including:
[0138] S610: Obtain a first answer text.
[0139] The first answer text is an answer text generated by a language model. By invoking the first large language model as an Agent (intelligent agent) for answering questions, for ease of description, the Agent is defined as Agent1.
[0140] Exemplarily, the large language model called by Agent1 kernel is the glm-4 model of Zhipu Qingyan. First, Agent1 is asked a scientific question. This scientific question has no restrictions on the field and difficulty. The scientific question is: Please explain the entire process of rocket launch, requiring knowledge in the following fields: Physical mechanics: Describe how the rocket overcomes the earth's gravity and rises, involving Newton's laws and conservation of momentum. Chemistry: Explain the combustion reaction of rocket fuel and how the generated energy propels the rocket. Astronomy: During the launch process, how to consider the influence of the earth's rotation and gravitational changes on orbit insertion.
[0141] Agent1 will answer according to the question raised and generate the answering text, that is, the first answering text. When the large language model answers questions, the generated text is not in a standardized format. For the convenience of subsequent processing, in some embodiments, the first answering text is preprocessed to obtain the preprocessed first answering text, and then the preprocessed first answering text is segmented to generate a list of statements. Among them, the preprocessing includes removing irregular symbols in the first answering text through regularization and performing screening and extraction on specific symbols, and then performing segmentation.
[0142] S620: Segment the first answering text to generate a list of statements.
[0143] After segmenting the first answering text, the generated list of statements includes multiple statements. The statements are independent statements. First, these independent statements are tagged, that is, a tag symbol is assigned to each independent statement.
[0144] S630: Based on the list of statements, generate a clustering result through a hierarchical clustering method based on complex entropy.
[0145] After calculating the feature vector of each sentence, execute step S100 of the clustering method based on complex entropy, and continue to execute steps S200 - S400.
[0146] Through the above process, the obtained clustering result is a tree structure. After each round of clustering ends, for example, after generating the first clustering result, as Figure 7 and Figure 8 shown, each "round" represents a round of clustering. Each bracket in the figure represents a category. These categories will enter the next round of hierarchical clustering process based on complex entropy until the termination condition is met.
[0147] To verify the clustering result, the following experiment is used to verify the effect:
[0148] As Figure 9As shown, the input sentence group 7 is about the application of the chain of thought in the query process, but these three sentences are irrelevant. Among them, the sentence labels of sentence group 7 are [0], [1], and [2]. The input sentence group 8 is about having breakfast in the morning, and the three sentences are relevant. The sentence labels of sentence group 8 are [3], [4], and [5].
[0149] As Figure 10 shown, at round 1, [3], [4], and [5] are merged into the same category, while [0], [1], and [2] are not merged.
[0150] Through the above experiments, it can be proved that the method provided by the present application can generate a tree structure diagram and can cluster strongly related sentences into one category, but it is not sufficient to demonstrate whether the obtained tree structure has robustness and anti-interference ability.
[0151] Through the following experiments, the aim is to verify the robustness and anti-interference ability of the provided method. Randomly select a certain sentence from the above tree structure, and generate a more detailed explanation through the Agent. This process can be understood as that the content of a certain passage may be in doubt or not clear enough. Therefore, a large language model is needed to further explain it.
[0152] To ensure that the answers of the LLM are more accurate, all historical information, that is, the previous sentences, will be used as input, which helps to provide richer context, but at the same time may destroy the original tree structure. Since the goal is to obtain detailed information about a certain sentence while maintaining the hierarchy of the tree structure unchanged. Therefore, the aim is that the newly generated sentences can be correctly clustered into the category in the original structure that is most relevant to them.
[0153] Specifically, the generated sentences will be highly similar to the selected target sentence, thus forming a new category containing multiple similar pieces of information. That is to say, it is required that the generated sentences can be smoothly incorporated into the category where the target sentence is located during the clustering process without affecting the integrity of the original structure. Therefore, the method provided by the present application can accurately cluster the generated sentences that are highly relevant to the target sentence into the same cluster without destroying the tree structure.
[0154] After passing through step S530, a tree structure is obtained, which reflects the hierarchical clustering results of all sentences to show the clustering situation at each level. In step S510, a question is asked to the large language model, expecting Agent1 to provide an answer to the question. However, inevitably, some of the answer sentences may not fully help to understand the question.
[0155] At this point, the parts that need more detailed explanation are selected from the generated sentences, and Agent 2 is called to make the model give more professional answers, and a more detailed answer can be obtained. At this point, three pieces of information have been obtained, namely all the sentences generated for the first time, the sentences that are selected to require further explanation, and the newly generated, more detailed explanation sentences.
[0156] Then, the newly generated detailed explanation statement of the third item is added to all the statements generated for the first time to form a new sequence, and the next round of label construction is carried out.
[0157] Repeating steps S510 to S530 can obtain a new hierarchical tree structure to verify that the proposed method can accurately cluster the generated sentences that are highly related to the target sentence into the same cluster without destroying the tree structure. Figure 11 as well as Figure 12 , are the results of two experiments.
[0158] Figure 11 The original labels are [0], [1], [2], [3], [4], [5], [6], [7], [8], and the newly added labels are [9],
[10] ,
[11] , where [9],
[10] ,
[11] are the results of expanding the label [4] sentence. That is to say, [9],
[10] ,
[11] are most likely related to label [4]. It can be seen from the expanded sentence clustering tracking information that in Round4, labels [4], [9],
[10] are merged into the same category. It is understandable that in the process of generating the large model, some greetings or formatting languages will be generated, and label
[11] is not merged into the same category with other labels. Therefore, label
[11] can be some greetings or formatting languages.
[0159] Figure 12 In round 1, the original labels are [0], [1], [2], [3], [4], [5], [6], [7], [8], [9],
[10] ,
[11] ,
[12] ,
[13] ,
[14] ,
[15] ,
[16] ,
[17] ,
[18] , and the newly added labels are
[19] ,
[20] ,
[21] . The selected sentence label is
[14] , that is, the sentence
[14] is expanded to get the labels
[19] ,
[20] ,
[21] . In round 1, the label
[14] is clustered with the highly related labels
[19] ,
[20] ,
[21] to form a new category [20, 21, 14, 19].
[0160] The above results clearly show that the method provided by the present application can accurately cluster the generated sentences that are highly relevant to the target sentence into the same cluster without destroying the tree structure.
[0161] For the similar parts between the embodiments provided in this application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of this application and do not constitute a limitation on the protection scope of this application. For those skilled in the art, any other embodiments extended based on the solution of this application without creative efforts belong to the protection scope of this application.
Claims
1. A clustering method based on complex entropy, characterized in that: include: Initializing features of elements in a first category, and calculating a category centroid based on the initialized elements; Calculating similarity based on the centroid of the first category to obtain a first merged category; If the number of categories of the first merged category is greater than or equal to a first preset value, calculating a first entropy value by complex entropy, wherein the first entropy value is calculated by similarity difference; A clustering result is generated based on the first entropy value, the clustering result includes at least a first clustering result and a second clustering result, and the clustering result is a tree structure.
2. The complex entropy-based clustering method according to claim 1, characterized in that: Initializing the features of the elements in the first category and calculating the category centroid based on the initialized elements includes: Get the vector representation of the elements in the first category; Initialize the features of the elements in the first category, wherein the initialization is an element embedding vector representation; Obtaining the number of elements in the first category and the sample feature vector; Based on the number of elements and the sample feature vector, a first class centroid is calculated.
3. The complex entropy-based clustering method according to claim 2, characterized in that: The calculating the similarity based on the first category centroid to obtain a first merged category includes: Obtain a second category and calculate the second category centroid, wherein the similarity between the second category centroid and the first category centroid is greater than a similarity threshold; Merging the first category and the second category to obtain a first category pair; Based on the first category pair, a third category is obtained, and a centroid of the third category is calculated, wherein a similarity between the centroid of the third category and the first category pair is greater than a similarity threshold; The first category pair and the third category are merged to obtain a first merged category.
4. The complex entropy-based clustering method according to claim 3, characterized in that: The first preset value is three; If the number of categories of the first merged category is greater than or equal to a first preset value, calculating a first entropy value by complex entropy includes: If the number of categories of the first merged category is greater than or equal to three, obtaining a fourth category and calculating the centroid of the fourth category; Based on the first category centroid, the second category centroid, the third category centroid, and the fourth category centroid, the first entropy value is calculated by the following formula: H=-∑ 1≤i<j≤n ∑ 1≤k<1≤n,(i,k)<(k,l) log(1-|S ij -S kl |); Among them, i is the first category, j is the second category, k is the third category, l is the fourth category, S ij is the similarity between the centroid of the first category and the centroid of the second category, S kl is the similarity between the centroid of the third category and the centroid of the fourth category; |S ij -S kl | is the absolute difference in distance between the first category pair and the second category pair, wherein the second category pair is the third category and the fourth category.
5. The complex entropy-based clustering method according to claim 4, characterized in that: The generating a clustering result based on the first entropy value includes: If the first entropy value is less than the entropy value threshold, marking the first merged category as a reserved category to generate a first clustering result; If the first entropy value is greater than or equal to the entropy value threshold, marking the first category pair as a first merged category to generate a first round of clustering results; The centroid of the first round of clustering results is calculated, and similarity is calculated based on the centroid to generate a second clustering result.
6. The complex entropy-based clustering method according to claim 5, characterized in that: The calculating the similarity based on the centroid to generate a second clustering result includes: Calculate similarity based on the centroid to obtain a second merged category; If the second merged category is obtained, calculating a second entropy value; If the second merged category is not obtained, or if the second entropy value is greater than or equal to the entropy value threshold, a second clustering result is generated.
7. The complex entropy-based clustering method according to claim 6, characterized in that: After generating the clustering result based on the first entropy value, the method further includes: Traversing the clustering results to obtain assigned labels, the assigned labels including a first label and a second label, the first label being a label of the first clustering result, and the second label being a label of the second clustering result; Based on the assigned label matching the initial label, updating the tracking information, the initial label is a label generated by the statement list and is an index of the statement, and the tracking information is used to record the initial label of the first category; A hierarchical clustering result is generated, wherein the hierarchical clustering result includes a clustering result and tracking information.
8. A clustering system based on complex entropy, characterized in that: include: A centroid calculation module, used for initializing the features of the elements in the first category and calculating the category centroid based on the initialized elements; A merging module, configured to calculate a similarity based on the first category centroid to obtain a first merged category; If the number of categories of the first merged category is greater than or equal to a first preset value, the complex entropy calculation module is used to calculate a first entropy value by complex entropy, where the first entropy value is obtained by calculating the similarity difference; A clustering result generating module is used to generate a clustering result based on the first entropy value, wherein the clustering result at least includes a first clustering result and a second clustering result, and the clustering result is a tree structure.
9. A method for constructing a complex information hierarchy, characterized in that: include: Obtaining a first answer text, wherein the first answer text is an answer text generated by a language model; Segmenting the first answer text to generate a statement list, wherein the statement list includes a plurality of statements; Based on the statement list, a clustering result is generated by a hierarchical clustering method based on complex entropy.
10. The complex information hierarchy construction method according to claim 9, characterized in that: The first answer text is segmented to generate a statement list, including: Preprocessing the first answer text to obtain a preprocessed first answer text, wherein the preprocessing includes removing irregular symbols in the first answer text by regularization and performing screening and extraction on specific symbols; The preprocessed first answer text is segmented to generate a statement list.
Citation Information
Cited By
Advanced mathematics problem solving model reasoning strengthening method based on hierarchical thinking chain
CN120654837A