An API dependency intelligent decoupling method, system, device and medium based on hyperplane multi-cluster space analysis
Patent Information
- Application Number
- CN202610756023.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-01
AI Technical Summary
[0007]因此,本发明解决的技术问题是:如何解决现有API依赖关系分析中因依赖单中心聚类与统计特征度量而导致的对多模态非线性依赖结构识别精度不足、依赖强度量化粗糙及解耦策略缺乏自适应能力的问题
[0018]本发明提供了一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现所述的一种基于超平面多簇空间分析的API依赖关系智能解耦方法的步骤。
Smart Images

Figure CN122673701A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of API dependency analysis technology, specifically to an intelligent decoupling method, system, device, and medium for API dependencies based on hyperplane multi-cluster space analysis. Background Technology
[0002] Existing API dependency analysis methods primarily rely on call logs or API description documents, determining dependencies through static topology analysis or rule matching. These methods typically assume that dependencies between interfaces are linear and stable. However, when faced with scenarios where API dependency structures exhibit high-dimensional non-linear characteristics across systems and multiple business domains, they struggle to capture implicit coupling relationships that evolve dynamically.
[0003] At the clustering analysis level, existing methods mostly use single-center clustering algorithms to group APIs. When faced with multimodal feature distributions or non-convex cluster structures, clustering distortion is easily generated, and it is impossible to accurately distinguish the potential dependency boundaries between different functional modules.
[0004] At the dependency measurement level, existing methods mostly rely on call frequency or parameter similarity for rough evaluation, without establishing explicit mapping of dependencies at the geometric space level, and lacking quantitative description of the distribution structure between clusters, resulting in insufficient accuracy in coupling degree calculation.
[0005] At the decoupling strategy level, existing methods mostly remain at the stage of static evaluation and manual decision-making, lacking the ability to adaptively adjust strategies in response to system operation. When the system structure changes, manual intervention is required. Summary of the Invention
[0006] In view of the above-mentioned problems, the present invention provides an API dependency intelligent decoupling method, system, device and medium based on hyperplane multi-cluster space analysis.
[0007] Therefore, the technical problem solved by this invention is: how to solve the problems of insufficient accuracy in identifying multimodal nonlinear dependency structures, coarse quantification of dependency strength, and lack of adaptive capability of decoupling strategies in existing API dependency analysis due to reliance on single-center clustering and statistical feature measurement.
[0008] To address the aforementioned technical problems, this invention provides the following technical solution: an intelligent decoupling method for API dependencies based on hyperplane multi-cluster space analysis, comprising, Feature extraction and fusion processing are performed on the associated data of the API to construct a unified feature matrix; Based on a unified feature matrix, several cluster centers are constructed using a clustering algorithm to form a set of cluster centers; Based on the relationship between each API sample and each cluster center, determine the membership degree of each API sample to each cluster center; For each pair of cluster centers in the cluster center set, construct a separating hyperplane, and combine the membership degree with the separating hyperplane to determine the dependency strength between each pair of cluster centers, and construct a dependency strength matrix. Decoupling strategies are generated based on dependency strength matrices.
[0009] As a preferred embodiment of the API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis described in this invention, the construction of the unified feature matrix includes: Structured encoding is performed on the associated data to obtain multi-dimensional feature representations for each API; Based on the degree of correlation between each feature dimension and API dependency in historical dependency samples, the dependency correlation weight corresponding to each feature dimension is determined. The multi-dimensional feature representations are weighted and fused based on dependency association weights to generate a unified feature matrix and the corresponding feature correlation matrix.
[0010] As a preferred embodiment of the API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis described in this invention, the step of constructing several cluster centers based on a unified feature matrix and a clustering algorithm to form a cluster center set includes: The number of target clusters is determined based on a unified feature matrix and clustering evaluation criteria. Perform multi-center clustering partitioning on API samples in the unified feature matrix, dividing the API samples into API dependency clusters corresponding to the number of target clusters; The feature vector of the API sample that achieves the optimal feature aggregation degree within each API dependency cluster is determined as the corresponding cluster center, and the cluster center set is composed of all cluster centers.
[0011] As a preferred embodiment of the intelligent decoupling method for API dependencies based on hyperplane multi-cluster space analysis described in this invention, wherein: determining the membership degree of each API sample to each cluster center according to the relationship between each API sample and each cluster center includes: Calculate the feature space distance between the feature vector of each API sample and each cluster center in the cluster center set; Based on the feature space distance, the membership degree of each API sample to each cluster center is obtained through probability normalization, and the sum of the membership degrees of each API sample to all cluster centers is one.
[0012] As a preferred embodiment of the API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis described in this invention, the following steps are included: constructing a separating hyperplane for each pair of cluster centers in the cluster center set, determining the dependency strength between each pair of cluster centers by combining membership degree and the separating hyperplane, and constructing a dependency strength matrix: For each pair of cluster centers in the cluster center set, a separating hyperplane is constructed in the feature space where the unified feature matrix is located to separate the corresponding two API-dependent clusters; Calculate the projection distance of each API sample to the separating hyperplane, perform a nonlinear decay transformation on the projection distance, and obtain the local dependency contribution of each API sample to the corresponding cluster pair; The local dependency contribution is aggregated based on membership degree to obtain the dependency strength between each pair of cluster centers. The dependency strength matrix is composed of the dependency strengths of all cluster pairs.
[0013] As a preferred embodiment of the intelligent decoupling method for API dependencies based on hyperplane multi-cluster space analysis described in this invention, the decoupling strategy based on dependency strength matrix generation includes: Based on the dependency strength matrix, cluster pairs whose dependency strength satisfies the coupling condition are identified as target coupled cluster pairs; Construct a hierarchical decoupling graph based on the dependencies between target coupled cluster pairs; Generate a set of candidate decoupling strategies based on a hierarchical decoupling graph; By combining decoupling constraints, the target decoupling strategy is determined from the set of candidate decoupling strategies through strategy evaluation.
[0014] As a preferred embodiment of the intelligent decoupling method for API dependencies based on hyperplane multi-cluster space analysis described in this invention, the determination of the target decoupling strategy includes: Obtain the newly added API call data and calculate the distribution offset between the feature distribution corresponding to the newly added API call data and the feature distribution corresponding to the unified feature matrix. In response to the distribution offset exceeding the drift threshold, the cluster center set and the separating hyperplane are updated based on the newly added API call data. The dependency strength matrix is reconstructed based on the updated cluster center set and separating hyperplane, and the target decoupling strategy is regenerated based on the reconstructed dependency strength matrix.
[0015] This invention provides an intelligent decoupling system for API dependencies based on hyperplane multi-cluster space analysis.
[0016] To solve the above technical problems, the present invention provides the following technical solution: an intelligent decoupling system for API dependencies based on hyperplane multi-cluster space analysis, comprising: a feature matrix construction module, a set construction module, a calculation module, a dependency matrix construction module, and a generation strategy module; The feature matrix construction module is used to extract and fuse features from the associated data of the API to construct a unified feature matrix. The set construction module is used to construct several cluster centers based on a unified feature matrix and a clustering algorithm, forming a cluster center set. The calculation module is used to determine the membership degree of each API sample to each cluster center based on the relationship between each API sample and each cluster center; The dependency matrix construction module is used to construct a separating hyperplane for each pair of cluster centers in the cluster center set, and combine the membership degree with the separating hyperplane to determine the dependency strength between each pair of cluster centers, and construct a dependency strength matrix. The generation strategy module is used to generate a decoupling strategy based on the dependency strength matrix.
[0017] The present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the intelligent decoupling method for API dependencies based on hyperplane multi-cluster space analysis.
[0018] The present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the aforementioned intelligent decoupling method for API dependencies based on hyperplane multi-cluster space analysis.
[0019] The beneficial effects of this invention are as follows: This invention constructs a unified feature matrix by extracting and fusing features from associated API data, mapping API data from different sources and dimensions to the same feature space, thus eliminating the inconsistency in feature representation caused by differences in format and scale among multi-source data. By constructing multiple cluster centers through a clustering algorithm to form a cluster center set, compared to single-center clustering which can only identify single-modality clustering structures, this invention can identify multiple locally dependent clusters in the API feature space, providing more accurate grouping of APIs under multimodal and non-convex distribution structures, and reducing misjudgments of dependency boundaries caused by clustering distortion.
[0020] By constructing a separating hyperplane for each pair of cluster centers and combining membership to determine dependency strength, the dependency relationship between APIs is transformed from a statistical indicator to a distance relationship in geometric space for quantification. This gives the measurement of dependency strength a clear geometric meaning. Compared with a rough evaluation based on call frequency or parameter similarity, it can make a more refined distinction of the strength of nonlinear coupling structures in high-dimensional space.
[0021] This invention generates a decoupling strategy based on a dependency strength matrix, enabling decoupling decisions to be directly driven by quantified dependency strength data, reducing reliance on human experience and improving the matching degree between the decoupling strategy and the actual dependency structure. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 The above is a flowchart of an intelligent decoupling method for API dependencies based on hyperplane multi-cluster space analysis, provided as an embodiment of the present invention.
[0024] Figure 2 The flowchart illustrates the determination of target coupled cluster pairs in an API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis, as provided in one embodiment of the present invention.
[0025] Figure 3 The flowchart illustrates the target decoupling strategy of an API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis, as provided in one embodiment of the present invention. Detailed Implementation
[0026] To make the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0027] Example 1, referring to Figure 1 This is one embodiment of the present invention, which provides an intelligent decoupling method for API dependencies based on hyperplane multi-cluster space analysis, including: In a microservice architecture, complex dependencies exist between API interfaces within the system, often exhibiting multimodal and nonlinear distribution characteristics. Existing dependency analysis methods typically rely on statistical features and single-center clustering for dependency determination, which struggles to accurately characterize the complex coupling structures in high-dimensional space. This embodiment constructs a unified feature space, performs multi-cluster centroid clustering within this space to identify the local dependency cluster structure of APIs, then constructs a separating hyperplane for each cluster pair, quantifying the dependencies into distance relationships in geometric space. Finally, the generation of decoupling strategies is driven by the quantified dependency strength.
[0028] S1. Extract and fuse features from the associated data of the API to construct a unified feature matrix; S2. Based on a unified feature matrix, several cluster centers are constructed using a clustering algorithm to form a set of cluster centers; S3. Determine the membership degree of each API sample to each cluster center based on the relationship between each API sample and each cluster center; S4. For each pair of cluster centers in the cluster center set, construct a separating hyperplane. Combine the membership degree with the separating hyperplane to determine the dependency strength between each pair of cluster centers and construct the dependency strength matrix. S5. Generate decoupling strategy based on dependency strength matrix.
[0029] Example 2, an embodiment of the present invention, provides an intelligent decoupling method for API dependencies based on hyperplane multi-cluster space analysis, based on the previous embodiment, including: In step S1, feature extraction and fusion processing are performed on the associated data of the API to construct a unified feature matrix, including the following steps S11~S13: It should be noted that the current stage is the foundational layer of the entire method. The core objective is to standardize, extract, and fuse various types of input data to construct a unified high-dimensional feature space. This provides a consistent analytical basis for multi-cluster clustering and hyperplane metric, ensuring that the analysis in all stages is based on a unified feature standard and avoiding analytical biases caused by feature heterogeneity.
[0030] The API-related data involved in this phase includes: API interface call log setℒ, which originates from the system's log collection module and directly reflects the API's operational status. The set is in the form of... Where n represents the total number of API interfaces, and each log Includes key information about the API call process, including call frequency, call latency, request parameters, and response status; service dependency topology. It originates from the system's service registry and topology parsing module, and exists in the form of a graph structure. Nodes represent microservice modules in the system, and edges represent the calling relationships between modules.
[0031] Furthermore, this topology clarifies the service module to which each API belongs and the relationship between that service and other services, providing a basis for extracting the structural dependency features of the API; API semantic and structural description dataset. The API documentation and interface registration information, including the functional description, interface type, parameter constraints, and return value format of each API, are used to extract semantic features of the APIs and assist in identifying functional dependencies between APIs; historical dependency sample library. It is derived from the system's historical operation records and manually labeled data, including confirmed dependencies, dependency types, and dependency failure cases between APIs over a period of time, and is used to assist in feature weight calculation and decoupling strategy optimization.
[0032] S11. Perform structured encoding on the associated data to obtain multi-dimensional feature representations corresponding to each API.
[0033] The input call logs, service dependency topology, and API semantic and structural description data are structured and encoded to define feature extraction functions. Extract the corresponding feature representation for each API: , Among them, E i Let l be the multi-dimensional feature representation obtained after the i-th API is structured and encoded. i For the call log of the i-th API, This is a dataset describing the semantics and structure of APIs. This is a service dependency topology. The feature extraction function extracts behavioral features (including call frequency and statistical distribution of call latency) from the call logs, structural features (including the topological location of the module to which the API belongs and the connectivity with adjacent modules) from the service dependency topology, and semantic features (including semantic vector representation of interface functions and similarity encoding of parameter structures) from the semantic and structural description data. These three types of features together constitute the multi-dimensional feature representation of each API.
[0034] S12. Based on the degree of correlation between each feature dimension and API dependency in historical dependency samples, determine the dependency association weight corresponding to each feature dimension.
[0035] The dependency association weight βk is not manually set, but automatically calculated by the feature importance assessment module. This module analyzes a historical dependency sample library. The correlation between different feature dimensions and API dependencies is analyzed, with higher correlation resulting in higher weights. Combined with the information gain algorithm, the dependency weight βk of each feature dimension is automatically output to ensure the rationality of weight allocation and avoid subjective bias caused by manual setting.
[0036] Specifically, the information gain algorithm measures the importance of a feature dimension by calculating the reduction in information entropy of each feature dimension in relation to the API dependency determination result: the greater the reduction in information entropy, the stronger the feature dimension's ability to distinguish dependencies, and the larger the corresponding dependency association weight βk. Besides the information gain algorithm, dependency association weights can also be determined using correlation analysis based on the Pearson correlation coefficient. This method measures feature importance by calculating the linear correlation between the numerical sequences of each feature dimension and the dependency label sequences; the larger the absolute value of the correlation coefficient, the greater the corresponding dependency association weight.
[0037] S13. Based on the dependency association weight, the multi-dimensional feature representations are weighted and fused to generate a unified feature matrix and the corresponding feature correlation matrix.
[0038] The various feature dimensions are standardized and weighted to construct a unified feature matrix F: , in, This represents the dependency association weight corresponding to the k-th feature dimension, and it is not manually set but automatically calculated by the feature importance evaluation module built into this invention. This module analyzes the historical dependency sample library. In the model, the correlation between different feature dimensions and API dependencies (the higher the correlation, the greater the weight) is considered. Combined with the information gain algorithm, the weight βk of each feature dimension is automatically output to ensure the rationality of weight allocation and avoid subjective bias caused by manual setting. Ek is the feature encoding matrix corresponding to the k-th feature dimension, and K is the total number of feature dimensions.
[0039] After the above three steps, this stage outputs two core results: High-dimensional unified feature space Where n represents the total number of API interfaces and d represents the total number of feature dimensions. Each row of this feature matrix corresponds to a feature vector of an API, and each column corresponds to a feature dimension, realizing a standardized and unified representation of all API features.
[0040] Feature correlation matrix , where d represents the total number of feature dimensions. The elements of this matrix... The correlation coefficient represents the relationship between the k-th and m-th feature dimensions, ranging from -1 to 1. A coefficient closer to 1 or -1 indicates a stronger correlation between the two feature dimensions, while a coefficient closer to 0 indicates a weaker correlation. This result can be used to remove highly correlated redundant features. For example, when the absolute value of the correlation coefficient between two feature dimensions exceeds a preset redundancy threshold (e.g., 0.9), one of the feature dimensions can be removed, reducing the computational load of subsequent analysis and improving clustering accuracy.
[0041] The high-dimensional feature representation constructed in this stage provides a unified numerical space for the multi-cluster centroid clustering analysis in the subsequent step S2, enabling the model to simultaneously understand the structural, semantic, and behavioral dependency features between interfaces.
[0042] In step S2, based on the unified feature matrix, several cluster centers are constructed using a clustering algorithm to form a set of cluster centers, including the following steps S21~S23: Traditional single-center clustering methods can only identify cluster structures with a single modality. However, API dependencies in complex microservice systems often exhibit multimodal characteristics. APIs from different business modules form different dependency clusters, and APIs within the same module may form further subdivided dependency clusters. Single-center clustering cannot accurately capture such complex dependency structures. Therefore, this stage adopts a multi-center clustering mechanism to address the above problems.
[0043] S21. Based on the unified feature matrix, determine the number of target clusters through clustering evaluation criteria.
[0044] The number of target clusters C is not manually specified, but automatically determined by clustering evaluation criteria. The clustering evaluation criteria analyze the distribution characteristics of API samples in the unified feature matrix F, evaluate the clustering quality under different numbers of clusters, and select the number of clusters corresponding to the optimal clustering quality as the number of target clusters C.
[0045] Specifically, the clustering evaluation criterion can be achieved through the BIC criterion: The BIC criterion determines the number of target clusters by balancing model fit and model complexity. The corresponding BIC value is calculated for different numbers of candidate clusters. The BIC value comprehensively considers the degree of model fit to the data and the complexity penalty brought by the number of model parameters. The number of candidate clusters with the optimal BIC value is selected as the number of target clusters C.
[0046] Clustering evaluation criteria can also be achieved through the Silhouette coefficient: The Silhouette coefficient evaluates the clustering quality by measuring the closeness of each API sample to other samples in its own cluster and the degree of separation from the nearest neighbor cluster samples. The average Silhouette coefficient of all API samples is calculated for different numbers of candidate clusters, and the number of candidate clusters corresponding to the largest average Silhouette coefficient is selected as the number of target clusters C.
[0047] S22. Perform multi-center clustering partitioning on API samples in the unified feature matrix, and divide the API samples into API dependency clusters corresponding to the number of target clusters.
[0048] Specifically, after determining the number of target clusters C, a multi-center clustering partitioning algorithm is used based on the unified feature matrix F to divide all API samples into C API dependency clusters.
[0049] Multi-center clustering can be achieved using a Gaussian mixture model (GMM). The GMM assumes that the API samples in the unified feature matrix F follow a probability distribution composed of a mixture of C Gaussian distributions. Through an expectation-maximization iterative process, the mean, covariance, and mixture weight parameters of each Gaussian component are alternately estimated until the parameters converge. Finally, each API sample is assigned to the API dependency cluster corresponding to the Gaussian component with the highest posterior probability. This approach is suitable for scenarios where the API feature distribution is Gaussian.
[0050] Furthermore, multi-center clustering can also be implemented using the K-Medoids algorithm: The K-Medoids algorithm selects C actual sample points from the API samples in the unified feature matrix F as initial centers. By iteratively exchanging center points and non-center points, it minimizes the sum of distances from all samples within each cluster to its center, and finally assigns each API sample to the API dependency cluster corresponding to the nearest center. This method is suitable for scenarios with outliers, such as when API call logs contain abnormal delay data. Compared to mean-based clustering algorithms, the K-Medoids algorithm has stronger robustness.
[0051] S23. Determine the feature vector of the API sample that achieves the optimal feature aggregation degree within each API dependency cluster as the corresponding cluster center, and form a cluster center set from all cluster centers.
[0052] Specifically, after completing the multi-center clustering partitioning, it is necessary to determine the cluster centers corresponding to each API dependency cluster. The calculation of cluster centers follows the minimum distance principle, and its core logic is: each cluster center... (j from 1 to 0) is the sample point in the unified feature matrix F that minimizes the sum of distances from all samples in the cluster to the center, i.e.: , Where argminx∈F represents finding the sample point x in the feature matrix F that minimizes the value of the subsequent expression; this sample point is the cluster center cj; F i This represents the i-th sample point in the feature matrix F; Fi represents the square of the Euclidean distance between sample point x and Fi, used to measure the similarity between two API feature vectors; the smaller the distance, the higher the similarity. This represents the sum of squared Euclidean distances between sample point x and all API feature vectors within the cluster. The sample point with the smallest value is the center of the cluster, which best represents the common features of the APIs within the cluster.
[0053] Furthermore, the set of cluster centers is composed of all C cluster centers: , Among them, each cluster center c jIt is a d-dimensional feature vector, where d is the total number of feature dimensions, representing the core features of the API within this cluster.
[0054] The cluster center set output in this stage provides a benchmark reference point for the calculation of membership in step S3, and also provides a geometric basis for the construction of the hyperplane in step S4. Subsequently, a separating hyperplane will be constructed for each pair of cluster centers in the cluster center set to realize the quantification of inter-cluster dependencies.
[0055] In step S3, based on the relationship between each API sample and each cluster center, the membership degree of each API sample to each cluster center is determined, including the following steps S31~S32: The current step calculates the membership degree between each API sample and each cluster center based on the cluster center set output in step S2. Membership degree measures the degree to which an API sample belongs to a cluster, with a value ranging from [0,1]. A higher membership degree indicates a stronger association between the API sample and the corresponding cluster center. Unlike hard clustering where each sample belongs to only a single cluster, this stage uses a soft membership mechanism, allowing the same API sample to belong to multiple clusters to varying degrees. This enables the depiction of boundary interfaces or cross-module interfaces that are simultaneously associated with multiple dependent clusters.
[0056] S31. Calculate the feature space distance between the feature vector of each API sample and each cluster center in the cluster center set.
[0057] Specifically, for each API sample Fi (i from 1 to n) in the unified feature matrix F, its eigenvector and cluster center set are calculated respectively. The center of each cluster c j The feature space distance (j from 1 to 0) is measured using the square of the Euclidean distance. Among them, F i Let c be the feature vector of the i-th API sample. j Let j be the eigenvector of the j-th cluster center. This represents the squared Euclidean distance between two feature vectors in a unified feature space. The smaller this value, the closer the features of the i-th API sample are to the features of the j-th cluster center.
[0058] Through this step, each API sample will obtain C feature space distance values, which correspond to the distance relationship between it and C cluster centers, forming the distance distribution vector of the API sample.
[0059] S32. Based on the feature space distance, the membership degree of each API sample to each cluster center is obtained through probability normalization. The sum of the membership degrees of each API sample to all cluster centers is one.
[0060] Based on the feature space distance calculated in step S31, the distance value is converted into a membership degree through probability normalization. The core logic of membership degree calculation is: the closer the feature space distance, the higher the membership degree; the farther the feature space distance, the lower the membership degree.
[0061] Specifically, probability normalization can be achieved using the softmax function, which is calculated as follows: , in, This represents the membership degree of the i-th API sample to the j-th cluster center; σ represents an exponential function used to convert distance into a positive value; 2 The bandwidth parameter, σ, controls the concentration of membership distribution. 2 The value of σ is automatically determined by the variance of the unified feature matrix F, and is usually set to a preset multiple of the feature variance (e.g., 1.2 times). 2 The larger the membership degree, the more uniform the distribution. 2 The smaller the membership degree, the more concentrated the distribution; This represents the summation of the exponent terms of all C cluster centers for normalization, ensuring that the membership values are in the range of [0,1] and that the sum of the memberships of each API sample to all cluster centers is one.
[0062] Furthermore, probability normalization can also be achieved through Gaussian kernel normalization: Gaussian kernel normalization uses the feature space distance between each API sample and each cluster center as the input of the Gaussian kernel function, calculates the Gaussian kernel response value corresponding to each distance value, and then divides each kernel response value by its sum to obtain the normalized membership degree. The difference between Gaussian kernel normalization and the softmax function lies in the different decay patterns of their kernel functions. Gaussian kernel normalization decays more gently when the distance is large, making it suitable for scenarios where the boundaries between API-dependent clusters are relatively blurred.
[0063] After the above two steps, this stage outputs the membership matrix. , where n is the total number of APIs and C is the number of clusters. Matrix element P ij This represents the membership degree of the i-th API to the j-th cluster.
[0064] The membership matrix P will be used in conjunction with the separating hyperplane in the subsequent step S4: when calculating the inter-cluster dependency strength, the membership degree needs to be combined to distinguish the contribution of different APIs to the inter-cluster dependency. The higher the membership degree of the API sample, the greater its contribution to the dependency strength of its cluster pair, thereby ensuring the accuracy of the dependency measurement.
[0065] In step S4, a separating hyperplane is constructed for each pair of cluster centers in the cluster center set. The dependency strength between each pair of cluster centers is determined by combining the membership degree with the separating hyperplane, and a dependency strength matrix is constructed, including the following steps S41~S43. S41. For each pair of cluster centers in the set of cluster centers, construct a separating hyperplane in the feature space where the unified feature matrix is located to separate the corresponding two API-dependent clusters.
[0066] For each pair of cluster centers in the cluster center set, in the d-dimensional feature space containing the unified feature matrix F, construct an optimal separating hyperplane to separate the corresponding two API-dependent clusters. The mathematical representation of this hyperplane is as follows: , in, The normal vector of the hyperplane determines the orientation of the hyperplane in the feature space; The bias term of the hyperplane determines its position in the feature space; x is any point in the feature space. Geometrically, this hyperplane places the two corresponding API dependency clusters on opposite sides of the hyperplane, forming the geometric separation boundary between the two clusters. For the case where the set of cluster centers contains C cluster centers, a total of C(C-1) / 2 separating hyperplanes need to be constructed, corresponding to all cluster pair combinations.
[0067] Hyperplane parameters and The determination is based on the positional relationship between the two corresponding cluster centers: normal vector The bias term is determined along the line connecting the centers of the two clusters. j The hyperplane is positioned between the two cluster centers based on the midpoint of the two cluster centers, thus achieving a balanced separation of the two API-dependent clusters.
[0068] S42. Calculate the projection distance of each API sample to the separating hyperplane, perform nonlinear decay transformation on the projection distance, and obtain the local dependency contribution of each API sample to the corresponding cluster pair. For each separating hyperplane ( , ), calculate the unified feature matrix F for each API sample Geometric projection distance to the hyperplane: , in, API Sample The function value after substituting into the hyperplane equation; |·| is the absolute value operation to ensure that the distance is non-negative; Normal vector The L2 norm is used to normalize the function value to the actual geometric distance. This projected distance reflects the degree to which the API sample deviates from the inter-cluster separation boundary in the feature space: the smaller the projected distance, the closer the API sample is to the boundary region of the two clusters, and the stronger its association with both clusters; the larger the projected distance, the farther the API sample is from the boundary, and the deeper it is into the interior of a cluster.
[0069] After obtaining the projected distance, a nonlinear decay transformation is performed, mapping the distance value to a local dependency contribution. The core logic of the nonlinear decay transformation is: the smaller the projected distance of an API sample, the greater its dependency contribution to the corresponding cluster pair; the larger the projected distance of an API sample, the more rapidly its contribution decays to near zero. The transformation uses an exponential decay function. The exponential decay function takes the square of the projected distance as the negative input of the exponent, so that API samples close to the hyperplane produce a contribution close to 1, while API samples far from the hyperplane produce a contribution close to 0, thus focusing on the API in the inter-cluster boundary region.
[0070] S43. Aggregate the local dependency contributions based on membership degree to obtain the dependency strength between each pair of cluster centers. The dependency strength matrix is composed of the dependency strengths of all cluster pairs.
[0071] After obtaining the local dependency contribution of each API sample to each cluster pair in step S42, the local dependency contribution is weighted and aggregated in combination with the membership matrix P output in step S3 to determine the dependency strength between each pair of cluster centers.
[0072] For cluster center pairs Its dependence strength The calculation method is as follows: , Where n is the total number of API samples, For the k-th API sample to the cluster pair The projected distance corresponding to the separating hyperplane. This represents the contribution to the corresponding local dependencies.
[0073] It should be further explained that during the aggregation process, the elements in the membership matrix P... and Represent the cluster of the k-th API sample respectively with cluster The degree of membership of API samples determines the strength of their dependence on the cluster pair. API samples with higher membership contribute more to the dependence strength of the cluster pair, while API samples with lower membership contribute less. This avoids interference from API samples with weaker association with the cluster pair in the calculation of dependence strength.
[0074] Iterate through all C(C-1) / 2 cluster pairs in the cluster centroid set and calculate the dependency strength for each cluster pair. Assemble all dependency strengths into a dependency strength matrix. The matrix is a symmetric matrix, and the element in the i-th row and j-th column is a cluster. with cluster j Dependence strength between The diagonal elements represent the degree of self-dependency within the cluster. A larger value indicates a tighter coupling relationship between the two clusters. The smaller the value, the weaker the dependency between the two clusters.
[0075] Step S5 involves generating a decoupling strategy based on the dependency strength matrix, including the following steps S51-S54: It should be noted that the dependency strength matrix output in step S4 Expert Knowledge Base This content is derived from the experience and industry standards of domain experts, including basic decoupling principles (such as the principle of minimum invasiveness and the principle of business continuity), decoupling methods for common coupling types (such as interface aggregation decoupling, middleware decoupling, and data redundancy decoupling), and decoupling constraints (such as system performance constraints, cost constraints, and schedule constraints); historical optimization records. It originates from the system's historical decoupling records and optimization results, including past decoupling strategies, decoupling objects, decoupling effects, and problems and solutions encountered during the decoupling process.
[0076] Reference Figure 2 S51. Based on the dependency strength matrix, cluster pairs whose dependency strength satisfies the coupling condition are identified as target coupled cluster pairs.
[0077] Specifically, based on the dependency strength matrix output in step S4 traverse the dependency strength of all cluster pairs in the matrix Cluster pairs whose dependency strength satisfies the coupling condition are identified as target coupled cluster pairs, forming a set of target coupled cluster pairs: , Where τ is a preset dependency strength threshold, For clusters with cluster The strength of the dependency between them, when When the threshold exceeds τ, it indicates a strong coupling relationship between the cluster pairs, requiring decoupling treatment. The value of the dependency strength threshold τ can be set according to the coupling tolerance of the actual system, or it can be adaptively determined based on the statistical distribution of all dependency strength values in the dependency strength matrix. For example, the mean of all dependency strength values plus a standard deviation can be used as the threshold.
[0078] S52. Construct a hierarchical decoupling graph based on the dependency relationship between target coupled cluster pairs.
[0079] Specifically, based on the target coupled cluster pair set Phigh determined in step S51, a hierarchical decoupling graph is constructed. A hierarchical decoupling graph is a directed graph structure where nodes correspond to API dependency clusters within a target coupling cluster pair, edges correspond to dependencies between target coupling cluster pairs, and the weight of an edge represents the strength of that dependency. .
[0080] Furthermore, in the hierarchical decoupling diagram, the coupling clusters of each target are hierarchically divided according to the strength of their dependencies: clusters with higher dependency strength are located at the upper level of the decoupling diagram, representing priority decoupling objects; clusters with lower dependency strength are located at the lower level of the decoupling diagram, representing secondary priority decoupling objects. This hierarchical structure allows the decoupling process to proceed layer by layer in descending order of dependency strength, avoiding cascading effects caused by improper decoupling order.
[0081] S53. Generate a set of candidate decoupling strategies based on the hierarchical decoupling graph.
[0082] The hierarchical decoupling diagram constructed based on step S52 For the target coupling clusters at each level in the diagram, combined with the expert knowledge base The common coupling types and decoupling methods recorded in the document are used to generate a set of candidate decoupling strategies.
[0083] Specifically, each candidate decoupling strategy includes the decoupling object (i.e., the target coupling cluster pair), the decoupling method (such as interface aggregation decoupling, middleware decoupling, and data redundancy decoupling), the decoupling order (determined based on the hierarchical relationship in the decoupling diagram), and the expected effect evaluation.
[0084] For the same target coupled cluster pair, there may be multiple feasible decoupling methods. Each method differs in decoupling effect, implementation cost, and impact on system operation. Therefore, multiple candidate decoupling strategies are formed, which together constitute a candidate decoupling strategy set.
[0085] S54. Combining decoupling constraints, determine the target decoupling strategy from the set of candidate decoupling strategies through strategy evaluation.
[0086] Specifically, based on the candidate decoupling strategy set generated in step S53, and combined with the decoupling constraints, the target decoupling strategy is determined from the candidate decoupling strategy set through strategy evaluation. The decoupling constraints are derived from an expert knowledge base and include system performance constraints, cost constraints, and schedule constraints. These constraints are used to limit the scope of decoupling strategy generation and ensure that the decoupling strategy meets actual business needs.
[0087] The objective of strategy evaluation is expressed as: , in, represents the target decoupling strategy; argmax(π) represents finding the strategy that maximizes the expected reward value among all candidate decoupling strategies π; E[·] represents the expectation operation; Representative strategy π in hierarchical decoupling diagram With expert knowledge base The reward function under constraints. The calculation of the reward function comprehensively considers three core indicators: the decrease in dependency strength, the improvement rate of system performance, and the decoupling cost. Each of the three indicators is assigned a corresponding evaluation weight. The higher the reward value, the better the decoupling effect of the strategy and the lower the cost.
[0088] Furthermore, policy evaluation can be implemented using reinforcement learning algorithms, modeling the selection of candidate decoupling policies as a sequential decision-making process, using a hierarchical decoupling graph. The current state is used as the state input, candidate decoupling strategies are used as optional actions, and the reward function R is used as the reward signal. Through iterative optimization of the policy gradient, the strategy with the maximum expected reward is output as the target decoupling strategy after convergence. Policy evaluation can also be achieved using a multi-objective genetic algorithm: each strategy in the candidate decoupling strategy set is encoded as an individual, and multiple optimization objectives are used, including the decrease in dependency strength, the improvement rate of system performance, and the decoupling cost. Through iterative evolution via genetic operations such as selection, crossover, and mutation, the strategy with the best overall performance is finally selected from the Pareto front as the target decoupling strategy.
[0089] This phase outputs three core results: the target decoupling strategy, the hierarchical dependency decomposition scheme, and the strategy execution report. It records the specific content of the target decoupling strategy, the decoupling objects, the decoupling order, the expected effects, and the satisfaction of constraints, providing an operational basis for decoupling execution.
[0090] This stage transforms the dependency measurement results output from step S4 into actionable optimization schemes. At the same time, the output target decoupling strategy will serve as input for the subsequent model self-learning and dynamic optimization stages, used to evaluate the decoupling effect and provide feedback for optimizing model parameters.
[0091] It should be further explained that, with reference to Figure 3 Determining the target decoupling strategy is a real-time, dynamic process.
[0092] Specifically, to obtain the data from the newly added API calls Following the feature extraction and fusion processing method in step S1, the newly added API call data is structured and fused to obtain the feature distribution corresponding to the new data. The feature distribution Pt+1 is compared with the original feature distribution Pt corresponding to the unified feature matrix F in step S1, and the distribution offset between the two is calculated.
[0093] Distribution shift is achieved using KL divergence: KL divergence measures the degree of difference between two probability distributions, calculating the original feature distribution Pt relative to the newly added feature distribution. The amount of information loss is determined by the KL divergence value. A larger KL divergence value indicates a greater difference between the two distributions, meaning a more severe drift in the feature distribution.
[0094] Specifically, the judgment criteria are as follows: , in, The preset drift threshold is used. When the distribution offset exceeds δ, it indicates that the system's operating state has changed significantly, the current model parameters are no longer applicable, and a model update needs to be triggered.
[0095] It should be noted that the preset drift threshold It can be set based on historical data and experience values.
[0096] In response to the distribution offset exceeding the preset drift threshold δ, the model parameters are dynamically adjusted based on the newly added API call data. The newly added API call data is incorporated into the unified feature matrix F, and the multi-center clustering is re-executed according to the clustering algorithm in step S2, updating the position of each cluster center in the cluster center set.
[0097] Furthermore, based on the updated cluster center set, the separating hyperplanes corresponding to each cluster pair are reconstructed according to the hyperplane construction method in step S4, and the normal vectors and bias term parameters in the hyperplane mapping parameter set 𝒲 are updated.
[0098] Simultaneously, a self-learning mechanism is introduced to update the model parameter weights. The parameter update method is as follows: , in, For the current model parameter set, For the updated model parameter set, For learning rate, For dependency loss function For the current parameter The gradient is used to adjust the model parameters along the direction that decreases the value of the dependency loss function, so that the updated model can more accurately reflect the current dependency structure of the system.
[0099] Based on the updated set of cluster centers and the set of parameters of the separating hyperplane, the membership degree of each API sample to each cluster center is recalculated according to step S3. The projection distance of each API sample to the separating hyperplane is recalculated according to step S4, and nonlinear decay transformation and membership degree weighted aggregation are performed to obtain the updated dependency strength matrix. Based on the updated dependency strength matrix, the target decoupling strategy is regenerated according to steps S51 to S54.
[0100] Example 3 is an embodiment of the present invention. This embodiment provides an intelligent decoupling system for API dependencies based on hyperplane multi-cluster space analysis, including: a feature matrix construction module, a set construction module, a calculation module, a dependency matrix construction module, and a generation strategy module; The feature matrix construction module is used to extract and fuse features from the associated data of the API to construct a unified feature matrix. The set construction module is used to construct several cluster centers based on a unified feature matrix and a clustering algorithm, forming a cluster center set. The calculation module is used to determine the membership degree of each API sample to each cluster center based on the relationship between each API sample and each cluster center; The dependency matrix construction module is used to construct a separating hyperplane for each pair of cluster centers in the cluster center set, and combine the membership degree with the separating hyperplane to determine the dependency strength between each pair of cluster centers, and construct a dependency strength matrix. The generation strategy module is used to generate a decoupling strategy based on the dependency strength matrix.
[0101] This embodiment also provides an electronic device applicable to a method for intelligent decoupling of API dependencies based on hyperplane multi-cluster space analysis, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for intelligent decoupling of API dependencies based on hyperplane multi-cluster space analysis proposed in the above embodiment.
[0102] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements an API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis as proposed in the above embodiment.
[0103] The storage medium proposed in this embodiment and the method for intelligent decoupling of API dependencies based on hyperplane multi-cluster space analysis proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0104] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0105] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for intelligent decoupling of API dependencies based on hyperplane multi-cluster space analysis, characterized in that, include: Feature extraction and fusion processing are performed on the associated data of the API to construct a unified feature matrix; Based on a unified feature matrix, several cluster centers are constructed using a clustering algorithm to form a set of cluster centers; Based on the relationship between each API sample and each cluster center, determine the membership degree of each API sample to each cluster center; For each pair of cluster centers in the cluster center set, construct a separating hyperplane, and combine the membership degree with the separating hyperplane to determine the dependency strength between each pair of cluster centers, and construct a dependency strength matrix. Decoupling strategies are generated based on dependency strength matrices.
2. The API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis as described in claim 1, characterized in that, The construction of the unified feature matrix includes: Structured encoding is performed on the associated data to obtain multi-dimensional feature representations for each API; Based on the degree of correlation between each feature dimension and API dependency in historical dependency samples, the dependency correlation weight corresponding to each feature dimension is determined. The multi-dimensional feature representations are weighted and fused based on dependency association weights to generate a unified feature matrix and the corresponding feature correlation matrix.
3. The API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis as described in claim 2, characterized in that, The process of constructing several cluster centers based on a unified feature matrix using a clustering algorithm to form a set of cluster centers includes: The number of target clusters is determined based on a unified feature matrix and clustering evaluation criteria. Perform multi-center clustering partitioning on API samples in the unified feature matrix, dividing the API samples into API dependency clusters corresponding to the number of target clusters; The feature vector of the API sample that achieves the optimal feature aggregation degree within each API dependency cluster is determined as the corresponding cluster center, and the cluster center set is composed of all cluster centers.
4. The API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis as described in claim 3, characterized in that, The process of determining the membership degree of each API sample to each cluster center based on the relationship between each API sample and each cluster center includes: Calculate the feature space distance between the feature vector of each API sample and each cluster center in the cluster center set; Based on the feature space distance, the membership degree of each API sample to each cluster center is obtained through probability normalization, and the sum of the membership degrees of each API sample to all cluster centers is one.
5. The API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis as described in claim 4, characterized in that, The process involves constructing a separating hyperplane for each pair of cluster centers in the cluster center set, and determining the dependency strength between each pair of cluster centers by combining the membership degree with the separating hyperplane. The construction of the dependency strength matrix includes: For each pair of cluster centers in the cluster center set, a separating hyperplane is constructed in the feature space where the unified feature matrix is located to separate the corresponding two API-dependent clusters; Calculate the projection distance of each API sample to the separating hyperplane, perform a nonlinear decay transformation on the projection distance, and obtain the local dependency contribution of each API sample to the corresponding cluster pair. The local dependency contribution is aggregated based on membership degree to obtain the dependency strength between each pair of cluster centers. The dependency strength matrix is composed of the dependency strengths of all cluster pairs.
6. The API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis as described in claim 5, characterized in that, The decoupling strategy based on the dependency strength matrix includes: Based on the dependency strength matrix, cluster pairs whose dependency strength satisfies the coupling condition are identified as target coupled cluster pairs; Construct a hierarchical decoupling graph based on the dependencies between target coupled cluster pairs; Generate a set of candidate decoupling strategies based on a hierarchical decoupling graph; By combining decoupling constraints, the target decoupling strategy is determined from the set of candidate decoupling strategies through strategy evaluation.
7. The API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis as described in claim 6, characterized in that, The target decoupling strategy includes: Obtain new API call data and calculate the distribution offset between the feature distribution corresponding to the new API call data and the feature distribution corresponding to the unified feature matrix. In response to the distribution offset exceeding the drift threshold, the cluster center set and the separating hyperplane are updated based on the newly added API call data. The dependency strength matrix is reconstructed based on the updated cluster center set and separating hyperplane, and the target decoupling strategy is regenerated based on the reconstructed dependency strength matrix.
8. An intelligent API dependency decoupling system based on hyperplane multi-cluster space analysis, employing the intelligent API dependency decoupling method based on hyperplane multi-cluster space analysis as described in any one of claims 1 to 7, characterized in that, include: The module includes a feature matrix construction module, a set construction module, a calculation module, a dependency matrix construction module, and a generation strategy module. The feature matrix construction module is used to extract and fuse features from the associated data of the API to construct a unified feature matrix. The set construction module is used to construct several cluster centers based on a unified feature matrix and a clustering algorithm, forming a cluster center set. The calculation module is used to determine the membership degree of each API sample to each cluster center based on the relationship between each API sample and each cluster center; The dependency matrix construction module is used to construct a separating hyperplane for each pair of cluster centers in the cluster center set, and combine the membership degree with the separating hyperplane to determine the dependency strength between each pair of cluster centers, and construct a dependency strength matrix. The generation strategy module is used to generate a decoupling strategy based on the dependency strength matrix.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the API dependency intelligent decoupling method based on hyperplane multi-cluster space analysis according to any one of claims 1 to 7.