Shadow api governance method based on self-supervised contrast and graph reasoning
Patent Information
- Application Number
- CN202610257390.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-04
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-03-04
AI Technical Summary
基于规则匹配的流量扫描方法依赖预设关键字、正则表达式或已知攻击特征进行识别,在面对未知接口或未显式暴露特征的影子API时难以准确识别;基于接口文档比对的资产盘点方法依赖OpenAPI或Swagger先验接口文档进行对照分析,当接口文档滞后或未更新时,无法覆盖真实运行中的接口资产;基于监督学习的异常检测方法则依赖大量已标注的异常样本,在实际生产环境中影子API样本数量稀缺且难以人工准确标注,导致监督模型难以训练出有效的判别边界
[0065](1) This invention constructs a graph-structured modulation TimesURL self-supervised contrastive learning mechanism in an unlabeled environment, which significantly improves the ability to characterize small distortions in high-dimensional sparse API loads. By introducing a modulation factor composed of normalized node topology centrality and node community legitimacy confidence during the enhancement stage, the frequency domain enhancement intensity and time domain perturbation intensity adapt to the position and legitimacy status of the API in the global topology. At the same time, dual-domain Universum hard negative samples are generated jointly by instance dimension, sequence dimension and neighborhood feature dimension, and the contrastive learning loss function is used to normalize the similarity constraint between the positive sample embedding vector and the multi-class hard negative sample embedding vector. A high-density discrimination region around the legitimate business boundary is formed in the low-dimensional embedding space. It can form a more refined embedding separation effect for shadow APIs that are reused in the framework but have very small differences in internal payload or call frequency without relying on manual annotation. The recognition recall rate of low-amplitude load distortion is significantly improved.
Smart Images

Figure CN122263125B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of shadow API governance technology, and in particular to a shadow API governance method based on self-supervised comparison and graph reasoning. Background Technology
[0002] With the widespread adoption of microservice architecture and cloud-native technologies, enterprise internal systems are increasingly composed of numerous fine-grained APIs. APIs, as the core channels for data interaction and business collaboration between systems, have become fundamental building blocks of modern application architectures. In this highly distributed environment, the number of APIs is growing exponentially, and their lifecycles change frequently. Development, testing, canary releases, and legacy system migrations often result in APIs that are not managed by a unified gateway or registered in the API documentation. These are commonly referred to as shadow APIs. Due to the lack of unified authentication control, access policy auditing, and asset registration, shadow APIs are highly susceptible to becoming entry points for data leaks and privilege escalation.
[0003] Existing technologies for API asset discovery and anomaly detection mainly include rule-matching-based traffic scanning methods, interface document comparison-based asset inventory methods, and supervised learning-based anomaly detection methods. Rule-matching-based traffic scanning methods rely on preset keywords, regular expressions, or known attack characteristics for identification, making them difficult to accurately identify unknown interfaces or shadow APIs that do not explicitly expose characteristics. Interface document comparison-based asset inventory methods rely on prior interface documents from OpenAPI or Swagger for comparative analysis; when these documents are outdated or not updated, they cannot cover interface assets in actual operation. Supervised learning-based anomaly detection methods rely on a large number of labeled anomaly samples; however, in real-world production environments, the number of shadow API samples is scarce and difficult to accurately label manually, making it difficult for supervised models to train effective discrimination boundaries.
[0004] In unlabeled traffic scenarios, some existing self-supervised time series learning methods attempt to learn representations of traffic data through random masking, random pruning, or contrastive learning. However, API call traffic typically has high-dimensional, sparse, and structured characteristics. Random masking or random pruning operations may destroy the semantic structure of key fields, such as authentication fields or path parameter structures, leading to distorted learned representations. Meanwhile, the negative samples constructed by traditional contrastive learning methods are often too different from normal traffic, making it easy for the model to learn broad decision boundaries and failing to characterize the subtle distortions of shadow APIs in high-dimensional load spaces. Summary of the Invention
[0005] One objective of this invention is to propose a shadow API governance method based on self-supervised comparison and graph reasoning. This invention can achieve a more refined embedding and separation effect for shadow APIs that are reused in the framework but have very small differences in internal payload or call frequency without relying on manual annotation, and significantly improves the recognition and recall rate for low-amplitude load distortion.
[0006] A shadow API governance method based on self-supervised comparison and graph reasoning according to an embodiment of the present invention includes:
[0007] By capturing API request and response traffic through gateway-side bypass, a multivariate time series tensor matrix is constructed and an initial API global spatiotemporal heterogeneous graph is generated.
[0008] Node topological centrality and community legitimacy confidence are calculated on the initial API global spatiotemporal heterogeneous graph. The node topological centrality and community legitimacy confidence are used as modulation factors to improve the TimesURL self-supervised feature learning network. Graph structure-guided frequency domain enhancement and time domain enhancement are performed on the mounted historical multivariate time series tensor matrix, and the modulated enhanced sample set is output.
[0009] An improved TimesURL self-supervised feature learning network is used to generate dual-domain Universum hard negative samples by jointly combining instance dimension, sequence dimension and neighborhood feature dimension.
[0010] The enhanced sample set and the dual-domain Universum hard negative samples are simultaneously input into the improved TimesURL self-supervised feature learning network to perform contrastive learning to obtain low-dimensional embedding vectors for API endpoints.
[0011] Write back the low-dimensional embedding vector of the API endpoint to the initial API global spatiotemporal heterogeneous graph as a node attribute to obtain the embedding-enhanced API global spatiotemporal heterogeneous graph, and output the fused node feature vector.
[0012] A spatiotemporal anomaly potential energy function is constructed based on the feature vector of the fusion node. The potential energy value is calculated for all API endpoints in the global spatiotemporal heterogeneous graph of the embedded enhanced API. API endpoints with potential energy values greater than a preset threshold are marked as shadow API candidate nodes.
[0013] Dynamic graph difference analysis is performed on the embedded enhanced API global spatiotemporal heterogeneous graph of candidate nodes in multi-time slices to screen the target nodes of shadow API that meet the dynamic abnormal evolution conditions and map them to the security governance strategy engine. Access control policies are generated according to preset blocking rules and distributed to the network gateway to implement automated circuit breaking, rate limiting or isolation governance of target nodes of shadow API.
[0014] Optionally, the step of constructing a multivariate time series tensor matrix and generating an initial API global spatiotemporal heterogeneous graph includes:
[0015] On the network gateway side, lossless bypass capture is performed on real-time API request and response streams. According to the unified protocol parsing rules, the URL path characteristics, Header parameter entropy, Query parameter change rate, JSONPayload nesting depth, and timestamps of the real-time API request and response streams are encoded into multivariate time series tensor matrices. A corresponding historical multivariate time series tensor matrix is established for each API endpoint. Based on the historical multivariate time series tensor matrix, node abstraction is performed on API endpoints, calling source IP entities, authentication identity entities, and microservice module instances. Edge abstraction is performed based on call relationships, parameter passing relationships, and permission links to construct an initial global spatiotemporal heterogeneous graph of APIs. The historical multivariate time series tensor matrix is mounted as the original feature of the nodes on the initial global spatiotemporal heterogeneous graph of APIs.
[0016] Optionally, the step of using node topological centrality and community legitimacy confidence as modulation factors to improve the TimesURL self-supervised feature learning network includes:
[0017] A topological graph structure is constructed based on the initial API global spatiotemporal heterogeneous graph. The topological graph structure consists of a set of nodes and a set of edges.
[0018] Calculate the topological centrality of any node in the node set, and calculate the normalized topological centrality.
[0019] Divide the node set into several community sets, and calculate the community legitimacy confidence score for any community set;
[0020] For any node in the node set, determine the community set to which the node belongs, and assign the community legitimacy confidence of the corresponding community set to the node community legitimacy confidence of the corresponding node. The normalized node topological centrality and the node community legitimacy confidence are weighted and summed to obtain the modulation factor.
[0021] The enhancement strength coefficient is obtained by multiplying the value obtained by subtracting the modulation factor by a preset enhancement reference strength coefficient, and then calculating with the historical multivariate time series tensor matrix to obtain the frequency domain enhanced tensor matrix.
[0022] The TimesURL self-supervised feature learning network is improved based on the enhancement intensity coefficient input, and graph-structure-guided temporal enhancement is performed to obtain the temporally enhanced tensor matrix.
[0023] The frequency-domain enhanced tensor matrix and the time-domain enhanced tensor matrix are aggregated as modulated enhanced samples of the same historical multivariate time series tensor matrix to obtain a set of modulated enhanced samples.
[0024] Optionally, the improved TimesURL self-supervised feature learning network is used in conjunction with the instance dimension, sequence dimension, and neighborhood feature dimension, including:
[0025] When generating time-series Universum hard negative samples at the instance dimension, boundary approximation interpolation coefficients are obtained by linearly mapping the modulation factor between the lower bound and the upper bound of the preset interpolation coefficients.
[0026] The frequency-domain enhanced tensor matrix corresponding to the target node and the time-domain enhanced tensor matrix corresponding to the reference node are linearly combined element-wise according to the boundary approximation interpolation coefficients to obtain the instance-level perturbation tensor matrix.
[0027] When generating temporal domain Universum hard negative samples in the sequence dimension, the sequence-level perturbation segmentation index is calculated based on the modulation factor;
[0028] The time segment of the time domain-enhanced tensor matrix corresponding to the target node before the sequence-level perturbation segmentation index is retained, and the time segment of the frequency domain-enhanced tensor matrix corresponding to the reference node after the sequence-level perturbation segmentation index is concatenated to the time segment of the target node to obtain the sequence-level perturbation tensor matrix.
[0029] When generating temporal domain Universum hard negative samples in the neighborhood feature dimension, the first-order neighborhood node set of the target node is determined based on the initial API global spatiotemporal heterogeneous graph, and the neighborhood node is selected. The frequency domain enhanced tensor matrix corresponding to the target node and the frequency domain enhanced tensor matrix corresponding to the neighborhood node are linearly combined element by element according to the preset neighborhood feature perturbation interpolation coefficients to obtain the neighborhood feature perturbation tensor matrix.
[0030] The instance-level perturbation tensor matrix, the sequence-level perturbation tensor matrix, and the neighborhood feature perturbation tensor matrix are converged to obtain the time-series domain Universum hard negative sample set;
[0031] Based on the initial API global spatiotemporal heterogeneous graph, a pseudo-topological graph is constructed.
[0032] The pseudo-topological graph is used as a set of hard negative samples in the topological domain Universum and hard negative samples in the temporal domain to generate dual-domain Universum hard negative samples.
[0033] Optionally, the step of simultaneously inputting the enhanced sample set and dual-domain Universum hard negative samples into the improved TimesURL self-supervised feature learning network includes:
[0034] The instance-level perturbation tensor matrix, sequence-level perturbation tensor matrix, and neighborhood feature perturbation tensor matrix corresponding to the target node in the temporal domain Universum hard negative sample set are respectively input into the encoder of the TimesURL self-supervised feature learning network to obtain the instance-dimensional temporal domain Universum hard negative sample embedding vector, the sequence-dimensional temporal domain Universum hard negative sample embedding vector, and the neighborhood feature-dimensional temporal domain Universum hard negative sample embedding vector.
[0035] Extract the temporal domain Universum hard negative sample embedding vector corresponding to the neighborhood node set into the topological domain Universum hard negative sample embedding vector.
[0036] Positive sample pairs are formed by the embedding vectors of the target node in the frequency domain enhancement perspective and the embedding vectors of the target node in the time domain enhancement perspective. Negative sample sets are formed by the instance-dimensional time domain Universum hard negative sample embedding vectors, the sequence-dimensional time domain Universum hard negative sample embedding vectors, the neighborhood feature-dimensional time domain Universum hard negative sample embedding vectors, and the topology domain Universum hard negative sample embedding vectors. A contrastive learning loss function is constructed.
[0037] The contrastive learning loss function is used to calculate the corresponding contrastive learning loss value for each target node in the node set, and the contrastive learning loss values of all target nodes are summed to obtain the batch loss value.
[0038] Based on the batch loss value, the set of network weight parameters of the TimesURL self-supervised feature learning network is updated by gradient to obtain the updated set of network weight parameters.
[0039] After the network weight parameter set is updated, the updated network weight parameter set is used to perform forward inference on the modulated and enhanced sample corresponding to any API endpoint in the node set to obtain the low-dimensional embedding vector of the API endpoint.
[0040] Optionally, the step of writing back the low-dimensional embedding vector of the API endpoint to the initial API global spatiotemporal heterogeneous graph as a node attribute includes:
[0041] Write back the low-dimensional embedding vector of the API endpoint to the node attributes of the topological graph structure to obtain the global spatiotemporal heterogeneous graph of the embedding-enhanced API.
[0042] An anomaly function is constructed on the global spatiotemporal heterogeneous graph of embedded enhanced APIs to represent the anomaly degree of the low-dimensional embedding vector of API endpoints, and the anomaly degree is calculated based on the anomaly function for the low-dimensional embedding vector of any API endpoint in the node set.
[0043] The neighborhood gating coefficient is obtained by multiplying the anomaly by the gating decay coefficient and then taking the negative exponential function.
[0044] Based on the global spatiotemporal heterogeneous graph of the embedded enhancement API, the first-order neighboring node set of any node is determined, and the latent feature vector of the node is calculated by combining the neighborhood gating coefficient.
[0045] By using the latent feature vectors of nodes output at a preset number of layers, a fused node feature vector is constructed.
[0046] Optionally, the step of constructing a spatiotemporal anomaly potential energy function based on the fusion node feature vectors and calculating the potential energy value for all API endpoints within the global spatiotemporal heterogeneous graph of the embedded enhanced API includes:
[0047] Based on the feature vectors of the fused nodes, calculate the business baseline deviation of the nodes corresponding to the API endpoints;
[0048] Based on the embedded enhanced API global spatiotemporal heterogeneous graph, the first-order neighbor node set of the node corresponding to the API endpoint is determined, and the neighborhood consistency violation degree is calculated based on the fused node feature vector.
[0049] Based on the normalized node topological centrality and the node community legitimacy confidence, calculate the spoofing consistency of the nodes corresponding to the API endpoints;
[0050] The spatiotemporal anomaly potential energy function is obtained by weighted summation based on the deviation from the business baseline, the degree of neighborhood consistency violation, and the degree of spoofing consistency, and the potential energy value is calculated.
[0051] Calculate the potential energy value for each node corresponding to all API endpoints in the global spatiotemporal heterogeneous graph of the embedded enhanced API, and mark the nodes corresponding to API endpoints with potential energy values greater than a preset threshold as candidate nodes for shadow APIs.
[0052] Optionally, the dynamic graph difference analysis performed on the embedded augmented API global spatiotemporal heterogeneous graph of the shadow API candidate nodes in multi-time slices includes:
[0053] The global spatiotemporal heterogeneous graph of the embedded enhancement API is divided sequentially according to a fixed time window to obtain multiple time window corresponding to the slice sequence of the global spatiotemporal heterogeneous graph of the embedded enhancement API.
[0054] For any shadow API candidate node in the shadow API candidate node set, calculate the island density index in the global spatiotemporal heterogeneous graph slice sequence of the embedding enhancement API corresponding to multiple time windows;
[0055] The island density change rate is obtained by subtracting the island density index in the previous time window from the island density index in the current time window. When the island density change rate is less than zero and the island density change rate is continuously less than the preset island density change threshold in several consecutive time windows, it is determined to be an island density collapse trend.
[0056] The call frequency growth rate is obtained by subtracting the call frequency index in the previous time window from the call frequency index in the current time window. When the call frequency growth rate is greater than the preset call frequency growth threshold, and the call frequency growth rate continues to be greater than the preset call frequency growth threshold in several consecutive time windows, it is determined to be a slow penetration growth trend.
[0057] Candidate nodes that simultaneously satisfy the island density collapse trend and slow penetration growth trend are selected as the target node set for the Shadow API.
[0058] Map the topology location, permission link risk, and potential data access scope of any shadow API target node in the shadow API target node set to the security governance policy engine;
[0059] The security governance strategy engine generates access control policies based on preset blocking rules and distributes these policies to network gateways to implement automated circuit breaking, rate limiting, or isolation governance.
[0060] Optionally, the preset blocking rules include:
[0061] Access circuit breaker rules: When the risk of the access control link is greater than the preset risk threshold and the duration of the island density collapse trend is greater than the preset time threshold, an automated circuit breaker strategy will be executed.
[0062] Access rate limiting rules: When the risk of the access control link is less than the preset risk threshold and the duration of the slow penetration growth trend is greater than the preset time threshold, the rate limiting policy will be executed.
[0063] Access Isolation Rule: When the topological location of the target node of the Shadow API belongs to a community set with low community legitimacy confidence and the potential data access scope includes core data nodes, an isolation governance strategy is executed.
[0064] The beneficial effects of this invention are:
[0065] (1) This invention constructs a graph-structured modulation TimesURL self-supervised contrastive learning mechanism in an unlabeled environment, which significantly improves the ability to characterize small distortions in high-dimensional sparse API loads. By introducing a modulation factor composed of normalized node topology centrality and node community legitimacy confidence during the enhancement stage, the frequency domain enhancement intensity and time domain perturbation intensity adapt to the position and legitimacy status of the API in the global topology. At the same time, dual-domain Universum hard negative samples are generated jointly by instance dimension, sequence dimension and neighborhood feature dimension, and the contrastive learning loss function is used to normalize the similarity constraint between the positive sample embedding vector and the multi-class hard negative sample embedding vector. A high-density discrimination region around the legitimate business boundary is formed in the low-dimensional embedding space. It can form a more refined embedding separation effect for shadow APIs that are reused in the framework but have very small differences in internal payload or call frequency without relying on manual annotation. The recognition recall rate of low-amplitude load distortion is significantly improved.
[0066] (2) This invention breaks through the topological blind zone problem of single-point temporal detection. It writes back the low-dimensional embedding vector of API endpoint as graph node attributes, constructs an anomaly function on the global spatiotemporal heterogeneous graph of embedded enhanced API, obtains the business baseline deviation by calculating the cosine similarity between the feature vector of the fused node and the community embedding center vector, and generates a neighborhood gating coefficient based on the anomaly. It performs exponential decay control on the contribution of neighborhood information in the message passing process of graph neural network, so that the API with high anomaly is automatically weakened in the graph propagation process. Thus, the temporal feature consistency information, spatial connection consistency information and topological integrity information are retained in the feature vector of the fused node at the same time. It realizes anomaly-aware message passing, enabling graph inference to identify shadow APIs that are normal in temporal performance but have discontinuities in topological support, and significantly reduces the false detection rate of topologically hidden shadow APIs.
[0067] (3) This invention proposes a spatiotemporal anomaly potential function that combines business baseline deviation, neighborhood consistency violation, and masquerade consistency to achieve accurate identification of masquerade shadow APIs. The masquerade consistency is introduced into the potential function. By coupling the business consistency value with the topology support missing degree, node community legitimacy confidence, and normalized node topology centrality, a nonlinear characterization of the composite anomaly pattern is formed. This allows the potential value to increase significantly even when the business baseline deviation is low but the community legitimacy and topology centrality are both low, thus including highly concealed shadow APIs in the candidate range. The spatiotemporal anomaly potential function has significantly enhanced the ability to distinguish masquerade shadow APIs in the microservice environment, realizing refined risk quantification and screening in a multidimensional spatiotemporal feature space. Attached Figure Description
[0068] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0069] Figure 1 This is a flowchart of a shadow API governance method based on self-supervised comparison and graph reasoning proposed in this invention. Detailed Implementation
[0070] Example 1: Reference Figure 1 A shadow API governance method based on self-supervised comparison and graph reasoning includes:
[0071] By capturing API request and response traffic through gateway-side bypass, a multivariate time series tensor matrix is constructed and an initial API global spatiotemporal heterogeneous graph is generated.
[0072] In this embodiment, constructing a multivariate time series tensor matrix and generating an initial API global spatiotemporal heterogeneous graph includes:
[0073] On the network gateway side, lossless bypass capture is performed on real-time API request and response streams. According to the unified protocol parsing rules, the URL path characteristics, Header parameter entropy, Query parameter change rate, JSONPayload nesting depth, and timestamps of the real-time API request and response streams are encoded into multivariate time series tensor matrices. A corresponding historical multivariate time series tensor matrix is established for each API endpoint. Based on the historical multivariate time series tensor matrix, node abstraction is performed on API endpoints, calling source IP entities, authentication identity entities, and microservice module instances. Edge abstraction is performed based on call relationships, parameter passing relationships, and permission links to construct an initial global spatiotemporal heterogeneous graph of APIs. The historical multivariate time series tensor matrix is mounted as the original feature of the nodes on the initial global spatiotemporal heterogeneous graph of APIs.
[0074] Node topological centrality and community legitimacy confidence are calculated on the initial API global spatiotemporal heterogeneous graph. The node topological centrality and community legitimacy confidence are used as modulation factors to improve the TimesURL self-supervised feature learning network. Graph structure-guided frequency domain enhancement and time domain enhancement are performed on the mounted historical multivariate time series tensor matrix, and the modulated enhanced sample set is output.
[0075] In this embodiment, node topological centrality and community legitimacy confidence are used as modulation factors to improve the TimesURL self-supervised feature learning network, including:
[0076] A topological graph structure is constructed based on the initial API global spatiotemporal heterogeneous graph. The topological graph structure consists of a set of nodes and a set of edges.
[0077] The node set includes API endpoints, source IP entities, authentication entities, and microservice module instances, while the edge set includes call relationships, parameter passing relationships, and permission chains.
[0078] Calculate the topological centrality of any node in the node set, and calculate the normalized topological centrality.
[0079] In Example 1, the node topological centrality is obtained by counting the number of connections of a node in the edge set and dividing by the total number of nodes minus one. After obtaining the node topological centrality of all nodes, the minimum and maximum values of the node topological centrality are extracted. The node topological centrality of any node is then reduced by the minimum value and divided by the difference between the maximum and the minimum value to obtain the normalized node topological centrality, which is then placed within a closed interval between zero and one.
[0080] Divide the node set into several community sets, and calculate the community legitimacy confidence score for any community set;
[0081] In Example 1, each community set consists of several nodes, and the edges within a community set are edges where both ends of the node belong to the same community set. The community legitimacy confidence is obtained by counting the number of edges in the community set that satisfy the existence and consistency checks of the authentication fields and dividing by the total number of edges in the community set.
[0082] For any node in the node set, determine the community set to which the node belongs, and assign the community legitimacy confidence of the corresponding community set to the node community legitimacy confidence of the corresponding node. The normalized node topological centrality and the node community legitimacy confidence are weighted and summed to obtain the modulation factor.
[0083] The enhancement strength coefficient is obtained by multiplying the value obtained by subtracting the modulation factor by a preset enhancement reference strength coefficient, and then calculating with the historical multivariate time series tensor matrix to obtain the frequency domain enhanced tensor matrix.
[0084] In Example 1, the historical multivariate time series tensor matrix mounted on the initial API global spatiotemporal heterogeneous graph node is uniformly represented. The historical multivariate time series tensor matrix consists of the number of discrete time steps within the historical time window and the number of feature dimensions composed of URL path features, Header parameter entropy value, Query parameter change rate, JSONPayload nesting depth and timestamp.
[0085] The frequency domain representation is obtained by performing a frequency domain transformation on the historical multivariate time series tensor matrix along the time dimension. A preset high-frequency band position is selected in the frequency domain representation and multiplied element-wise with a frequency domain perturbation tensor of the same shape. Then, the result is multiplied by an enhancement intensity coefficient to obtain the frequency domain perturbation result. The frequency domain perturbation result is added to the original frequency domain representation and then an inverse frequency domain transformation is performed to obtain the frequency domain enhanced tensor matrix.
[0086] The TimesURL self-supervised feature learning network is improved based on the enhancement intensity coefficient input, and graph-structure-guided temporal enhancement is performed to obtain the temporally enhanced tensor matrix.
[0087] In Example 1, temporal augmentation applies constrained perturbations to the sampling positions and corresponding eigenvalues in the time dimension without changing the number of feature dimensions and time steps of the historical multivariate time series tensor matrix. The constrained perturbations satisfy that the absolute value of the difference between the augmented element value and the original element value at any time index and any feature dimension index does not exceed the augmentation intensity coefficient, thus obtaining the temporally augmented tensor matrix.
[0088] The frequency-domain enhanced tensor matrix and the time-domain enhanced tensor matrix are aggregated as modulated enhanced samples of the same historical multivariate time series tensor matrix to obtain a set of modulated enhanced samples.
[0089] An improved TimesURL self-supervised feature learning network is used to generate dual-domain Universum hard negative samples by jointly combining instance dimension, sequence dimension and neighborhood feature dimension.
[0090] In this embodiment, an improved TimesURL self-supervised feature learning network is used to jointly address the instance dimension, sequence dimension, and neighborhood feature dimension, including:
[0091] When generating time-series Universum hard negative samples at the instance dimension, boundary approximation interpolation coefficients are obtained by linearly mapping the modulation factor between the lower bound and the upper bound of the preset interpolation coefficients.
[0092] In Example 1, the boundary approximation interpolation coefficients are used to control the linear combination ratio of the frequency-domain enhanced tensor matrix and the time-domain enhanced tensor matrix corresponding to the reference node, so that the boundary approximation interpolation coefficients change with the changes in the node topological centrality and the node community legitimacy confidence.
[0093] ;
[0094] in, Indicates the lower bound of the preset interpolation coefficients and satisfies Indicates the upper bound of the preset interpolation coefficients and satisfies and , For modulation factor, This represents the boundary approximation interpolation coefficients.
[0095] The frequency-domain enhanced tensor matrix corresponding to the target node and the time-domain enhanced tensor matrix corresponding to the reference node are linearly combined element-wise according to the boundary approximation interpolation coefficients to obtain the instance-level perturbation tensor matrix.
[0096] In Example 1, based on the modulated enhanced sample set and the initial API global spatiotemporal heterogeneous map, the corresponding frequency-domain enhanced tensor matrix and time-domain enhanced tensor matrix are determined for any node in the node set, and the frequency-domain enhanced tensor matrix and the time-domain enhanced tensor matrix are uniformly constrained to be the same type of tensor matrix as the historical multivariate time series tensor matrix.
[0097] The instance-level perturbation tensor matrix maintains the same discrete time steps within the historical time window as the historical multivariate time series tensor matrix, along with the feature dimensions consisting of URL path features, Header parameter entropy, Query parameter change rate, JSONPayload nesting depth, and timestamps. This is used to generate time-domain Universum hard negative samples that closely approximate the boundaries of legitimate business operations in the shadow API governance scenario.
[0098] When generating temporal domain Universum hard negative samples in the sequence dimension, the sequence-level perturbation segmentation index is calculated based on the modulation factor;
[0099] The sequence-level perturbation segmentation index is used to segment and concatenate the time-domain enhanced tensor matrix corresponding to the target node and the frequency-domain enhanced tensor matrix corresponding to the reference node in the time dimension, while keeping the number of discrete time steps within the historical time window unchanged.
[0100] ;
[0101] in, This indicates the floor operator. This is the sequence-level perturbation segmentation index, where T represents the time step.
[0102] The time segment of the time domain-enhanced tensor matrix corresponding to the target node before the sequence-level perturbation segmentation index is retained, and the time segment of the frequency domain-enhanced tensor matrix corresponding to the reference node after the sequence-level perturbation segmentation index is concatenated to the time segment of the target node to obtain the sequence-level perturbation tensor matrix.
[0103] When generating temporal domain Universum hard negative samples in the neighborhood feature dimension, the first-order neighborhood node set of the target node is determined based on the initial API global spatiotemporal heterogeneous graph, and the neighborhood node is selected. The frequency domain enhanced tensor matrix corresponding to the target node and the frequency domain enhanced tensor matrix corresponding to the neighborhood node are linearly combined element by element according to the preset neighborhood feature perturbation interpolation coefficients to obtain the neighborhood feature perturbation tensor matrix.
[0104] The neighborhood feature perturbation tensor matrix is used to boundary-mix the business features of the same call chain or permission chain neighborhood with the business features of the target node in the shadow API governance scenario.
[0105] The instance-level perturbation tensor matrix, the sequence-level perturbation tensor matrix, and the neighborhood feature perturbation tensor matrix are converged to obtain the time-series domain Universum hard negative sample set;
[0106] Based on the initial API global spatiotemporal heterogeneous graph, a pseudo-topological graph is constructed.
[0107] In Example 1, an abnormal edge deletion set is obtained by deleting edges in the edge set that do not meet the existence and consistency verification of the authentication field. Then, a pseudo-legitimate edge insertion set is obtained by inserting directed edges that do not exist in the original edge set into the community set that meet the community legitimacy confidence level of not less than the preset community legitimacy threshold. The abnormal edge deletion set is removed from the original edge set and the pseudo-legitimate edge insertion set is added to the original edge set to form a pseudo-topology graph.
[0108] The pseudo-topological graph is used as a set of hard negative samples in the topological domain Universum and hard negative samples in the temporal domain to generate dual-domain Universum hard negative samples.
[0109] The enhanced sample set and the dual-domain Universum hard negative samples are simultaneously input into the improved TimesURL self-supervised feature learning network to perform contrastive learning to obtain low-dimensional embedding vectors for API endpoints.
[0110] In this embodiment, the enhanced sample set and dual-domain Universum hard negative samples are simultaneously input into the improved TimesURL self-supervised feature learning network, including:
[0111] The instance-level perturbation tensor matrix, sequence-level perturbation tensor matrix, and neighborhood feature perturbation tensor matrix corresponding to the target node in the temporal domain Universum hard negative sample set are respectively input into the encoder of the TimesURL self-supervised feature learning network to obtain the instance-dimensional temporal domain Universum hard negative sample embedding vector, the sequence-dimensional temporal domain Universum hard negative sample embedding vector, and the neighborhood feature-dimensional temporal domain Universum hard negative sample embedding vector.
[0112] Extract the temporal domain Universum hard negative sample embedding vector corresponding to the neighborhood node set into the topological domain Universum hard negative sample embedding vector.
[0113] The hard negative sample embedding vector of the topological domain Universum corresponds one-to-one with the neighboring nodes of the target node formed in the pseudo-topological graph through the insertion of pseudo-legal edges or the deletion of abnormal edges.
[0114] Positive sample pairs are formed by the embedding vectors of the target node in the frequency domain enhancement perspective and the embedding vectors of the target node in the time domain enhancement perspective. Negative sample sets are formed by the instance-dimensional time domain Universum hard negative sample embedding vectors, the sequence-dimensional time domain Universum hard negative sample embedding vectors, the neighborhood feature-dimensional time domain Universum hard negative sample embedding vectors, and the topology domain Universum hard negative sample embedding vectors. A contrastive learning loss function is constructed.
[0115] In Example 1, the contrastive learning loss function calculates the similarity value between positive sample pairs and compares the similarity value with the normalized similarity value between the positive sample pair and the embedding vector of each negative sample in the negative sample set to obtain the contrastive learning loss value of the target node.
[0116] The contrastive learning loss function is used to calculate the corresponding contrastive learning loss value for each target node in the node set, and the contrastive learning loss values of all target nodes are summed to obtain the batch loss value.
[0117] Batch loss values are used to measure the ability of the current set of network weight parameters to distinguish between legitimate business boundaries and dual-domain Universum hard negative sample boundaries in the shadow API governance scenario.
[0118] Based on the batch loss value, the set of network weight parameters of the TimesURL self-supervised feature learning network is updated by gradient, resulting in the updated set of network weight parameters.
[0119] In Example 1, gradient update is achieved by taking the partial derivative of the batch loss value with respect to the set of network weight parameters, scaling the gradient with the learning rate, and then subtracting it from the current set of network weight parameters to obtain the updated set of network weight parameters.
[0120] After the network weight parameter set is updated, the updated network weight parameter set is used to perform forward inference on the modulated and enhanced sample corresponding to any API endpoint in the node set to obtain the low-dimensional embedding vector of the API endpoint.
[0121] The low-dimensional embedding vector of an API endpoint is a real vector with the embedding dimension K. Each API endpoint corresponds to a unique low-dimensional embedding vector in the node set. The embedding dimension K represents the number of dimensions of the low-dimensional representation space used to represent the boundary of the shadow API business load.
[0122] Write back the low-dimensional embedding vector of the API endpoint to the initial API global spatiotemporal heterogeneous graph as a node attribute to obtain the embedding-enhanced API global spatiotemporal heterogeneous graph, and output the fused node feature vector.
[0123] In this embodiment, writing back the low-dimensional embedding vector of the API endpoint to the initial API global spatiotemporal heterogeneous graph as a node attribute includes:
[0124] Write back the low-dimensional embedding vector of the API endpoint to the node attributes of the topological graph structure to obtain the global spatiotemporal heterogeneous graph of the embedding-enhanced API.
[0125] In Example 1, the global spatiotemporal heterogeneous graph of the embedded enhanced API is composed of a set of nodes, a set of edges, and a set of embedded nodes. The set of nodes includes API endpoints, calling source IP entities, authentication identity entities, and microservice module instances. The set of edges includes calling relationships, parameter passing relationships, and permission links.
[0126] An anomaly function is constructed on the global spatiotemporal heterogeneous graph of embedded enhanced APIs to represent the anomaly degree of the low-dimensional embedding vector of API endpoints, and the anomaly degree is calculated based on the anomaly function for the low-dimensional embedding vector of any API endpoint in the node set.
[0127] Anomaly score is calculated as follows: Determine the community set to which the API endpoint belongs, calculate the arithmetic mean of the low-dimensional embedding vectors of all API endpoints in the community set, obtain the community embedding center vector of the community set, calculate the cosine similarity value between the low-dimensional embedding vector of the API endpoint and the community embedding center vector, and subtract the cosine similarity value from the numerical value to obtain the anomaly score of the corresponding API endpoint. The larger the anomaly score, the greater the deviation of the low-dimensional embedding vector of the corresponding API endpoint from the group embedding structure of its community set.
[0128] The community embedding center vector is obtained by summing the element-wise low-dimensional embedding vectors of all API endpoints in the community set and dividing by the number of nodes in the community set.
[0129] The neighborhood gating coefficient is obtained by multiplying the anomaly by the gating decay coefficient and then taking the negative exponential function.
[0130] The neighborhood gating coefficient, as a neighborhood aggregation weight in the message passing mechanism of the graph neural network, participates in the calculation of the latent features of the node. The gating decay coefficient is a preset positive number, which is used to control the degree of decay of the anomaly degree on the intensity of neighborhood information propagation.
[0131] The neighborhood gating coefficient is a dimensionless quantity between zero and one. The greater the anomaly, the smaller the neighborhood gating coefficient. It is used to implement attenuation control on the neighborhood information contribution of API endpoints with high anomaly in the shadow API governance scenario.
[0132] Based on the global spatiotemporal heterogeneous graph of the embedded enhancement API, the first-order neighboring node set of any node is determined, and the latent feature vector of the node is calculated by combining the neighborhood gating coefficient.
[0133] In Example 1, the low-dimensional embedding vector of the API endpoint is used as the initial node latent feature vector. The node latent feature vector of each neighboring node in the first-order neighboring node set is multiplied by the corresponding neighborhood gating coefficient. The weighted neighboring node latent feature vector is summed with the current latent feature vector of the target node. The summation result is multiplied by the learnable weight matrix of the current layer. A nonlinear activation function is applied to the matrix multiplication result to obtain the node latent feature vector of the next layer.
[0134] Construct a fused node feature vector by using the node latent feature vectors output at a preset number of layers;
[0135] In Example 1, the fused node feature vector is equal to the node latent feature vector after neighborhood aggregation of a preset number of layers. The fused node feature vector and the API endpoint low-dimensional embedding vector have the same embedding dimension K, and simultaneously contain temporal feature consistency information, spatial connectivity consistency information and topological integrity information.
[0136] A spatiotemporal anomaly potential energy function is constructed based on the feature vector of the fusion node. The potential energy value is calculated for all API endpoints in the global spatiotemporal heterogeneous graph of the embedded enhanced API. API endpoints with potential energy values greater than a preset threshold are marked as shadow API candidate nodes.
[0137] In this embodiment, a spatiotemporal anomaly potential energy function is constructed based on the feature vectors of the fused nodes, and the potential energy value is calculated for all API endpoints within the global spatiotemporal heterogeneous graph of the embedded enhanced API, including:
[0138] Based on the feature vectors of the fused nodes, calculate the business baseline deviation of the nodes corresponding to the API endpoints;
[0139] In Example 1, the cosine similarity value between the feature vector of the fusion node and the community embedding center vector is calculated. The cosine similarity value is subtracted from the numerical value to obtain the business baseline deviation, which is used to represent the degree of deviation of the API endpoint from the group business baseline embedding structure of its community set in the shadow API governance scenario.
[0140] Based on the embedded enhanced API global spatiotemporal heterogeneous graph, the first-order neighbor node set of the node corresponding to the API endpoint is determined, and the neighborhood consistency violation degree is calculated based on the fused node feature vector.
[0141] In Example 1, for each neighboring node in the first-order neighboring node set, the cosine similarity value between the fused node feature vector of the target API endpoint and the fused node feature vector of the neighboring node is calculated. Each cosine similarity value is subtracted from the first value to obtain multiple consistency deviation values. The multiple consistency deviation values are summed and divided by the number of nodes in the first-order neighboring node set to obtain the neighborhood consistency violation degree, which is used to represent the degree of consistency violation between the API endpoint and its call chain or permission chain neighborhood in the fused node feature vector space.
[0142] Based on the normalized node topological centrality and the node community legitimacy confidence, calculate the spoofing consistency of the nodes corresponding to the API endpoints;
[0143] In Example 1, the business baseline deviation is subtracted from the value to obtain the business consistency value. The business consistency value is multiplied by the topology support missing degree and then multiplied by the result of subtracting the node community legitimacy confidence from the value. The result is then multiplied by the result of subtracting the normalized node topology centrality from the value to obtain the fake consistency degree. This degree is used to represent the coupling strength of API endpoints in the shadow API governance scenario, which are disguised as community group embedding structures with temporal performance, but exhibit a discontinuity in topology and community legitimacy support.
[0144] The spatiotemporal anomaly potential energy function is obtained by weighted summation based on the deviation from the business baseline, the degree of neighborhood consistency violation, and the degree of spoofing consistency, and the potential energy value is calculated.
[0145] Calculate the potential energy value for each node corresponding to all API endpoints in the global spatiotemporal heterogeneous graph of the embedded enhanced API, and mark the nodes corresponding to API endpoints with potential energy values greater than a preset threshold as candidate nodes for shadow APIs.
[0146] Dynamic graph difference analysis is performed on the embedded enhanced API global spatiotemporal heterogeneous graph of candidate nodes in multi-time slices to screen the target nodes of shadow API that meet the dynamic abnormal evolution conditions and map them to the security governance strategy engine. Access control policies are generated according to preset blocking rules and distributed to the network gateway to implement automated circuit breaking, rate limiting or isolation governance of target nodes of shadow API.
[0147] In this embodiment, dynamic graph difference analysis is performed on the global spatiotemporal heterogeneous graph of the embedded enhancement API in multi-time slices of the shadow API candidate nodes, including:
[0148] The global spatiotemporal heterogeneous graph of the embedded enhancement API is divided sequentially according to a fixed time window to obtain multiple time window corresponding to the slice sequence of the global spatiotemporal heterogeneous graph of the embedded enhancement API.
[0149] For any shadow API candidate node in the shadow API candidate node set, calculate the island density index in the global spatiotemporal heterogeneous graph slice sequence of the embedding enhancement API corresponding to multiple time windows;
[0150] In Example 1, within a certain time window, the number of incoming edges pointing to the candidate nodes of the shadow API is counted, and the number of outgoing edges issued by the candidate nodes of the shadow API is counted. The number of incoming edges is divided by the number of outgoing edges plus one to obtain the island density index within the time window. The island density index is used to represent the call support density of the candidate nodes of the shadow API in the shadow API governance scenario.
[0151] The island density change rate is obtained by subtracting the island density index in the previous time window from the island density index in the current time window. When the island density change rate is less than zero and the island density change rate is continuously less than the preset island density change threshold in several consecutive time windows, it is determined to be an island density collapse trend.
[0152] The call frequency growth rate is obtained by subtracting the call frequency index in the previous time window from the call frequency index in the current time window. When the call frequency growth rate is greater than the preset call frequency growth threshold, and the call frequency growth rate continues to be greater than the preset call frequency growth threshold in several consecutive time windows, it is determined to be a slow penetration growth trend.
[0153] Candidate nodes that simultaneously satisfy the island density collapse trend and slow penetration growth trend are selected as the target node set for the Shadow API.
[0154] The target node set of the shadow API is a set of all candidate nodes that simultaneously satisfy the island density collapse trend and the slow penetration growth trend in the global spatiotemporal heterogeneous graph slice sequence of the embedded enhancement API corresponding to multiple time windows.
[0155] Map the topology location, permission link risk, and potential data access scope of any shadow API target node in the shadow API target node set to the security governance policy engine;
[0156] The topological location includes the community index of the shadow API target node in the embedded enhanced API global spatiotemporal heterogeneous graph and the set of first-order neighboring nodes; the permission link risk is the probability of violation of the permission link associated with the target node; the potential data access range is the set of all downstream nodes reachable through the edge set of the shadow API target node.
[0157] The security governance strategy engine generates access control policies based on preset blocking rules and distributes these policies to network gateways to implement automated circuit breaking, rate limiting, or isolation governance.
[0158] In this embodiment, the preset blocking rules include:
[0159] Access circuit breaker rules: When the risk of the access control link is greater than the preset risk threshold and the duration of the island density collapse trend is greater than the preset time threshold, an automated circuit breaker strategy will be executed.
[0160] Access rate limiting rules: When the risk of the access control link is less than the preset risk threshold and the duration of the slow penetration growth trend is greater than the preset time threshold, the rate limiting policy will be executed.
[0161] Access Isolation Rule: When the topological location of the target node of the Shadow API belongs to a community set with low community legitimacy confidence and the potential data access scope includes core data nodes, an isolation governance strategy is executed.
[0162] Example 2:
[0163] In the production environment of a large enterprise-level microservice platform, the API gateway processes tens of millions of requests daily. The system contains more than 3,000 API endpoints that are exposed to the outside or called internally. Due to frequent iterations of historical versions, legacy test interfaces, and incomplete recovery of gray-scale environments, the platform has gradually developed several API endpoints that are not registered in the interface asset list, but are still accessible at the log level.
[0164] Within a continuous running cycle, the implementer performs bypass collection of all API request and response streams. During the collection cycle, approximately 8,420,000 API request records, approximately 1,260 request source IPs, and approximately 940 authenticated entities were acquired. By analyzing features such as URL path characteristics, Header parameter entropy, Query parameter change rate, JSON Payload nesting depth, and response time, historical multivariate time series tensor matrices corresponding to 3,214 API endpoints were constructed. Each tensor matrix has 1,200 time steps and 15 feature dimensions.
[0165] In the initial stage of constructing the global spatiotemporal heterogeneous graph of APIs, 3214 API endpoints, 1260 source IP entities, 940 authentication entities, and 210 microservice module instances were abstracted into node sets, with approximately 98,000 call relationship edges, approximately 46,000 parameter passing relationship edges, and approximately 22,000 permission link edges. After processing by a community partitioning algorithm, the node sets were divided into 27 community sets. One community set contains 118 API endpoints with a community legitimacy confidence score of 0.93; another community set contains 46 API endpoints with a community legitimacy confidence score of 0.41.
[0166] During the TimesURL self-supervised feature learning phase, the system is trained on a modulated and enhanced sample set. A total of 9642 sets of hard negative samples for the time-series domain Universum are generated from the instance dimension, sequence dimension, and neighborhood feature dimension. In the topology domain, 3180 edges with an authentication consistency indicator of 0 are deleted, and 2406 pseudo-legitimate edges are inserted into the community set with a legality confidence level higher than the threshold to construct a pseudo-topology graph.
[0167] During training, the embedding dimension was set to 128, the batch size was 256, and a total of 100 training rounds were conducted. Compared with traditional self-supervised models that construct negative samples based solely on random masks, the negative samples constructed by the method of this invention have an average cosine similarity of 0.58 with the positive samples, while the traditional method has an average cosine similarity of 0.21, indicating that the negative samples constructed by this invention are closer to the legal business boundary.
[0168] During the embedding generation stage, the cosine similarity between the frequency domain enhanced perspective embedding vector and the time domain enhanced perspective embedding vector of a certain API endpoint A is 0.92, indicating that its business baseline is stable. However, the cosine similarity between the community embedding center vector of the community where API endpoint A is located and its embedding vector is 0.61, resulting in a business baseline deviation of 0.39. At the same time, its normalized node topology centrality is 0.07, and the node community legitimacy confidence is 0.38.
[0169] When calculating anomaly degree on the global spatiotemporal heterogeneous graph of the embedded enhancement API, the system counted 4 first-order neighbor nodes. Among them, the cosine similarity between the fused node feature vector of 3 neighbor nodes and the target node was less than 0.55, and the neighborhood consistency violation degree was 0.44. Further calculation showed that its topological support missing degree was 0.62.
[0170] When constructing the masquerade consistency score, the business consistency value of 0.61 is multiplied by the topology support deficiency score of 0.62, then multiplied by (1 - node community legitimacy confidence) (0.62), and finally multiplied by (1 - normalized node topology centrality) (0.93), resulting in a masquerade consistency score of approximately 0.22. The business baseline deviation, neighborhood consistency violation, and masquerade consistency score are then weighted and summed with weighting coefficients of 0.4, 0.3, and 0.3, respectively, yielding a spatiotemporal anomaly potential value of 0.47.
[0171] Of the 3214 API endpoints, 41 had spatiotemporal anomaly potential values greater than the preset threshold of 0.45, and were marked as candidate nodes for shadow APIs. Among them, candidate node B had a business baseline deviation of only 0.05 and a neighborhood consistency violation of 0.12, but a masquerade consistency of 0.51 and a comprehensive potential value of 0.49. This indicates that it is similar to the community group in terms of temporal performance, but there is a significant gap in topological support and community legitimacy.
[0172] Multi-time-slice dynamic graph differential analysis was performed, dividing the global spatiotemporal heterogeneous graph of the embedded enhanced API into 30 slices according to time windows. The island density index was calculated for candidate node B, which was 0.76 in the first time window, decreasing to 0.48 in the 15th time window, and further decreasing to 0.29 in the 30th time window, with the island density change rate remaining below -0.01 for 20 consecutive time windows. Simultaneously, its call frequency was analyzed, increasing from 28 times per day in the initial time window to 312 times per day in the final time window, with a growth rate consistently exceeding 8%.
[0173] The system determined that candidate node B simultaneously satisfies both the island density collapse trend and the slow penetration growth trend, and therefore selected it as a target node for the shadow API. The node's topological location indicates that it has not formed a stable connection with the high-legitimacy community, and its permission chain contains two edges that do not meet authentication consistency verification. Its potential data access scope includes two core data service nodes.
[0174] The security governance strategy engine generates access control policies based on preset blocking rules. When the risk score of the access control link is higher than 0.7 and the slow penetration growth lasts for more than 10 time windows, the access circuit breaker policy is executed. The system automatically sends blocking rules to the network gateway to restrict all external access requests to the API endpoint. After the policy takes effect, monitoring data shows that the number of abnormal access requests to the API endpoint has decreased from an average of 312 per day to 0.
[0175] In the comparative experiment, the traditional method based on single-point temporal anomaly detection, which only detects based on call frequency and response latency, identified only 18 candidate nodes on the same dataset, of which 12 were genuine shadow APIs and 6 were false positives. In contrast, the method of this invention identified 41 candidate nodes, of which 37 were genuine shadow APIs and 4 were false positives. The traditional method failed to identify candidate node B, while the method of this invention successfully identified it.
[0176] Furthermore, in the evaluation of embedding space separation, the mean inter-class distance of the traditional method is 0.34, while that of the method of this invention is 0.68. The convergence time of the traditional method after 100 training rounds is 2.8 hours, while that of the method of this invention is 3.1 hours. However, the method of this invention is significantly better than the traditional method in terms of anomaly detection accuracy.
[0177] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A shadow API governance method based on self-supervised comparison and graph reasoning, characterized in that, include: By capturing API request and response traffic through gateway-side bypass, a multivariate time series tensor matrix is constructed and an initial API global spatiotemporal heterogeneous graph is generated. Node topological centrality and community legitimacy confidence are calculated on the initial API global spatiotemporal heterogeneous graph. The node topological centrality and community legitimacy confidence are used as modulation factors and input into the TimesURL self-supervised feature learning network. Graph structure-guided frequency domain enhancement and time domain enhancement are performed on the mounted historical multivariate time series tensor matrix, and the modulated enhanced sample set is output. Dual-domain Universum hard negative samples are generated jointly from the instance dimension, sequence dimension, and neighborhood feature dimension. The enhanced sample set and the dual-domain Universum hard negative samples are simultaneously input into the TimesURL self-supervised feature learning network to perform contrastive learning to obtain the low-dimensional embedding vector of the API endpoint. Write back the low-dimensional embedding vector of the API endpoint to the initial API global spatiotemporal heterogeneous graph as a node attribute to obtain the embedding-enhanced API global spatiotemporal heterogeneous graph, and output the fused node feature vector. A spatiotemporal anomaly potential energy function is constructed based on the feature vector of the fusion node. The potential energy value is calculated for all API endpoints in the global spatiotemporal heterogeneous graph of the embedded enhanced API. API endpoints with potential energy values greater than a preset threshold are marked as shadow API candidate nodes. Dynamic graph difference analysis is performed on the embedded enhanced API global spatiotemporal heterogeneous graph of candidate nodes in multi-time slices to screen the target nodes of shadow API that meet the dynamic abnormal evolution conditions and map them to the security governance strategy engine. Access control policies are generated according to preset blocking rules and distributed to the network gateway to implement automated circuit breaking, rate limiting or isolation governance of target nodes of shadow API.
2. The shadow API governance method based on self-supervised comparison and graph reasoning according to claim 1, characterized in that, The construction of the multivariate time series tensor matrix and the generation of the initial API global spatiotemporal heterogeneous graph include: On the network gateway side, lossless bypass capture is performed on real-time API request and response streams. According to the unified protocol parsing rules, the URL path characteristics, Header parameter entropy, Query parameter change rate, JSONPayload nesting depth, and timestamps of the real-time API request and response streams are encoded into multivariate time series tensor matrices. A corresponding historical multivariate time series tensor matrix is established for each API endpoint. Based on the historical multivariate time series tensor matrix, node abstraction is performed on API endpoints, calling source IP entities, authentication identity entities, and microservice module instances. Edge abstraction is performed based on call relationships, parameter passing relationships, and permission links to construct an initial global spatiotemporal heterogeneous graph of APIs. The historical multivariate time series tensor matrix is mounted as the original feature of the nodes on the initial global spatiotemporal heterogeneous graph of APIs.
3. The shadow API governance method based on self-supervised comparison and graph reasoning according to claim 1, characterized in that, The output modulated enhanced sample set includes: A topological graph structure is constructed based on the initial API global spatiotemporal heterogeneous graph. The topological graph structure consists of a set of nodes and a set of edges. Calculate the topological centrality of any node in the node set, and calculate the normalized topological centrality. Divide the node set into several community sets, and calculate the community legitimacy confidence score for any community set; For any node in the node set, determine the community set to which the node belongs, and assign the community legitimacy confidence of the corresponding community set to the node community legitimacy confidence of the corresponding node. The normalized node topological centrality and the node community legitimacy confidence are weighted and summed to obtain the modulation factor. The enhancement strength coefficient is obtained by multiplying the value obtained by subtracting the modulation factor by a preset enhancement reference strength coefficient, and then calculating with the historical multivariate time series tensor matrix to obtain the frequency domain enhanced tensor matrix. The enhancement intensity coefficients are input into the TimesURL self-supervised feature learning network, and graph-structure-guided temporal enhancement is performed to obtain the temporally enhanced tensor matrix. The frequency-domain enhanced tensor matrix and the time-domain enhanced tensor matrix are aggregated as modulated enhanced samples of the same historical multivariate time series tensor matrix to obtain a set of modulated enhanced samples.
4. The shadow API governance method based on self-supervised comparison and graph reasoning according to claim 3, characterized in that, The generation of dual-domain Universum hard negative samples includes: When generating time-series Universum hard negative samples at the instance dimension, boundary approximation interpolation coefficients are obtained by linearly mapping the modulation factor between the lower bound and the upper bound of the preset interpolation coefficients. The frequency-domain enhanced tensor matrix corresponding to the target node and the time-domain enhanced tensor matrix corresponding to the reference node are linearly combined element-wise according to the boundary approximation interpolation coefficients to obtain the instance-level perturbation tensor matrix. When generating temporal domain Universum hard negative samples in the sequence dimension, the sequence-level perturbation segmentation index is calculated based on the modulation factor; The time segment of the time domain-enhanced tensor matrix corresponding to the target node before the sequence-level perturbation segmentation index is retained, and the time segment of the frequency domain-enhanced tensor matrix corresponding to the reference node after the sequence-level perturbation segmentation index is concatenated to the time segment of the target node to obtain the sequence-level perturbation tensor matrix. When generating temporal domain Universum hard negative samples in the neighborhood feature dimension, the first-order neighborhood node set of the target node is determined based on the initial API global spatiotemporal heterogeneous graph, and the neighborhood node is selected. The frequency domain enhanced tensor matrix corresponding to the target node and the frequency domain enhanced tensor matrix corresponding to the neighborhood node are linearly combined element by element according to the preset neighborhood feature perturbation interpolation coefficients to obtain the neighborhood feature perturbation tensor matrix. The instance-level perturbation tensor matrix, the sequence-level perturbation tensor matrix, and the neighborhood feature perturbation tensor matrix are converged to obtain the time-series domain Universum hard negative sample set; Based on the initial API global spatiotemporal heterogeneous graph, a pseudo-topological graph is constructed. The pseudo-topological graph is used as a set of hard negative samples in the topological domain Universum and hard negative samples in the temporal domain to generate dual-domain Universum hard negative samples.
5. The shadow API governance method based on self-supervised comparison and graph reasoning according to claim 4, characterized in that, The execution of contrastive learning to obtain low-dimensional embedding vectors for API endpoints includes: The instance-level perturbation tensor matrix, sequence-level perturbation tensor matrix, and neighborhood feature perturbation tensor matrix corresponding to the target node in the temporal domain Universum hard negative sample set are respectively input into the encoder of the TimesURL self-supervised feature learning network to obtain the instance-dimensional temporal domain Universum hard negative sample embedding vector, the sequence-dimensional temporal domain Universum hard negative sample embedding vector, and the neighborhood feature-dimensional temporal domain Universum hard negative sample embedding vector. Extract the temporal domain Universum hard negative sample embedding vector corresponding to the neighborhood node set into the topological domain Universum hard negative sample embedding vector. Positive sample pairs are formed by the embedding vectors of the target node in the frequency domain enhancement perspective and the embedding vectors of the target node in the time domain enhancement perspective. Negative sample sets are formed by the instance-dimensional time domain Universum hard negative sample embedding vectors, the sequence-dimensional time domain Universum hard negative sample embedding vectors, the neighborhood feature-dimensional time domain Universum hard negative sample embedding vectors, and the topology domain Universum hard negative sample embedding vectors. A contrastive learning loss function is constructed. The contrastive learning loss function is used to calculate the corresponding contrastive learning loss value for each target node in the node set, and the contrastive learning loss values of all target nodes are summed to obtain the batch loss value. Based on the batch loss value, the network weight parameter set of the TimesURL self-supervised feature learning network is updated by gradient to obtain the updated network weight parameter set. After the network weight parameter set is updated, the updated network weight parameter set is used to perform forward inference on the modulated and enhanced sample corresponding to any API endpoint in the node set to obtain the low-dimensional embedding vector of the API endpoint.
6. The shadow API governance method based on self-supervised comparison and graph reasoning according to claim 1, characterized in that, The step of writing back the low-dimensional embedding vector of the API endpoint to the initial API global spatiotemporal heterogeneous graph as a node attribute includes: Write back the low-dimensional embedding vector of the API endpoint to the node attributes of the topological graph structure to obtain the global spatiotemporal heterogeneous graph of the embedding-enhanced API. An anomaly function is constructed on the global spatiotemporal heterogeneous graph of embedded enhanced APIs to represent the anomaly degree of the low-dimensional embedding vector of API endpoints, and the anomaly degree is calculated based on the anomaly function for the low-dimensional embedding vector of any API endpoint in the node set. The neighborhood gating coefficient is obtained by multiplying the anomaly by the gating decay coefficient and then taking the negative exponential function. Based on the global spatiotemporal heterogeneous graph of the embedded enhancement API, the first-order neighboring node set of any node is determined, and the latent feature vector of the node is calculated by combining the neighborhood gating coefficient. By using the latent feature vectors of nodes output at a preset number of layers, a fused node feature vector is constructed.
7. The shadow API governance method based on self-supervised comparison and graph reasoning according to claim 1, characterized in that, The step of constructing a spatiotemporal anomaly potential energy function based on the feature vectors of fused nodes, and calculating the potential energy value for all API endpoints within the global spatiotemporal heterogeneous graph of the embedded enhanced API, includes: Based on the feature vectors of the fused nodes, calculate the business baseline deviation of the nodes corresponding to the API endpoints; Based on the embedded enhanced API global spatiotemporal heterogeneous graph, the first-order neighbor node set of the node corresponding to the API endpoint is determined, and the neighborhood consistency violation degree is calculated based on the fused node feature vector. Based on the normalized node topological centrality and the node community legitimacy confidence, calculate the spoofing consistency of the nodes corresponding to the API endpoints; The spatiotemporal anomaly potential energy function is obtained by weighted summation based on the deviation from the business baseline, the degree of neighborhood consistency violation, and the degree of spoofing consistency, and the potential energy value is calculated. Calculate the potential energy value for each node corresponding to all API endpoints in the global spatiotemporal heterogeneous graph of the embedded enhanced API, and mark the nodes corresponding to API endpoints with potential energy values greater than a preset threshold as candidate nodes for shadow APIs.
8. The shadow API governance method based on self-supervised comparison and graph reasoning according to claim 1, characterized in that, The dynamic graph difference analysis performed on the embedded augmented API global spatiotemporal heterogeneous graph of the shadow API candidate nodes in multi-time slices includes: The global spatiotemporal heterogeneous graph of the embedded enhancement API is divided sequentially according to a fixed time window to obtain multiple time window corresponding to the slice sequence of the global spatiotemporal heterogeneous graph of the embedded enhancement API. For any shadow API candidate node in the shadow API candidate node set, calculate the island density index in the global spatiotemporal heterogeneous graph slice sequence of the embedding enhancement API corresponding to multiple time windows; The island density change rate is obtained by subtracting the island density index in the previous time window from the island density index in the current time window. When the island density change rate is less than zero and the island density change rate is continuously less than the preset island density change threshold in several consecutive time windows, it is determined to be an island density collapse trend. The call frequency growth rate is obtained by subtracting the call frequency index in the previous time window from the call frequency index in the current time window. When the call frequency growth rate is greater than the preset call frequency growth threshold, and the call frequency growth rate continues to be greater than the preset call frequency growth threshold in several consecutive time windows, it is determined to be a slow penetration growth trend. Candidate nodes that simultaneously satisfy the island density collapse trend and slow penetration growth trend are selected as the target node set for the Shadow API. Map the topology location, permission link risk, and potential data access scope of any shadow API target node in the shadow API target node set to the security governance policy engine; The security governance strategy engine generates access control policies based on preset blocking rules and distributes these policies to network gateways to implement automated circuit breaking, rate limiting, or isolation governance.
9. The shadow API governance method based on self-supervised comparison and graph reasoning according to claim 8, characterized in that, The preset blocking rules include: Access circuit breaker rules: When the risk of the access control link is greater than the preset risk threshold and the duration of the island density collapse trend is greater than the preset time threshold, an automated circuit breaker strategy will be executed. Access rate limiting rules: When the risk of the access control link is less than the preset risk threshold and the duration of the slow penetration growth trend is greater than the preset time threshold, the rate limiting policy will be implemented. Access Isolation Rule: When the community legitimacy confidence of the community set to which the target node of the shadow API belongs is lower than the preset community legitimacy threshold, and the potential data access scope includes core data nodes, an isolation governance strategy is executed.
Citation Information
Patent Citations
Petrochemical production process anomaly diagnosis and optimization method and system integrated with knowledge graph
CN119668245A
Digital power grid asset classification and evolution monitoring method, system and device based on self-supervised comparative learning and storage medium
CN120951021A