A node zero-trust trusted access method for a computing power network
By introducing the zero-trust principle and a multi-step dependency anomaly detection model into the computing power network, and combining multi-dimensional feature fusion and time decay mechanism, the real-time and accuracy problems of node behavior evaluation in the computing power network are solved, and the security and reliability of the system are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies lack dynamic monitoring and real-time evaluation of the long-term behavior of nodes in computing networks, making it difficult to capture security changes in a timely manner and posing potential security risks. Furthermore, traditional log anomaly detection methods cannot effectively capture potential behavioral anomalies in complex and ever-changing computing network environments, resulting in inaccurate evaluation results.
Based on the zero-trust principle, a trust assessment is performed during the node registration and task execution phases. Semantic features, temporal features, parameter features, and quantitative features are extracted and weighted feature fusion is performed using an attention mechanism. This constructs a multi-step dependency anomaly detection model, introduces time decay and penalty mechanisms for comprehensive scoring, and monitors node behavior in real time.
It significantly improves the security and adaptability of the computing network, enables accurate dynamic evaluation of node behavior, identifies abnormal nodes at an early stage, and improves the accuracy of anomaly detection and the robustness of the system.
Smart Images

Figure CN120750671B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network security technology, and in particular to a zero-trust trusted access method for nodes in computing power networks. Background Technology
[0002] With the booming development of various computing-intensive businesses and the large-scale application of AI in various industries, the demand for computing resources is increasing daily. Computing networks, as an innovative solution for resource integration and network collaboration, are gradually becoming an important component of future digital infrastructure. By efficiently connecting dispersed computing and storage resources, computing networks not only improve resource utilization efficiency but also provide flexible and optimized computing services for different business scenarios. Computing networks distribute resource information such as computing power, storage, and algorithms from service nodes through the network and, combined with network information, provide optimal resource allocation and network connection solutions for different user needs, thereby achieving optimal use of network resources. The rapid development of computing networks has also brought new security challenges. These challenges mainly stem from the new characteristics brought about by computing power interconnection and computing-network convergence, leading to an expanded attack surface; changes in network architecture and the introduction of new network element entities have also exacerbated security risks, increasing management complexity and operational risks. Security risks during node access and task execution phases are gradually becoming important issues affecting network stability and trustworthiness.
[0003] Existing research primarily focuses on security during the node access phase, typically ensuring the legitimacy of accessing nodes through authentication. However, this approach neglects the dynamic behavior of nodes during operation and its impact on overall system security. In the complex context of computing networks, a single assessment cannot fully reflect long-term behavioral changes as nodes continuously operate within the network. Therefore, this invention introduces the zero-trust principle, addressing not only security during access but also enabling continuous monitoring and real-time evaluation of node behavior. Through this mechanism, the computing network can promptly detect malicious nodes or abnormal behavior, improving overall security. The zero-trust architecture requires real-time, trusted evaluation of node behavior during task execution to prevent potential security threats from persisting within the network.
[0004] In computing networks, the trustworthiness of node behavior is not only a crucial indicator of network health but also a vital link in ensuring network security. Here, "behavior" refers to the operations and states exhibited by nodes during the execution of computing tasks, data processing, and resource requests. For example, a node's behavioral trustworthiness can be assessed through the stability of its computing task execution (e.g., whether tasks are successfully completed), resource usage (e.g., whether memory and CPU usage exceed expectations), response time (e.g., whether task requests are responded to promptly), and the presence of abnormal operations (e.g., unauthorized resource access or frequent task failures). With changes in task load, dynamic adjustments to network topology, and interference from the external environment, node behavior can change significantly, and a single assessment cannot capture these changes in a timely manner. To address this issue, this invention proposes a dynamic and continuous trusted node access scheme based on the zero-trust principle. Under the zero-trust architecture, it is no longer assumed that any node remains trustworthy after joining the network; instead, the behavior of each node is continuously monitored and evaluated in real time. This means that regardless of the node's access status, every interaction with the network and every task execution requires a trust assessment to ensure that the node always meets security requirements.
[0005] Logs generated during node operation record the node's detailed status and key events, serving as the primary basis for assessing the trustworthiness of node behavior. Traditional log anomaly detection methods identify potential node behavior problems by analyzing abnormal patterns in the logs. While these methods are effective to some extent, they still have limitations in comprehensively assessing the trustworthiness of node behavior. Log anomalies are only one part of evaluating node trustworthiness; more multi-dimensional behavioral characteristics (such as resource consumption, task execution quality, and user feedback) need to be incorporated into the trustworthiness assessment to achieve more comprehensive and accurate results.
[0006] In log anomaly detection, traditional methods rely on manual processing, which is clearly impractical given the ever-increasing volume of log data in computing network environments. With the rapid development of computing networks, the continuous expansion of system scale, and the widespread application of distributed technologies, the amount of logs generated by the system has increased dramatically. Simultaneously, the interleaving of log data and the complex dependencies between nodes make manual detection even more difficult. Therefore, traditional manual log analysis methods are no longer suitable for the needs of modern computing networks. To address this issue, researchers have proposed various automated log anomaly detection methods. Early log anomaly detection methods primarily extracted semantic features from logs and combined them with machine learning algorithms for anomaly detection. However, these methods have limitations in feature capture and struggle to effectively reveal deep-seated patterns in log data. With the rapid development of deep learning technology, deep learning-based log anomaly detection methods have gradually become mainstream. These methods can automatically extract more complex and hidden features from large-scale log data, significantly improving the accuracy and efficiency of anomaly detection. Nevertheless, log anomaly detection for computing network nodes still faces many challenges, especially when dealing with complex dependencies between nodes and massive amounts of log data; existing anomaly detection technologies still have considerable room for optimization.
[0007] In terms of trust assessment, especially in computing power network environments, the heterogeneity of tasks and the differences in resource configuration among nodes require not only relying on log data for anomaly detection, but also dynamically assessing the trustworthiness of node behavior. Only by taking into account the diversity of tasks, the stability of execution, and the influence of external factors can a more robust and efficient trust assessment be provided for computing power networks.
[0008] Therefore, it is necessary to design a zero-trust trusted access method for nodes in computing power networks to solve the technical challenge of conducting robust and efficient security assessments in computing power network scenarios. Summary of the Invention
[0009] This invention addresses the shortcomings of existing technologies by proposing a zero-trust trusted access method for nodes in computing power networks. Based on the zero-trust principle, this scheme performs trust assessments during both the registration and operation phases of computing power nodes. After obtaining an initial trust value upon registration, the behavior of computing power nodes is monitored in real time, and the trustworthiness of node behavior is comprehensively estimated. This enables early detection and handling of nodes exhibiting abnormal behavior, improving the overall security and efficiency of the computing power network. This scheme extracts features from different dimensions of computing power network logs by introducing semantic feature extraction, temporal feature extraction, parametric feature extraction, and quantitative feature extraction. Effective feature fusion is performed based on attention, and a multi-step dependency anomaly detection model is used to capture long-distance relationships, significantly improving the accuracy of node log anomaly detection in the context of computing power networks. Furthermore, the anomaly severity of logs over a period of time is comprehensively scored and integrated with user feedback. By assessing the trustworthiness of computing power nodes and implementing risk responses, insecurity factors in the computing power network can be addressed promptly, improving the overall security and availability of the system.
[0010] To achieve the above objectives, the present invention provides the following technical solution:
[0011] A zero-trust trusted access method for nodes in a computing power network includes the following steps:
[0012] Step S1, Trust Assessment During Computing Node Registration: In a computing power network, node access is the first step for the stable operation of the entire system. Idle computing power devices register as computing power nodes with the scheduling center. When a new computing power device applies to register and join the computing power network, the network needs to verify the node's identity and assess its capabilities to ensure the node's trustworthiness, prevent malicious nodes from intruding into the network, and obtain an initial trust value.
[0013] Step S2, Trust Assessment Based on Behavioral Logs During the Task Execution Phase of Computing Power Nodes: Information collection probes deployed upon node access are used to track connected nodes in real time, collecting operational data to form a dynamic data stream. The scheduling system periodically requests uploaded computing power node logs, which are then parsed using a Brain parser to obtain structured logs. Semantic features, temporal features, parameter features, and quantity features are extracted from the obtained log templates, and these four types of vectors are weighted and fused. Considering that node behavior in computing power networks is often influenced by multi-step time dependencies, hidden states from multiple historical time steps are explicitly introduced to enhance the detection capability of abnormal patterns. After anomaly detection for each log entry, a time decay model is used to dynamically update the node's trust score, and penalties are supplemented based on the magnitude of the outlier. When a node exhibits abnormal behavior, the scheduling system can quickly identify and isolate the node to prevent potential threats from spreading throughout the network and adjust the corresponding payment standards.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0015] This invention addresses the current state of computing networks as a next-generation information infrastructure, which has become the cornerstone of societal digital transformation and a key project for major companies. However, existing technologies largely focus on trust assessment during node access, lacking dynamic monitoring and real-time evaluation of long-term node behavior. This makes it difficult to promptly capture changes in node security during task execution, posing potential security vulnerabilities. Based on the principle of zero trust, this invention, while conducting trust assessment during the access phase, also focuses on the security risks of computing network nodes during operation, implementing precise risk control throughout the node's entire lifecycle, thereby significantly improving the security and adaptability of computing networks.
[0016] 2. Based on the unique context of computing power networks, this invention proposes improved feature extraction methods. Existing log analysis methods typically rely on simple feature extraction, primarily focusing on anomaly analysis of single log features. These methods often fail to effectively capture potential behavioral anomalies in the complex and ever-changing computing power network environment, leading to inaccurate evaluation results. This invention extracts key semantic, temporal, parametric, and quantitative features from computing power networks and combines them with an attention mechanism for weighted feature fusion. This results in a more comprehensive extraction of node behavioral features from computing power network logs, ensuring accurate capture of potential anomalies in logs within a complex and ever-changing environment, and providing higher accuracy and reliability for node behavior analysis.
[0017] 3. This invention constructs an anomaly detection model based on multi-step dependencies. Addressing the characteristic in log analysis of computing networks that anomalies are typically not directly determined by log events at a single time step, but rather influenced by long-term dependencies, this invention explicitly uses the hidden states of multiple historical time steps as input to each time unit. This effectively overcomes the shortcomings of traditional methods (GRU, LSTM, etc.) in fully capturing long-term dependencies. This method can handle the temporal characteristics of log data in computing networks, improving the sensitivity and accuracy of anomaly detection.
[0018] 4. This invention introduces a time decay and penalty mechanism to comprehensively score the degree of log anomalies over a period of time, and integrates this score with user feedback scores to comprehensively evaluate the trustworthiness of nodes, thereby achieving fine-grained control over node access permissions. This addresses the shortcomings of existing trust assessment methods that only score log anomalies at a single point in time, lacking a dynamic adjustment mechanism for changes in abnormal behavior over time series and ignoring anomaly changes over long periods. This invention can dynamically adjust the trust score of nodes and integrate it with user feedback scores, providing a more refined assessment of node behavior trustworthiness, significantly improving the security and attack resistance of computing networks, ensuring precise control of node access permissions, and preventing potentially malicious nodes from lurking for extended periods.
[0019] In summary, the node behavior trustworthiness assessment method proposed in this invention improves the accuracy of anomaly detection while ensuring the robustness of the system. Furthermore, through intelligent stage scoring, it identifies and controls abnormal nodes in the early stages, greatly enhancing the security of the computing network. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0021] Figure 1 This is a diagram illustrating the operational mechanism of a zero-trust trusted access method for nodes in a computing power network, as provided in an embodiment of the present invention.
[0022] Figure 2 This is a macroscopic architecture diagram of a zero-trust trusted access method for nodes in a computing power network, provided in an embodiment of the present invention.
[0023] Figure 3 This is a detailed process diagram of the node computing power registration stage in a zero-trust trusted access method for computing power networks provided in an embodiment of the present invention.
[0024] Figure 4 This is a detailed process diagram of behavior-based trust assessment during the node task execution phase in a zero-trust trusted access method for computing power networks provided in an embodiment of the present invention. Detailed Implementation
[0025] To better understand this technical solution, the method of the present invention will be described in detail below with reference to the accompanying drawings.
[0026] The node behavior reliability assessment method proposed in this invention for computing power networks is deployed in the scheduling system of the computing power network. It interacts with computing power nodes to collect node behavior information and make reliability assessments and feedback, such as... Figure 1 As shown, the computing network comprises three components: "computing," "network," and a "brain." "Computing" refers to... Figure 1 The rightmost computing power provider produces computing power; the "network" connects these computing power; and the "brain," the scheduling system shown in the diagram, is responsible for the unified perception, orchestration, scheduling, and coordination of the computing power within the network. The leftmost computing power consumer requests computing power from the scheduling system to obtain usage rights and then rates the computing power nodes after use. The trust assessment module within the scheduling system periodically evaluates the trustworthiness of the computing power nodes' behavior and responds to risks.
[0027] The node behavior trust assessment method for computing power networks proposed in this invention mainly consists of two modules: trust assessment during the computing power node registration phase and trust assessment based on behavior logs during the computing power node task execution phase. The access phase is divided into three steps: identity trust assessment, capability trust assessment, and initial trust value calculation. The task execution phase is divided into three steps: log collection and parsing, multi-dimensional log feature extraction and fusion, and node trust assessment based on log features. Figure 2 As shown, it specifically includes the following two modules and six steps:
[0028] S1 Trust Assessment during Computing Node Registration Phase
[0029] In a computing power network, node access is the first step towards stable system operation. Idle computing devices register as computing power nodes with the scheduling center. When a new computing power device applies to register and join the computing power network, the network needs to verify the node's identity, credibility, and capabilities to ensure the node's trustworthiness and prevent malicious nodes from intruding into the network.
[0030] S1.1 Trusted Identity Verification
[0031] In computing power networks, node access requires strict compliance with both identity legitimacy and computing power capability trustworthiness. This invention designs a phased evaluation method. First, multi-factor authentication verifies the node's legitimacy; if it fails, access is directly rejected, and subsequent capability evaluations are not performed. Identity trustworthiness is used to determine the authenticity and trustworthiness of the node's registered identity and hardware identifier. The platform first verifies the legitimacy of the digital certificate submitted by the node, checking the validity of its certificate chain and the trusted issuing authority, and verifies the node's private key possession through digital signature. Finally, the integrity of the identity is determined by combining the node's model, MAC address, and network protocol stack fingerprint consistency. Defined as a concise binary threshold judgment, it is defined as 1 when all values pass, and 0 otherwise.
[0032] S1.2 Capability Credibility Assessment
[0033] After successful authentication, the platform directly collects capability information using a lightweight probe program deployed on the nodes, avoiding resource consumption introduced by benchmark tasks. Capability credibility assessment consists of two parts: first, capability size assessment, used to generate capability tags; second, capability consistency assessment, used to determine the degree of matching between the node's declared resources and the actual measured resources. Capability size indicators include CPU computing performance, GPU acceleration capability, memory capacity, and network bandwidth. Let the capability indicators measured by the probe be... , , The total ability score can then be defined as a weighted sum:
[0034] (1),
[0035] in Weights for each dimension. Size classification ability standard (2),
[0036] in This is the threshold for high capability level; nodes exceeding this threshold have a high capability label. Nodes with a low capability level threshold are labeled as low capability, while nodes with capabilities between the high and medium capability levels are labeled as medium capability.
[0037] Capability consistency is assessed by comparing the relative differences between node declared values and probe measurements, where the declared value is... The consistency score is then: (3),
[0038] The closer this value is to 1, the higher the consistency. The final node capability reliability is defined as... This is used for subsequent organizational reputation correction and access trust score generation. If identity verification fails, the system directly rejects node access and capability data is not collected, thus forming an access strategy of verifying identity first, then capability, and in stages. This ensures security, reduces computing power network access overhead, and provides capability tags and consistency indicators to support task scheduling.
[0039] S1.3 Node Trust Modeling Based on Bayesian Prior and Graph Regularization
[0040] In a computing network environment, the trustworthiness of node access depends not only on the node's capability assessment results during the access phase, but also on its long-term reputation within its organization, and there is potential behavioral consistency among nodes within the same organization. To achieve integrated modeling of organizational priors, node observations, and group smoothness, this patent proposes a node trust inference method based on Bayesian priors and graph regularization to generate the globally optimal trust score during the node registration phase.
[0041] Suppose that the computing power network contains a set of organizations. With node set And define heterogeneous bipartite graph : (4),
[0042] in For each node n, there is a node-organization affiliation edge. For each node n, its affiliated organization reputation value R... o (n) represents the average trust value of the nodes leased by the organization. Node capability consistency score. The probes deployed in the previous section collect data in real time upon access, providing an objective evaluation of the computing power performance. To reasonably integrate these two types of information, this patent first treats node trust as a potential probability variable x. n ∈[0,1], and a prior Beta distribution is set based on organizational reputation. (5),
[0043] Where prior parameters , Obtained through organizational reputation mapping: (6),
[0044] κ represents the prior confidence strength. The organizational credibility of node n is identified. In the absence of historical reliable data, this distribution degenerates into a uniform distribution. Subsequently, the system corrects the node credibility to a posterior expectation based on capability observations. Due to the capability distribution collected by the probe... This can be viewed as a continuous observation probability. This patent simplifies the modeling, with each node's prior correction confidence score... for: (7),
[0045] in The introduction of this is equivalent to an observation event, whose confidence contribution is less than that of long-term organizational priors, in order to balance short-term capabilities with historical credibility.
[0046] After obtaining the prior corrected confidence score for each node. Subsequently, to further express the consistency of node scores within the same organization, this patent introduces smoothing regularization into the global graph structure and defines an objective function. (8),
[0047] in For smoothness weights, This represents the graph regularization term, used to constrain the trust values between adjacent nodes to be similar, ensuring trust consistency in the network structure. The objective function ensures that node trust is both close to the organization's prior knowledge and its own capabilities, while maintaining reasonable consistency within the same organization. This optimization problem is equivalent to solving a linear equation:
[0048] (9),
[0049] in For the Laplacian matrix of the node-node graph, To correct the expected vector, the closed-form solution is: (10)
[0050] By modeling node trust based on Bayesian priors and graph regularization, the node trust score not only has probabilistic prior meaning but also satisfies the smooth coupling relationship of the organization-node bipartite graph in the computing power network, which significantly improves the interpretability and global consistency of the score.
[0051] S1.4 Access Decision and Dynamic Evolution
[0052] After completing the trust inference during the node access period, the platform uses the calculated node trust score. Based on capability tags, hierarchical decision-making management is implemented for node access status. The node access status is determined by the following decision function:
[0053] (11),
[0054] in This is the access rejection threshold; nodes below this threshold are denied access. This is the full access threshold; nodes exceeding this threshold can achieve full access. , It will dynamically adjust based on the current network load, task urgency, and available computing resources: (12)
[0055] in This indicates the current computing power utilization rate (i.e., load ratio). As a static baseline threshold, This represents the maximum dynamic adjustment range.
[0056] For access status in the edge area (located in) and For nodes (between), the platform can adopt a progressive access strategy, that is, first grant them limited permissions, allow them to execute low-risk or low-priority tasks, and continuously monitor their behavior during runtime to decide whether to upgrade permissions.
[0057] After a node is connected to the computing network, the platform continuously collects its runtime behavior data (such as resource usage, task completion status, and anomaly logs) through deployed probe programs, and dynamically evolves the node's trust status based on the following mechanism: The system aggregates the trust status of the organization's subordinate nodes every 24 hours and updates the organization's reputation. (13)
[0058] The median is used to enhance robustness to outlier nodes, where This represents the organization's trust value at the k-th period, used to indicate the organization's overall reputation level in the current period, and its value range is usually [0,1]. The trust value of an organization in the next cycle (the (k+1)th cycle) is obtained through a dynamic evolution mechanism. This is a smoothing coefficient, ranging from [0, 1], which controls the weighting of historical confidence values and current observations. A larger value indicates a greater reliance on historical values and a more conservative update; a smaller value indicates a greater reliance on new observations and a more sensitive response. The runtime behavior score of a node belonging to the organization at time t. This represents the set of all nodes in the organization.
[0059] S2, Trust assessment based on behavior logs during the task execution phase of computing power nodes.
[0060] S2.1 Log Collection and Analysis
[0061] S2.1.1 Log Collection
[0062] During node registration in step S1, the scheduling system deploys a data collection probe to each registered node. This probe runs continuously on each connected computing node, performing comprehensive real-time monitoring and data collection. The scheduling system can collect various log data generated by nodes during computational tasks. This log data typically includes, but is not limited to: task execution status (recording the start, execution, and completion of computational tasks and related status information), resource usage (including the consumption of computing, storage, and bandwidth resources, reflecting the efficiency of node resource utilization), and abnormal behavior records (abnormal situations that may occur during node operation, such as task failures or resource overload).
[0063] The following specific requirements apply to the collected logs during the log collection process: First, the collected logs originate from various nodes within the computing network and are uploaded periodically according to requests from the scheduling system. Second, the log collection frequency must meet the real-time requirements of task execution on the nodes within the computing network. Generally, the log collection frequency is once per second to ensure the capture of subtle changes in node behavior. For high-frequency tasks, the log collection frequency may be further increased to reflect node status in real time. Finally, to ensure the accuracy and reliability of the log data, the collection module employs consistency and integrity verification mechanisms. All log data is encrypted and verified during transmission and storage to prevent data loss or tampering. Simultaneously, the collected data must guarantee the accuracy of key information such as timestamps, node IDs, and task IDs to ensure data traceability.
[0064] Both the real-world HDFS and BGL datasets can be used as the log collection results of this invention. The HDFS dataset was generated by Hadoop-based map-reduce jobs on more than 200 Amazon EC2 nodes, containing a total of 111,756,290 raw log messages. The BGL dataset is an open dataset of logs collected from the Blue Gene / L supercomputer system at Lawrence Livermore National Laboratory, containing 4,747,963 raw log messages.
[0065] S2.1.2 Brain-based log parsing
[0066] Data acquisition analyzes unstructured log data output from edge computing nodes, employing state-of-the-art Brain methods for structuring. Log parsing removes redundant information and extracts log event time and content information. Preprocessing primarily includes word segmentation and filtering of common variables. Initial groups are created by selecting the longest word combination from combinations of words with frequencies greater than a frequency threshold. For each initial group, a bidirectional parallel tree is used to add nodes, outputting a log template.
[0067] S2.2 Multi-dimensional log feature extraction and fusion
[0068] This step extracts semantic features, temporal features, parameter features, and quantity features from the log template, and then performs multi-dimensional feature fusion, including:
[0069] S2.2.1 Semantic Feature Extraction Based on ALBERT and Contrastive Learning
[0070] In computing networks, node logs typically contain complex semantic information, which is crucial for understanding node behavior patterns and anomaly detection. For example, logs containing direct entries like "exception" or "hrown" indicate that an anomaly occurred during a task. However, log data often presents challenges such as ambiguity and log variation.
[0071] Unlike general system logs, logs from computing power networks are characterized by multi-task complexity, semantic ambiguity, and log variations depending on task type. Traditional rule-based feature extraction methods struggle to effectively address these issues. To more comprehensively capture the deep semantic information of computing power network logs and improve the model's robustness to log mutations, this invention employs ALBERT as a pre-trained language model, combined with contrastive learning, to achieve efficient semantic embedding of log templates. This method effectively addresses challenges such as polysemous words, multi-task log variations, and unknown log templates, thereby improving the accuracy of anomaly detection.
[0072] Log template A is split into M tokens, which are used as input to ALBERT, where [CLS] marks the starting position of the log template. Each token generates a corresponding semantic feature vector. The initial semantic embedding is calculated by averaging the hidden vectors of the penultimate encoding layer. .
[0073] To further enhance the discriminative power of log templates, this invention introduces contrastive learning to optimize the semantic representation of ALBERT. Log templates are defined. The embedding is represented as The formula is as follows: (14)
[0074] The goal is to bring log templates with similar semantics (positive samples) closer together and log templates with unrelated semantics (negative samples) further apart, as shown in the formula below: (15) (16)
[0075] loss function of contrastive learning The definition is as follows: (17)
[0076] in It is a temperature parameter that controls the sensitivity of sample distribution. This represents a log template vector with a similar context. It is a negative sample of an irrelevant log template.
[0077] Therefore, comparative learning adjustment This representation makes similar logs closer together and logs of different categories further apart, ultimately resulting in semantic embedding. This serves as input for subsequent anomaly detection.
[0078] S2.2.2 Extraction of Time Features Based on the Rate of Change of Time Difference
[0079] In most logging systems of computing network nodes, each log entry contains a timestamp, which can be used to calculate the time interval between log entries. For a stable system, the execution time of the same program path usually remains stable. When system performance issues occur, faulty components may cause abnormal changes in execution time, resulting in a significant increase in log intervals. For example, suppose a normal computing node takes 1 second to complete a specified operation, but on a certain node it takes 1000 seconds to complete the same operation, which indicates that an anomaly may have occurred.
[0080] The absolute value of the time difference between logs may have different dimensions due to the different time units (such as milliseconds, seconds, minutes) of different computing power nodes and their different log systems. The relative rate of change, being a dimensionless indicator, is more universal. Therefore, this invention introduces the relative rate of change of time difference as a time feature for anomaly detection. Through softmax normalization and high-dimensional embedding, it can capture subtle differences in time interval changes, improving detection accuracy.
[0081] First, calculate the timestamps of adjacent log entries in the log sequence. and Time difference between And generate a time difference sequence.
[0082] in This represents the time interval between adjacent log entries. Next, the relative rate of change of adjacent time differences in the time difference series is calculated. The formula can be expressed as: (18)
[0083] Based on this, a time difference relative change rate sequence is generated. relative rate of change of time difference Embedding encoding is performed to generate an embedding vector of the relative rate of change of time difference, as follows.
[0084] First, Multiplying by the randomly initialized weight vector W and adding the randomly initialized bias vector b yields the mapped high-dimensional vector. The high-dimensional vector z is normalized using the softmax function, scaling its values to the range [0,1] while ensuring that the sum of each dimension is 1. The formula can be expressed as: (19)
[0085] For each in the sequence By performing the above operations sequentially, the embedded sequence of relative change rate of time difference is finally obtained. .
[0086] S2.2.3 Numerical Parameter Feature Extraction Based on Normalization Enhancement
[0087] The values of certain parameters in the logs are also crucial for anomaly detection. For example, in a computing network, resource consumption parameters such as memory usage, CPU utilization, and network bandwidth in log events are important indicators of whether node behavior is normal. When the parameter values of certain log events are abnormal, it may mean that the node has performance problems or configuration errors. For example, in a computing network, suppose there are two log events that use the same template: "Request computing resources". However, in the second event, the log parameters show that the node's memory usage significantly exceeds the normal range (e.g., memory utilization reaches over 90%). This abnormal resource usage usually leads to a decrease in node performance or task failure.
[0088] To address the characteristics of log parameters in computing network environments, this invention proposes a normalized enhanced numerical parameter feature extraction method to improve anomaly detection capabilities for computing network logs. In computing network logs, parameters typically include information such as task identifiers, computing nodes, resource requests, and storage paths. Although these parameters may exhibit significant numerical differences across different logs, they generally follow a certain distribution pattern during the normal execution of computing tasks. Therefore, this invention enhances the expressive power of numerical parameters in anomaly detection through dynamic normalization and hierarchical embedding.
[0089] Because the parameter values in the computing power network logs vary greatly, directly inputting the raw values may lead to false positives in anomaly detection. Therefore, we adopt a sliding window normalization strategy to normalize the parameters within a local window to preserve their relative trends. (20)
[0090] in, These are the original parameter values. and These are the mean and standard deviation within the sliding window, respectively. To prevent division by zero for minute numbers, this method effectively adapts to dynamic task load changes in computing networks, making the model more robust.
[0091] Based on the normalized parameters, this invention employs a multilayer perceptron (MLP) to extract features from the parameters, thereby capturing numerical features at different levels, as shown in the following formula: (twenty one),
[0092] in, , , , For trainable parameters, ReLU is used as a non-linear activation function, enabling the model to learn complex mapping relationships between parameter values. This hierarchical embedding enhances the ability to perceive abnormal patterns and improves detection accuracy. Furthermore, to further improve the anomaly detection capability of numerical parameters, this invention introduces a global distribution constraint loss during training to ensure that normal log parameters in the embedding space remain compact, while anomalous parameters can be effectively distinguished. By minimizing the embedding variance of normal log parameters and maximizing the Euclidean distance between anomalous log parameters and normal parameters, the model's sensitivity to anomalous parameters is enhanced.
[0093] S2.2.4 Optimizing Quantity Feature Extraction Based on Category-Aware Counting
[0094] In computing power networks, log template sequences not only contain time, semantic, and parameter information, but also rich quantitative patterns. In log analysis of computing power networks, anomalies in quantitative characteristics refer to deviations in the frequency of certain templates from the expected normal pattern. These anomalies often reflect abnormalities or malfunctions in the system's operational status. Too many or too few tasks being executed, or abnormal resource usage, may be caused by these reasons, making them crucial for identifying potential problems in computing power networks.
[0095] Existing quantitative embedding methods detect anomalies by counting the occurrences of log templates using a sliding window. However, in computing power networks, the heterogeneity of computing power types and task types leads to significant differences in log quantity statistics. Logs from computing power networks originate from different types of computing tasks, including but not limited to AI training, parallel computing, and big data batch processing. The log behavior and quantitative characteristics of different task types can be drastically different, while existing quantitative statistics methods typically use a uniform statistical approach, ignoring the differences between task categories. This can lead to confusion of log features, thus affecting the accuracy of anomaly detection. For example, in batch processing tasks, the number of log events such as "data submission" and "job completion" usually follows a specific periodic pattern and is much less frequent than in AI training tasks. If these tasks are counted together with AI training tasks, the frequent "training progress" logs may be confused with the low-frequency "job completion" logs, leading to the failure of anomaly detection.
[0096] To address the aforementioned issues, this invention proposes a category-aware counting-optimized quantitative feature extraction method to improve the effectiveness of log quantity features in anomaly detection in computing power networks. Logs in computing power networks originate from different types of computing tasks; therefore, this invention introduces a category-aware counting method to count the number of log templates for different task categories, defining an optimized category-aware counting vector. The calculation process is as follows: (twenty two),
[0097] in, Represents the set of computational task categories. Indicate category Download log template The count value. By using category-aware counting, the log patterns of different tasks can be effectively distinguished, improving the targeting of anomaly detection.
[0098] S2.2.5 Multi-dimensional Log Feature Fusion
[0099] To detect various anomalies in log sequences, traditional methods typically involve either individually checking each feature for anomalies or directly concatenating different features before inputting them into the detection model. However, for computational networks requiring complex, multi-dimensional features, these methods have significant limitations. Directly concatenated vectors assign equal weights to different features, failing to highlight features more important for anomaly detection. Furthermore, concatenation increases the dimensionality of the input features, potentially leading to increased computational overhead and introducing redundant information, thus affecting the model's generalization ability. Therefore, this invention employs a weighted feature fusion method. This method adaptively assigns weights to log features through a learnable attention mechanism, thereby improving the anomaly detection capability of computational network logs.
[0100] Let the log data feature sequence of the computing power network be represented as follows: Each of them The log feature embedding represents a single dimension, where T is the sequence length, N is the number of sequences, and C is the number of channels. The specific fusion method is as follows: First, all log feature embeddings are concatenated to obtain a vector E. Then, a linear transformation layer is used to learn the mapping relationship between channels to obtain... The formula is shown below: (twenty three), (twenty four),
[0101] in, This is a learnable parameter matrix used to adjust the information weights of different channels. Softmax normalization is applied across different dimensions to calculate the importance of log feature embeddings in each dimension. Here, W serves as the attention weight matrix, dynamically adjusting the contribution of each feature to the final log vector. Finally, each log feature is weighted using element-wise multiplication, and the fused feature Y is calculated, as shown in the following formula: (25),
[0102] Y is the final fused log feature representation, used for anomaly detection in the next step S2.3.1.
[0103] Compared with traditional splicing methods, this invention, by learning attention weights, can highlight anomalies in computing power network log anomaly detection tasks. It is more important to detect the embedding scale, improve the accuracy of detection, and avoid the problem of high-dimensional feature redundancy caused by direct splicing, thereby improving the computational efficiency and generalization ability of the model and enhancing the ability to detect sudden anomalies and long-term trend anomalies.
[0104] S2.3 Node Trustworthiness Assessment Based on Log Features
[0105] S2.3.1 Anomaly Detection Based on Multi-Step Dependency
[0106] This invention uses a log sequence fused with multi-dimensional features as input to an anomaly detection model, applying recurrent neural networks (RNNs) to detect various anomalies from different information sources. In log analysis of computing networks, anomalies are usually not directly determined by log events at a single time step, but are influenced by long-term dependencies. Therefore, relying solely on traditional RNN variants such as GRU or LSTM may lead to poor detection performance due to short-term dependency issues. To address this problem, this invention uses a multi-step dependency network to construct the anomaly detection model.
[0107] In a computing power network, let's assume a log sequence after feature fusion. , where e represents the log event at time step t. We first construct the hidden state matrix H' of the current time step t and its K previous historical time steps, as shown in the following formula: (26)
[0108] in, For the front The hidden state at each time step It is a learnable residual matrix that captures the correlation between historical log features and current features.
[0109] To extract the contributions from different historical time steps, this invention employs an attention mechanism to calculate the time step weights for H' and performs a weighted summation. (27) (28) (29) (30)
[0110] (31),
[0111] The softmax algorithm ensures that the sum of the weights of all time steps is 1, thus maximizing the contribution of more important historical time steps and minimizing the impact of irrelevant log events. Finally, the current time step... With historical characteristics To enhance feature representation capabilities, fusion is performed.
[0112] (32), (33), (34)
[0113] Then the hidden state is updated to include information from the current time step and passed to the next computation. The final hidden state is obtained as follows. It includes an integration of all historical information for that time step. The hidden state of each log entry. These are the final representations of the log in the multi-step time dependency, capturing the relationships between logs and whether the log exhibits normal or abnormal behavior patterns. The formula is shown below: (35),
[0114] Next, by examining the hidden state Anomaly scoring is performed. To convert the model's output into a percentage score, the Sigmoid activation function is used to map the output to the range [0,1]. Then, it is multiplied by 100 to obtain the percentage anomaly score A for a single log entry, as shown below: (36).
[0115] S2.3.2 Trustworthiness Score of Node Log Sets Based on Time Decay and Penalty Mechanism
[0116] In step S2.3.2, the anomaly score of the single log output in step S2.3.1 is processed to obtain a comprehensive log anomaly score over a period of time. The initial score of a node is the authentication score at the time of access, and the trust score during operation will be gradually adjusted according to changes in the node's behavior. To better reflect the decreasing impact of log anomalies over time, this invention introduces a time decay mechanism. That is, anomalies closer to the current time point contribute more to the final trust score, while earlier anomalies gradually decrease in their contribution to the score. For the k-th log, the anomaly score after time decay adjustment... It can be calculated using the following formula: (37)
[0117] Where k is the sequential index of the log, k current Represents the sequence number of the current log. This is the anomaly score for the k-th log entry obtained by the anomaly detection model.
[0118] To further enhance the accuracy and reliability of the model, this invention introduces a penalty mechanism. This mechanism is based on the anomaly rate of the logs. To determine the abnormality level of a certain log entry. Exceeding a certain set threshold At that time, the abnormality of the log will be penalized, thereby making the overall score more reflective of the actual situation. The adjusted log abnormality level As shown below: (38)
[0119] Here, α is a penalty value that controls the severity of the penalty. It is a threshold. When When the threshold is exceeded, the penalty mechanism will be activated, and the anomaly level will be adjusted by a coefficient α, with additional penalties imposed on logs with very high anomaly levels.
[0120] Ultimately, the overall score is... It is obtained by weighting the anomaly score after time decay and the score after penalty adjustment. For example, consider a time window of length T containing N log entries. The formula for calculating the comprehensive score is as follows:
[0121] (39)
[0122] in, The anomaly score for the k-th log entry after time decay adjustment. The penalty-adjusted score is given for the k-th log entry, where N is the number of log entries within the window. To facilitate merging the log anomaly scores with user scores in the next step, the min function is used to keep the overall value below 100. This results in the comprehensive log score. It reflects the degree of anomaly throughout the entire time window and can effectively assess the degree of behavioral anomalies reported by computing network nodes in the logs.
[0123] S2.3.3 Comprehensive Credibility Calculation Based on User Feedback
[0124] In a computing power network, users obtain computing power services through a scheduling system. Although they do not directly perceive the computing power nodes providing the services, they can rate the overall service quality through feedback. For example... Figure 1 As shown, after receiving computing power allocated by the scheduling system, computing power consumers can rate the actual usage of that computing power. This invention uses user ratings as a supplementary evaluation dimension, reflecting the service performance of computing power network nodes in actual operation, and conducting quantitative evaluation based on user feedback and specific node operating data. This rating dimension provides a more intuitive and comprehensive assessment of node trustworthiness.
[0125] For example, user ratings include a node’s performance in executing computational tasks, including the accuracy, completeness, and timeliness of the tasks, the node’s response speed to user requests, and whether performance degradation or crashes frequently occur during continuous operation.
[0126] Suppose a node receives m ratings within a time period T. In the special case where m equals 0, i.e., the computing node does not receive user feedback ratings during this period, the anomaly level of user feedback is set to 0, and the total anomaly ratings from user feedback are... It can be represented as:
[0127] (40)
[0128] To further improve the accuracy of node behavior credibility assessment, this invention combines user ratings with log anomaly detection results. Specifically, user ratings, as an independent dimension, are standardized and then weighted and fused with anomaly detection scores based on log data to determine the initial credibility of the node. Based on this, a comprehensive node credibility score is finally obtained. The calculation formula is as follows:
[0129] (41),
[0130] in, It is a log-based anomaly detection score, typically a value calculated by some deep learning model or statistical method. It is the average of all user ratings over a certain period of time.
[0131] It is a weighting coefficient that can be adjusted according to the actual situation; a higher weighting coefficient is preferable. This indicates a greater emphasis on log ratings, with lower ratings. This indicates a greater emphasis on user ratings. The higher the value, the less trustworthy the node is.
[0132] Finally, risk response is assessed through a comprehensive rating of the computing power nodes. The rating formula for computing power nodes is shown below: (42),
[0133] Among them Warning threshold This is the safety threshold. Anomaly scores below the safety threshold... Nodes with a level of 1 are considered safe and can continue to schedule computing tasks normally; nodes with an anomaly score greater than the safety threshold are considered safe. However, it has not yet reached the warning threshold. The node, with a level of 2, indicates a slight decrease in the node's behavioral credibility. Assign it a low-importance task and continue observation; the anomaly score has exceeded the warning threshold. A node with a level of 3 is considered untrusted, and the scheduling system will isolate that node.
Claims
1. A node zero-trust trusted access method for a computing power network, characterized in that, Comprise the following modules and steps: Module S1: trust evaluation in computing power node registration stage; Step S1.1: identity trusted verification; Step S1.2: ability trusted evaluation; Step S1.3: node trust modeling based on Bayesian prior and graph regularization; Step S1.4: access decision and dynamic evolution; Through identity trusted verification and ability trusted evaluation, trusted computing power nodes are ensured to access the computing power network, and the initial trust value is calculated, and information collection probes are deployed to the computing power nodes to facilitate real-time monitoring of the security of the nodes in the subsequent running process; Module S2: trust evaluation based on behavior log in computing power node task execution stage; Step S2.1: log collection and analysis; Step S2.2, multi-dimensional log feature extraction and fusion; Step S2.3, node trusted evaluation based on log features; Through log analysis, log templates are obtained, and semantic, time, parameter and quantity features specific to the computing power network are extracted, and the weights of each feature in the final log vector are dynamically adjusted, the long-term dependence between log events in the time series is captured, each log is accurately detected, then the time decay model and the punishment mechanism are used to obtain the log abnormal score in a period of time, and the user feedback score is fused to obtain the node comprehensive trust degree, according to the comprehensive trust degree, the scheduling system can quickly identify and isolate the node to prevent potential threats from spreading to the entire network.
2. The node zero-trust trusted access method for a computing power-oriented network according to claim 1, wherein, Step S1 arranges simple tasks for the computing power nodes requesting access to observe their completion to conduct dual authentication of ability and identity and calculate the initial trust value.
3. The node zero-trust trusted access method for a computing power-oriented network according to claim 1, wherein, The specific implementation process in step S2.2 is: Step S2.2.1: semantic feature extraction based on ALBERT and contrastive learning; Step S2.2.2: time feature extraction based on time difference change rate; Step S2.2.3: parameter feature extraction based on normalized enhancement; Step S2.2.4: quantity feature extraction based on category perception count optimization; Step S2.2.5: multi-dimensional feature fusion.
4. The node zero-trust trusted access method for a computing power-oriented network according to claim 1, wherein, The specific implementation process in step S2.3 is: Step S2.3.1: anomaly detection based on multi-step dependence; Step S2.3.2: node log set trust score based on time decay and punishment mechanism; Step S2.3.3: comprehensive trust degree calculation by fusing user feedback.
5. The node zero-trust trusted access method for a computing power-oriented network according to claim 3, characterized in that, In step S2.2.2, the dimension problem caused by different time units of different log systems of computing power nodes is solved by using time difference change rate as the time feature to identify abnormal changes in execution time when system performance problems occur.
6. The node zero-trust trusted access method for a computing power-oriented network according to claim 3, wherein, In step S2.2.3, a normalized enhanced numerical parameter feature extraction method is proposed for heterogeneous computing power networks to improve the anomaly detection capability of computing power network logs.
7. The node zero-trust trusted access method for a computing power-oriented network according to claim 3, characterized in that, In step S2.2.4, a category perception count optimization quantity feature extraction method is proposed to solve the problem of significant differences in log quantity features caused by heterogeneous computing power node types and task types, to improve the effectiveness of log quantity features in computing power network anomaly detection.
8. The node zero-trust trusted access method for a computing power-oriented network according to claim 4, wherein, Step S2.3.
1. The step S2.3.1 is aimed at the problem that the abnormality in the log analysis of the computing power network is usually not directly determined by the log events of a single time step, but is affected by long-time dependence. By explicitly inputting multiple hidden states, the long-distance abnormality in the computing power network is captured, and the abnormality degree of each log is outputted, so as to facilitate the credibility scoring.
9. The node zero-trust trusted access method for a computing power-oriented network according to claim 4, wherein, The module S2.3.2 introduces a time decay mechanism and a penalty mechanism to better reflect the decreasing influence of the log abnormality of the computing power network over time. The closer the abnormal event is to the current time point, the greater the contribution of the abnormal event to the final credibility score. When the abnormality degree of a log exceeds a certain threshold, the abnormality degree of the log is additionally punished, so that the comprehensive score can better reflect the actual situation.
10. The node zero-trust trusted access method for a computing power oriented network according to claim 4, wherein, Step S2.3.3 uses user scores as a supplementary evaluation dimension to reflect the service performance of the computing power network nodes in actual operation, and quantitatively evaluates according to the user feedback and the specific operation data of the nodes, so as to provide more intuitive and comprehensive node credibility evaluation.
Citation Information
Patent Citations
Construction and dynamic maintenance method of trusted group in electric power Internet of Things environment
CN114553458A
Security protection method and system for power terminal
WO2023216641A1