Edge intelligent ai application trusted computing power scheduling method for power internet of things

By using trusted computing root certificates and lightweight time-series prediction models in the power Internet of Things to predict the trust decay trend of edge nodes, the problem of unpredictable trust decay trends in existing technologies is solved, enabling forward-looking scheduling and lossless migration of edge nodes, and improving the security continuity of computing tasks and resource utilization efficiency.

CN122451908APending Publication Date: 2026-07-24STATE GRID HENAN INFORMATION & TELECOMM CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID HENAN INFORMATION & TELECOMM CO
Filing Date
2026-04-27
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies cannot effectively predict the trust decay trend of edge nodes in the power Internet of Things, leading to the interruption of AI inference services and the loss of data service continuity, and making it impossible to achieve forward-looking scheduling and lossless migration.

Method used

By performing integrity verification based on trusted computing root certificates and combining a lightweight time-series prediction model to predict trust decay trajectories of runtime logs and network jitter sequences, high-risk nodes are proactively identified and hot-start switching of instance runtime contexts is performed, enabling forward-looking scheduling and lossless migration of nodes.

Benefits of technology

It effectively avoids interruptions to AI inference services caused by passive responses, improves the security continuity and resource utilization efficiency of computing task scheduling, and realizes the forward-looking assessment of the trust status of edge nodes and the lossless hot migration of instance runtime context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122451908A_ABST
    Figure CN122451908A_ABST
Patent Text Reader

Abstract

The application discloses an edge intelligent AI application trusted computing power scheduling method for a power Internet of Things. A lightweight timing prediction model is used to fuse and model multi-dimensional features such as operation logs and network jitter. The decay trend is actively identified and the migration plan is triggered in advance before the node security situation falls below the threshold. Based on this concept, the AI inference business interruption and operation context loss caused by passive response are effectively avoided, the forward-looking research and judgment of the edge node trust state and the lossless hot migration of the instance operation context are realized, and the security continuity and resource utilization efficiency of the computing task scheduling in the multi-agent computing network environment are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent scheduling, and more specifically, to a trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things. Background Technology

[0002] As the power Internet of Things (IoT) evolves towards a multi-entity collaborative computing network architecture, edge AI computing tasks exhibit complex characteristics with multiple layers of complexity, including data privacy sensitivity, heterogeneous computing power requirements, and security and compliance constraints. Especially in high real-time scenarios such as power load forecasting in densely populated areas, computing tasks not only place stringent demands on node computing power, but their own runtime contexts, such as intermediate feature tensors, also possess high timeliness and state continuity. Any scheduling interruption can lead to regional-level data service failures. Therefore, it is urgent to construct a multi-dimensional, layered perception and intelligent decomposition mechanism for computing tasks in multi-entity computing networks.

[0003] Existing solutions typically employ static security authentication mechanisms based on fixed thresholds to determine privacy compliance for edge nodes and handle node anomalies using a reactive emergency response model. However, the core flaw of such mechanisms lies in their inherently passive security awareness chain, triggering a response only after a sudden change in the node's security posture. They cannot proactively predict or intervene in trust decay trends along a causal chain. When edge nodes are attacked or experience security degradation such as trust deterioration or network service interruptions, the system cannot detect the anomalies in a timely manner within the detection window. This results in the inability to extract and seamlessly migrate the runtime context of megabyte-scale AI applications hosted on the nodes, leading to the interruption of real-time AI inference services and the loss of data service continuity.

[0004] Therefore, we look forward to an optimized trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things. Summary of the Invention

[0005] To address the aforementioned technical issues, this application provides a trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things.

[0006] According to one aspect of this application, a trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things is provided, comprising: S1, based on the trusted computing root certificate, performs integrity verification on the firmware hash value and platform register metric value reported by the edge node, and performs sliding window statistics on the acquired environmental telemetry data stream to obtain the trusted node matrix and node baseline feature set; S2, based on the power data hierarchical desensitization mapping rules, performs security protocol-level constraint transformation on the data privacy labels carried in the task request to obtain execution constraints, and performs matrix mask screening on the trusted node matrix based on the execution constraints to obtain a privacy-compliant node pool; S3, within the optimization space of the privacy-compliant node pool, performs multi-objective scheduling optimization on the node baseline feature set and execution constraints to lock the target node identifier, and instantiates and deploys on the target node to continuously generate instance runtime context; S4. Based on the target node identifier, collect the corresponding node's operation log and network jitter sequence, and use a lightweight time series prediction model to predict the trust decay trajectory and determine the high risk of the above operation log and network jitter sequence to obtain the trust decay gradient vector and the high-risk node warning list. S5 responds to the high-risk node warning list, matches the takeover node from the privacy-compliant node pool based on the trust decay gradient vector, and silently clones and injects the instance runtime context into the takeover node to complete the hot start switch.

[0007] Compared with existing technologies, this application provides a trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things. This method uses a lightweight time-series prediction model to fuse and model multi-dimensional features such as operation logs and network jitter. It proactively identifies degradation trends and triggers migration plans in advance before the node's security status falls below a threshold. Based on this concept, it effectively avoids AI inference service interruptions and loss of runtime context caused by passive responses, achieving proactive assessment of the trust status of edge nodes and lossless hot migration of instance runtime contexts. This significantly improves the security continuity and resource utilization efficiency of computing task scheduling in multi-entity computing network environments. Attached Figure Description

[0008] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0009] Figure 1 A flowchart of a trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things according to an embodiment of this application; Figure 2 This is a schematic diagram of the data flow of a trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things according to an embodiment of this application; Figure 3 This is a flowchart of step S2 in the trusted computing power scheduling method for edge intelligent AI applications for the power Internet of Things according to an embodiment of this application; Figure 4 This is a flowchart of step S3 in the trusted computing power scheduling method for edge intelligent AI applications for the power Internet of Things according to an embodiment of this application. Detailed Implementation

[0010] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.

[0011] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0012] While this application makes various references to certain modules of the systems according to embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The modules described are merely illustrative, and different aspects of the systems and methods may use different modules.

[0013] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0014] The technical solution of this application proposes a trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things. Figure 1 This is a flowchart of a trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things, according to an embodiment of this application. Figure 2 This is a system architecture diagram of a trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things, according to an embodiment of this application. Figure 1 and Figure 2As shown, the trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things according to an embodiment of this application includes the following steps: S1, based on the trusted computing root certificate, performing integrity verification on the firmware hash value and platform register metric value reported by the edge nodes, and performing sliding window statistics on the acquired environmental telemetry data stream to obtain a trusted node matrix and a node baseline feature set; S2, based on the power data hierarchical desensitization mapping rules, performing security protocol-level constraint transformation on the data privacy label carried by the task request to obtain execution constraints, and performing matrix mask screening on the trusted node matrix based on the execution constraints to obtain a privacy-compliant node pool; S3, in the privacy-compliant node pool, searching... Within the optimal space, multi-objective scheduling optimization is performed on the node baseline feature set and execution constraints to lock the target node identifier, and instantiation and deployment are performed on the target node to continuously generate the instance runtime context; S4, based on the target node identifier, the runtime logs and network jitter sequences of the corresponding nodes are collected in a targeted manner, and a lightweight time series prediction model is used to predict the trust decay trajectory and determine high risk on the above runtime logs and network jitter sequences to obtain the trust decay gradient vector and the high-risk node warning list; S5, in response to the high-risk node warning list, a takeover node is matched from the privacy-compliant node pool based on the trust decay gradient vector, and the instance runtime context is silently cloned and injected into the takeover node to complete the hot start switch.

[0015] Specifically, S1, based on the trusted computing root certificate, performs integrity verification on the firmware hash value and platform register metric value reported by the edge nodes, and performs sliding window statistics on the acquired environmental telemetry data stream to obtain the trusted node matrix and node baseline feature set. It should be understood that in the real-world deployment scenario of the power IoT, the heterogeneity of edge devices is extremely prominent. Taking a typical 110kV substation as an example, multiple generations and vendors of intelligent devices are deployed simultaneously: older generation transmission line monitoring terminals have ample computing power, equipped with high-performance edge inference chips, and can smoothly run heavy AI models such as transmission line defect identification; while the new generation of secure access gateways, although equipped with State Grid-certified cryptographic machines and trusted execution environment chips, has extremely limited computing power because it is positioned for secure communication rather than computing, and cannot independently carry out AI inference tasks. In this highly heterogeneous edge device deployment environment, if the integrity of the underlying firmware and the platform's trusted status of each node are not strictly verified, nodes with maliciously tampered firmware or forged platform register metric values ​​will be mixed into the candidate resource pool, directly threatening the data security and execution reliability of subsequent AI inference tasks. At the same time, static and reliable verification results alone are not enough to support refined multi-objective scheduling decisions. The system also needs to continuously collect and statistically model the real-time operating environment characteristics of each trusted node in order to obtain baseline characteristic data reflecting key dimensions such as the current load level, thermal stability and power supply quality of the node. This provides reliable and timely node status input for the privacy compliance screening in step S2 and the multi-objective scheduling optimization in step S3.

[0016] In specific implementation, in S1-1, based on the trusted computing root certificate, the firmware hash value and platform register metric value reported by each edge node are subjected to low-level integrity verification to obtain the node trust decision identifier set. In the edge computing architecture of a multi-entity computing network, each edge node, during startup or operation, performs cryptographic hashing on its own firmware image through its built-in trusted platform module to generate a firmware hash value. Simultaneously, it sequentially expands and writes the metric results of each stage of the startup chain into the platform configuration register to form the platform register metric value. These two data sets are transmitted to the computing network scheduling center through a secure reporting channel. The computing network scheduling center holds the expected baseline values ​​for each node issued by the trusted computing root certificate. During this process, for each edge node in the computing network, its reported firmware hash value is subjected to a strict bit-by-bit cryptographic comparison with the corresponding firmware baseline hash value pre-stored in the trusted computing root certificate, and its reported platform register metric value is subjected to the same strict bit-by-bit comparison with the platform register baseline metric value pre-stored in the certificate.

[0017] Specifically, a node is considered a trusted node if and only if its firmware hash value and platform register metric value are both completely identical to the baseline value (i.e., both checks pass simultaneously). If either check fails (i.e., the firmware hash value deviates from the baseline or the platform register metric value deviates from the baseline), the node's trust decision flag is assigned a value of 0, and it is considered an untrusted node. The comparison operation in the above judgment logic is a strict bit-by-bit comparison of the cryptographic hash values. The two conditions are logically ANDed, meaning both must be satisfied simultaneously for a node to be considered trustworthy. After performing the above integrity verification operation on all edge nodes in the computing network, the trust decision flags of all nodes are summarized and arranged in order of node number, ultimately obtaining the set of node trust decision flags.

[0018] In S1-2, the trusted node matrix is ​​obtained by filtering the set of node trust decision identifiers. During this process, each element in the set of node trust decision identifiers is traversed, and all nodes with a trust decision identifier value of 1 are extracted, i.e., all trusted nodes that pass the integrity check are retained, and all untrusted nodes with a value of 0 are discarded. Assuming that a certain number of nodes pass the integrity check after filtering, these trusted nodes are renumbered according to their registration order in the computing network. For each trusted node, the system extracts its multi-dimensional attribute information from the computing network node registration database, including but not limited to node identifier, computing power specifications, memory capacity, network bandwidth, and security capability tags. The attribute information of all trusted nodes is arranged in rows to form a trusted node matrix. The row dimensions of this matrix correspond to each trusted node entity, and the column dimensions correspond to each attribute dimension of the node. Each element in the matrix represents the specific value of a trusted node in a certain attribute dimension.

[0019] In S1-3, based on the legitimate node identifiers in the trusted node matrix, the transparent telemetry data stream is precisely truncated, and the mean and unbiased variance of the truncated temperature and voltage time series are extracted using a sliding window algorithm to obtain the node baseline feature set. In the multi-entity computing network operating environment, all edge nodes (including trusted nodes that have passed integrity verification and untrusted nodes that have not passed verification) continuously transmit environmental telemetry data streams to the computing network scheduling center. This data stream contains multi-dimensional time series data such as temperature sensor readings, power supply voltage readings, CPU load rate, and memory usage rate of each node. Since the telemetry data reported by untrusted nodes may have been tampered with or injected with false information, directly using the global telemetry data stream will introduce untrusted data pollution.

[0020] Therefore, in this process, firstly, the set of legitimate node identifiers in the trusted node matrix is ​​used as the filtering index to perform a precise interception operation on the transparent telemetry data stream, retaining only the telemetry data reported by trusted nodes with matching identifiers and discarding the telemetry data of all untrusted nodes, thereby obtaining a pure subset of telemetry data after trustworthiness filtering.

[0021] Next, after accurate extraction, a sliding window algorithm is applied to extract statistical features from the temperature and voltage time series sequences of each trusted node. Specifically, a fixed-length time window is slid across the time series with a fixed step size. At each window position, local statistical operations are performed on all data points within the window's coverage area to extract statistical features for that time period. For the temperature time series sequence of each trusted node, mean and unbiased variance calculations are performed within each sliding window. Similarly, for the voltage time series sequence of each trusted node, mean and unbiased variance calculations are also performed within each sliding window.

[0022] Furthermore, for each trusted node, the four statistical measures extracted within the latest sliding window—mean temperature, unbiased temperature variance, mean voltage, and unbiased voltage variance—are combined to form the node's baseline feature vector. This vector fully characterizes the node's thermal stability and power supply quality features within the current time period. By summing the baseline feature vectors of all trusted nodes, the node baseline feature set can be obtained.

[0023] Specifically, S2, based on the power data hierarchical desensitization mapping rules, transforms the data privacy tags carried in the task request into security protocol-level constraints to obtain execution constraints, and then uses matrix masking to filter the trusted node matrix based on these execution constraints to obtain a privacy-compliant node pool. It should be understood that although the trusted node matrix has completed the underlying firmware integrity verification and trusted entity screening of all edge nodes in the computing network, ensuring the trustworthiness of each node at the hardware and platform levels, trustworthiness and privacy compliance are two different dimensions of security attributes. In the actual business scenarios of the power IoT, different types of AI computing tasks process data with varying levels of privacy sensitivity. For example, load forecasting tasks involving user electricity privacy and status monitoring tasks involving only equipment operating parameters have drastically different requirements for node privacy computing capabilities. The former may require nodes to have a trusted execution environment, federated learning components, or differential privacy injection capabilities, while the latter may only require basic data encryption transmission capabilities. Specifically, in the real deployment scenarios of the power IoT, the heterogeneity of edge devices is extremely prominent, with multiple generations and vendors of smart devices deployed simultaneously within the station, and these devices have significant differences in security capability configurations. Therefore, in the technical solution of this application, based on the data privacy label carried by the task itself, it is transformed into a quantifiable execution constraint. Using this as a benchmark, the privacy compliance capabilities of each node in the trusted node matrix are evaluated and screened one by one, eliminating nodes that do not meet the privacy constraints, and finally forming a candidate node pool that is both trusted and privacy compliant, providing a safe and compliant optimization space for the multi-objective scheduling optimization in the subsequent step S3.

[0024] Figure 3 This is a flowchart of step S2 in the trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things according to an embodiment of this application. Figure 3 As shown, in the first embodiment of this application, S2 includes: S2-1, performing affine transformation and truncation mapping on the data privacy label carried in the power grid AI task request based on a preset privacy weight mapping matrix to obtain execution constraints; S2-2, using the execution constraints, performing a node privacy support metric evaluation on the underlying privacy computing capability vector of each node in the trusted node matrix to obtain a node compliance mask vector; S2-3, based on the node compliance mask vector, performing set derivation mask filtering on the trusted node matrix to obtain a privacy-compliant node pool.

[0025] Specifically, in step S2-1, based on a preset privacy weight mapping matrix, affine transformation and truncation mapping are performed on the data privacy tags carried in the power grid AI task request to obtain execution constraints. In the multi-entity computing network architecture of the power Internet of Things, each AI computing task request carries a set of data privacy tags upon submission. These tags are labeled according to power data hierarchical desensitization mapping rules, describing the privacy-sensitive attributes of the data processed by the task. The data privacy tags cover multiple dimensions, such as whether the data involves users' personal electricity consumption information, the level of data desensitization, the type of encryption protocol required, whether a trusted execution environment isolation is required, whether federated learning framework support is needed, and whether differential privacy noise injection is required.

[0026] In this process, firstly, the data privacy tags carried in the task request are organized into a multi-dimensional vector, let this vector be... ,in For each component of the data privacy label, there are [number] dimensions. This indicates that the task is in the... Label values ​​on each privacy attribute dimension. The privacy weight mapping matrix is ​​the core quantitative carrier of the power data hierarchical de-identification mapping rules. This matrix is ​​pre-configured by the computing network management platform according to power industry data security standards and business privacy protection strategies. Let the privacy weight mapping matrix be... ,in The number of dimensions for implementing the constraints (corresponding to the various evaluation dimensions of the node's privacy computing capabilities). This represents the number of dimensions for the data privacy label. Each element in the matrix... Indicates the first The execution constraint dimension for the first The weight mapping coefficients for each privacy label dimension reflect the strength of the contribution of that privacy label attribute to that execution constraint dimension. A bias vector is also set. This is used to characterize the baseline threshold value for each execution constraint dimension, i.e., the minimum level of security constraints the system must meet even if the task does not carry any privacy label. Specifically, the affine transformation process can be expressed by the formula: in, This is a privacy weight mapping matrix. For bias vectors, This represents the original constraint vector after affine transformation. The essence of this affine transformation is to map the discrete or semi-structured privacy label space to a continuous execution constraint space through a linear weighted combination. The weight coefficients in the privacy weight mapping matrix determine the direction and intensity of the influence of each privacy label attribute on each execution constraint dimension.

[0027] Furthermore, since the output value of the affine transformation may exceed the valid range of values ​​for the constraint conditions (e.g., negative values ​​or values ​​exceeding the upper limit), a truncation mapping operation needs to be performed on the original constraint vector to restrict each component to a legal range. Specifically, for each component of the original constraint vector, values ​​less than the lower bound zero are truncated to zero, values ​​greater than the upper bound threshold are truncated to the upper bound threshold, and values ​​within the legal range remain unchanged. Let the ... The upper bound threshold for each execution constraint dimension is The process of truncating the mapping can be expressed by the formula: in, For the truncated mapping, the first One execution constraint component; This indicates taking the maximum value, used to ensure that the constraint value is not lower than zero; This indicates taking the minimum value, used to ensure that the constraint value does not exceed the upper bound threshold of that dimension. ; Indicates the first The original mapping values ​​on each execution constraint dimension. By combining the constraint components after truncating the mapping of all dimensions, the execution constraint vector is obtained. This vector fully encodes the data privacy protection requirements of the task request, and the truncated mapping ensures that all constraint values ​​are within a legal and valid range.

[0028] Specifically, in S2-2, by utilizing execution constraints, a node privacy support metric evaluation is performed on the underlying privacy computing capability vector of each node in the trusted node matrix to obtain a node compliance mask vector. In this process, firstly, for each trusted node, its underlying privacy computing capability vector is compared dimension-by-dimensionally with the execution constraint vector, determining in each dimension whether the node's actual capability value is greater than or equal to the constraint requirement value for that dimension. Specifically, for the... The trusted node is at the _th ... The calculation logic for single-dimensional compliance determination across multiple dimensions is as follows: in, Indicates the first The trusted node is at the _th ... Actual capability values ​​in each dimension of privacy computing capabilities Indicates the first The trusted node is at the _th ... The single-dimensional compliance judgment result on the privacy computing capability dimension is as follows: a value of 1 indicates that the actual capability of the node in this dimension meets or exceeds the constraints of the task, and a value of 0 indicates that the actual capability of the node in this dimension is insufficient to meet the constraints of the task.

[0029] Then, the judgment results from each dimension are aggregated to obtain the overall compliance judgment result for the node. Since privacy compliance requires nodes to meet the requirements in all constraint dimensions (i.e., the lack of capability in any one dimension will prevent the node from securely executing tasks), the overall compliance judgment uses a full-dimensional logical AND operation, that is, multiplying the single-dimensional compliance judgment results of all dimensions together. After performing the above quantitative evaluation operation on each trusted node in the trusted node matrix, the compliance mask values ​​of all nodes are summarized and arranged in order of node number, finally obtaining the node compliance mask vector. Specifically, this process can be expressed by the formula: in, For node compliance mask vectors, The vector represents the total number of trusted nodes in the trusted node matrix, and each element represents a node in the vector. With the first in the trusted node matrix Each trusted node corresponds to a private node. A value of 1 indicates that the node is a privacy-compliant node, and a value of 0 indicates that the node is a privacy-non-compliant node.

[0030] Specifically, in S2-3, based on the node compliance mask vector, a set-derived mask filtering is performed on the trusted node matrix to obtain a privacy-compliant node pool. In this process, starting from the first element of the node compliance mask vector, the mask value of each trusted node is checked sequentially. When the first element is checked... Mask value of each node When the system reads the complete attribute row vector of the node from the trusted node matrix, it writes it into the privacy-compliant node pool. When checking the node... Mask value of each node If the node is not selected, it is skipped without any retention, downgrade, or alternative marking operations, and the node and all its attribute information are excluded from the subsequent scheduling process. After traversal, the number of nodes in the privacy compliance node pool is equal to the total number of elements with a value of 1 in the node compliance mask vector.

[0031] However, in real-world deployments of the power Internet of Things (IoT), the heterogeneity of edge devices is extremely prominent. Taking a typical 110kV substation as an example, multiple generations and vendors of smart devices are deployed simultaneously within the station: older generation transmission line monitoring terminals have ample computing power and are equipped with high-performance edge inference chips, enabling them to smoothly run heavy AI models such as transmission line defect identification; while the new generation of secure access gateways, although equipped with State Grid-certified cryptographic machines and trusted execution environment chips, have extremely limited computing power because they are positioned for secure communication rather than computation, and cannot independently handle AI inference tasks.

[0032] The one-size-fits-all hard mask filtering logic used in sub-steps S2-3 of the first embodiment indiscriminately removes nodes with a mask value of 0 from the candidate resource pool. This mechanism leads to severe resource waste in the above scenario: the IED terminal with ample computing power is excluded entirely simply because it lacks a TEE security chip, leaving its valuable inference computing power idle. When the power grid faces AI tasks with high security constraints (such as load prediction involving user electricity privacy), the number of available nodes will precipitate sharply, severely impacting the overall system throughput and resource utilization, and in extreme cases, even making task scheduling impossible.

[0033] The root of the problem lies in the fact that the first embodiment completely ignores the objectively existing local topology compensation and cooperation relationships between heterogeneous edge nodes. Within the local area network of the same substation or distribution area, the communication latency between nodes is extremely low (usually within milliseconds), which provides a physical basis for collaborative computing with decoupled computing power and security: nodes with strong computing power but lacking security chips can be responsible for non-extremely sensitive computational tasks such as feature extraction, while nodes with complete security capabilities but insufficient computing power are specifically responsible for the encryption and aggregation of sensitive gradients. This complementary capability and physically adjacent cooperative mode is completely ignored in the first embodiment, which is the most critical design blind spot in the entire scheduling scheme.

[0034] In view of the above-mentioned technical defects, this application further proposes a second embodiment.

[0035] Specifically, firstly, based on the node compliance mask vector, the verified edge node matrix is ​​rigidly partitioned. Nodes with a mask value of 1 are assigned to the perfectly compliant node set, while nodes with a mask value of 0 are assigned to the node set to be compensated. Next, gap quantization is performed by calculating the dimension-by-dimensional difference between each node in the node set to be compensated and the task execution constraints to obtain the capability gap tensor of the damaged nodes.

[0036] Specifically, firstly, the verified edge node matrix is ​​partitioned using node compliance mask vectors. Nodes with a mask value of 1 are assigned to the perfectly compliant node set, while nodes with a mask value of 0 are assigned to the node set to be compensated, with the latter being retained instead of being discarded. Subsequently, for each node in the node set to be compensated... A normalized modified linear truncation function is introduced to calculate its relative capability gap with respect to task execution constraints across various privacy constraints, generating a capability gap tensor for damaged nodes, expressed as: in, For nodes The normalized capability gap vector; This is the task requirement vector extracted from the task execution constraints; For nodes The actual privacy-preserving computational capabilities it possesses; To correct linear units element by element, only the unsatisfied gap portions are retained; is the L1 norm of the demand vector, used to normalize the gap to the [0,1] interval; To prevent extremely small constants with a denominator of zero.

[0037] By quantifying the gaps instead of simply outputting a Boolean value of 0, the system can accurately record whether each damaged node is missing only a TEE or only a federated learning component, providing a calculable quantitative basis for subsequent precise jigsaw puzzle-style compensation. This is the information theory foundation upon which the entire second embodiment can stand.

[0038] Furthermore, based on the capability gap tensor of the damaged nodes, for each damaged node in the set of replacement nodes, candidate replacement nodes are searched in the set of perfectly compliant nodes to obtain a virtual compliant collaborative node set. It should be understood that capability complementarity alone is insufficient—if two complementary nodes belong to different substations, the high latency of cross-station communication will cause severe timeouts for AI inference tasks, making collaboration a burden. Therefore, when searching for replacement partners, both capability coverage and physical topological proximity must be considered simultaneously; neither can be neglected.

[0039] Specifically, for the damaged nodes in the cluster of nodes awaiting compensation Search for candidate compensation nodes in the perfectly compliant node cluster. The feasibility score for virtual node collaboration is calculated by combining the complementary coverage rate of joint capabilities with the topological communication distance in the power grid topology physical distance matrix, and is expressed as follows: in, For nodes With candidate compensation nodes The feasibility score for collaboration between them; The node generated in step one Capacity gap vector; Candidate compensation node Privacy-preserving computational capability vector; Represents a node Post-compensation node The remaining capacity gap, the smaller the numerator, the higher the compensation coverage rate; This is the delay-sensitive attenuation coefficient for power communication networks, reflecting the tolerance of services for communication delays; Nodes extracted from the physical distance matrix of the power grid topology and Topological communication distance (e.g., network hop count); exponential decay term The forced system only seeks partners within the same microgrid or local area network, fundamentally avoiding the latency risks associated with cross-site collaboration.

[0040] Furthermore, the score exceeds the preset collaboration threshold. The nodes are bound and aggregated to form a whole that logically satisfies all constraints, and pushed into a set of virtual compliant collaborative nodes. In this way, previously abandoned high-computing-power nodes can be brought back into the scheduling field of view in the form of virtual compliant nodes.

[0041] Subsequently, the perfectly compliant node set and the virtual compliance collaborative node set are expanded and merged to obtain a privacy compliance node pool. After the above process, the system simultaneously holds both a physically independent set of perfectly compliant nodes and a logically combined set of virtual compliance collaborative nodes. These two sets are then merged using a union aggregation operator to generate an expanded, highly resilient privacy compliance node pool, represented as: in, This is the final expanded version of the privacy-compliant node pool; A perfectly compliant set of nodes that requires no compensation; For the node With nodes The virtual node attribute vector is formed by the superposition of computing power and security capabilities. Characteristic combination operator; The system sets a minimum collaborative compensation threshold, and only virtual node pairs with scores exceeding this threshold are accepted into the resource pool.

[0042] At the same time, for the virtual collaborative nodes in the pool, computation-security decoupled microservice communication routing rules are injected into the original task execution constraints to specify the data flow within the nodes. and The internal forwarding graph between nodes upgrades the original constraints to enhanced task execution constraints, ensuring that downstream schedulers can correctly understand and distribute distributed collaborative tasks, and will not mistakenly treat virtual nodes as single physical nodes.

[0043] In particular, the second embodiment fundamentally breaks down the physical entity boundary limitations of edge computing power. In the real-world deployment environment of the power IoT, facing AI tasks with high security constraints, the number of available nodes no longer drastically decreases due to blanket filtering. Heterogeneous nodes with ample computing power but single-item shortcomings in security capabilities can be reactivated through a topology-aware compensatory collaboration mechanism, participating in scheduling competition as virtual compliant nodes. This mechanism significantly improves the overall resource utilization of the edge computing pool, reduces the scheduling failure rate of high-security-constrained tasks, and strictly constrains the collaborative communication overhead within the local area network through a topology distance penalty term, ensuring the real-time requirements of AI inference tasks. Ultimately, it achieves the synergistic attainment of the two goals of security compliance and efficient computing power at the edge of the power IoT.

[0044] Specifically, in step S3, within the optimization space of the privacy-compliant node pool, multi-objective scheduling optimization is performed on the node baseline feature set and execution constraints to lock the target node identifier, and instantiated and deployed on the target node to continuously generate instance runtime context. It should be understood that although the privacy-compliant node pool has completed a full-dimensional privacy computing capability assessment and hard mask filtering of the trusted node matrix, ensuring that each node in the pool passes the underlying firmware integrity verification and meets all privacy constraints of the task, the privacy-compliant node pool typically contains multiple candidate nodes that meet the conditions. These nodes differ significantly in operational status dimensions such as computing power specifications, memory capacity, network bandwidth, thermal stability, and power supply quality. In real-world deployments of the power Internet of Things (IoT), the heterogeneity of edge devices is extremely prominent. Taking a typical 110kV substation as an example, multiple generations and vendors of smart devices are deployed simultaneously. Older generation transmission line monitoring terminals have ample computing power, equipped with high-performance edge inference chips, enabling smooth operation of demanding AI models such as transmission line defect identification. However, while the new generation of secure access gateways incorporates State Grid-certified cryptographic machines and trusted execution environment chips, their computing power is extremely limited due to their focus on secure communication rather than computation, making them unable to independently handle AI inference tasks. In such a highly heterogeneous set of candidate nodes, randomly selecting nodes to execute tasks without refined multi-objective scheduling optimization can lead to problems such as excessively high AI inference latency, insufficient privacy constraints, or unstable node operating environments, severely impacting the execution quality and service continuity of computational tasks. Therefore, in the technical solution of this application, within the optimization space of the privacy compliance node pool, the real-time running status of each candidate node and the execution constraints of the task are comprehensively considered. The optimal target node is locked from the pool through a multi-objective scheduling optimization algorithm, and the instantiation deployment of the AI ​​model and the continuous generation of the running context are completed on the node. This provides transferable instance running status data for the trust decay trajectory prediction in the subsequent step S4 and the hot start switching in step S5.

[0045] Figure 4This is a flowchart of step S3 in the trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things according to an embodiment of this application. Figure 4 As shown, S3 includes: S3-1, using the legitimate node identifiers in the privacy-compliant node pool as indexes, matching and parsing the node baseline feature set to obtain the candidate node comprehensive state matrix; S3-2, based on the candidate node comprehensive state matrix, using the lowest inference latency as the objective function, substituting the execution constraints as linear penalty terms into a multi-objective optimization algorithm to calculate the fitness score and lock the node with the highest score to obtain the target node identifier; S3-3, based on the target node identifier, issuing deployment instruction messages to the corresponding physical device, deserializing and loading the AI ​​model within the allocated isolated memory interval, and starting a memory hook daemon process to capture stack information and feature tensors at high frequency, continuously generating instance runtime context.

[0046] Specifically, in step S3-1, the baseline feature set of nodes is matched and parsed using the legitimate node identifiers in the privacy-compliant node pool as indexes to obtain the comprehensive state matrix of candidate nodes. In this process, firstly, each legitimate node identifier in the privacy-compliant node pool is traversed, and a baseline feature vector that perfectly matches that identifier is searched in the node baseline feature set. The found baseline feature vector is then extracted.

[0047] Next, for each legitimate node in the privacy-compliant node pool, its attribute row vector in the privacy-compliant node pool includes attribute information such as computing power specifications, memory capacity, network bandwidth, and security capability labels. The corresponding baseline feature vector extracted from the node baseline feature set through identifier matching includes four statistics: mean temperature, unbiased temperature variance, mean voltage, and unbiased voltage variance.

[0048] Then, the attribute row vector of the node is concatenated end-to-end with the baseline feature vector to merge them into a longer comprehensive state row vector. After performing the above matching parsing and horizontal concatenation operations on all legitimate nodes in the privacy-compliant node pool, the comprehensive state row vectors of all nodes are arranged row by row to finally obtain the comprehensive state matrix of the candidate nodes.

[0049] Specifically, in S3-2, based on the candidate node comprehensive state matrix, the minimum inference latency is used as the objective function. The execution constraints are substituted as linear penalty terms into a multi-objective optimization algorithm to calculate the fitness score and lock the node with the highest score to obtain the target node identifier. In this process, for each candidate node in the candidate node comprehensive state matrix, the expected inference latency for the current AI inference task is estimated based on the computing power specifications, memory capacity, and network bandwidth attributes in its comprehensive state row vector. Specifically, the computational load of the current AI task is divided by the computing power specifications of the node to obtain the computational latency component; the data transmission volume of the current AI task is divided by the network bandwidth of the node to obtain the communication latency component; the computational latency component and the communication latency component are added to obtain the estimated expected inference latency value for the node. The smaller this latency estimate, the faster the node executes the AI ​​inference task, and the better its inference performance.

[0050] Next, since the multi-objective optimization algorithm needs to lock onto the node with the highest score, and inference latency is a metric where lower is better, the latency estimate is transformed into a score format where higher is better. Specifically, the inverse of the expected inference latency estimate is taken, i.e., 1 is divided by the expected inference latency estimate to obtain the inference performance score. The higher the score, the lower the inference latency and the better the inference performance of that node.

[0051] Furthermore, for each privacy-preserving computing capability dimension, the actual capability value of the node in that dimension is subtracted from the required value of the execution constraints in that dimension to obtain the constraint satisfaction margin for that dimension. Then, the constraint satisfaction margin for each dimension is multiplied by the corresponding penalty weight coefficient (pre-configured by the computing network management platform according to business priorities). Finally, the weighted margins of all dimensions are summed to obtain the privacy constraint margin score for that node. The essence of this linear penalty term is to use the execution constraints as a baseline and weightedly accumulate the capability margins of each candidate node that exceed the baseline. The larger the margin, the higher the score, thus favoring the selection of nodes with more robust privacy protection during the optimization process.

[0052] Subsequently, to incorporate the stability of the node's operating environment into the optimization considerations, an environmental stability score needs to be constructed based on the baseline feature information in the candidate node's comprehensive state matrix. The unbiased variance of temperature and voltage in the node's baseline feature vector reflects the degree of fluctuation in the node's operating environment; the smaller the variance, the more stable the environment. Finally, the inference performance score, privacy constraint margin score, and environmental stability score are weighted and combined to form the comprehensive fitness score for each candidate node. After calculating the comprehensive fitness score for each candidate node in the candidate node's comprehensive state matrix, the node with the highest score is selected as the target node.

[0053] Specifically, in step S3-3, a deployment instruction message is sent to the corresponding physical device based on the target node identifier. Within the allocated isolated memory region, the AI ​​model is deserialized and loaded, and a memory hook daemon is started to capture stack information and feature tensors at high frequency, continuously generating instance runtime context. In this process, firstly, the computing network scheduling center sends a deployment instruction message to the corresponding physical edge device based on the target node identifier. This deployment instruction message is a structured control message containing key information such as the task identifier, the storage address of the serialized binary file of the AI ​​model, the configuration parameters required for model operation, and the starting address and length of the isolated memory region allocated to the task. After receiving the deployment instruction message, the target node first allocates a dedicated isolated memory region in its physical memory space. This isolated memory region is strictly isolated from the memory spaces of other running tasks or system processes on the node, ensuring that the AI ​​model's runtime data is not illegally accessed or tampered with by other processes, and also preventing abnormal operation of the AI ​​model from affecting other services on the node. The size of the isolated memory region is determined by the length parameter specified in the deployment command message. This parameter is calculated by the computing network scheduling center based on the memory usage requirements of the AI ​​model and the available memory capacity of the target node.

[0054] Next, the target node deserializes and loads the AI ​​model within the allocated isolated memory region. The AI ​​model is stored and transmitted in the computing network as a serialized binary file, containing all information such as the model's network structure definition, layer weight parameters, optimizer state, and inference configuration. Specifically, the target node retrieves the serialized binary file of the AI ​​model from the storage address specified in the deployment command message, reads the file into the isolated memory region, and then parses and restores the binary data stream field by field according to the serialization protocol. This process reconstructs the model's network structure object, loads the layer weight parameters into the corresponding memory tensors, restores the optimizer state and inference configuration, and finally constructs a complete AI model instance capable of performing inference operations within the isolated memory region. After deserialization and loading are complete, the AI ​​model instance is in a ready state, capable of receiving input data and performing forward inference operations.

[0055] Subsequently, the target node initiates a memory hook daemon process to capture stack information and feature tensors during the execution of the AI ​​model instance at high frequency, continuously generating instance runtime context. The memory hook daemon process is a background monitoring process running in parallel with the AI ​​model instance. Its core function is to non-intrusively capture and snapshot key runtime state data within an isolated memory region at high frequency (e.g., millisecond intervals) without interfering with the normal inference computation of the AI ​​model. The data captured by the memory hook daemon process includes two main categories: the first is stack information, which is runtime control flow information such as function call stack state, local variable values, and program counter position during the execution of inference computation by the AI ​​model instance. This information records which step the model inference has reached and which layer of computation it is currently in; the second is feature tensors, which are intermediate computation result tensors generated by each intermediate layer of the AI ​​model during forward inference. These tensors contain the intermediate states of the model's layer-by-layer feature extraction and transformation of input data, and are the most crucial computational state data during model inference. By serializing and packaging the captured stack information and feature tensors within each capture cycle, a timestamped runtime context snapshot is generated. These snapshots accumulate chronologically to form the instance runtime context. The instance runtime context is a dynamically growing data structure that fully records the evolution of the AI ​​model instance's runtime state from deployment and startup to the current moment, with the latest snapshot reflecting the current runtime state of the model instance.

[0056] Specifically, in step S4, the corresponding node's operation logs and network jitter sequences are collected based on the target node identifier. A lightweight time-series prediction model is then used to predict the trust decay trajectory and determine high-risk status of the operation logs and network jitter sequences to obtain a trust decay gradient vector and a high-risk node warning list. It should be understood that the operating environment of edge nodes is not static. In real-world deployment scenarios of the power IoT, target nodes may face various security status changes during task execution, including remote injection of malicious code into firmware, side-channel attacks on trusted platform modules, continuous degradation of network link quality, and abnormal increases in node temperature leading to hardware performance degradation. Existing solutions employ a passive security authentication mechanism based on a fixed threshold. Its logical chain is: monitoring node status, and if the trust level falls below the threshold, blocking and triggering migration. However, in real physical environments, the time window from a node being "damaged" to "network connection loss / system crash" can be extremely short. If tens or even hundreds of megabytes of AI application context and model weights are only packaged when the node is deemed untrustworthy, the original node, due to its damage, may have its computation process suspended or its I / O interface cut off, resulting in the incomplete export of context data. Therefore, in the technical solution of this application, the security defense is upgraded from the traditional "0 / 1" state machine model (secure / insecure) to the concept of a continuous physical quantity, "trust decay gradient calculation." This combines implicit physical features such as network jitter and temperature / pressure drift with firmware hashes to calculate the "decay rate / gradient" of the trust value. Within the warning time window when a node is predicted to become untrustworthy, a backup node is selected in advance and a silent background clone is performed. Then, based on the target node identifier, the real-time operation log and network jitter sequence of the node are collected in a targeted manner. A lightweight time series prediction model is used to predict the trust decay trajectory and solve the gradient of the above multi-dimensional time series data. The high-risk nodes whose trust status is deteriorating rapidly are identified through the over-limit judgment mechanism, so as to provide early warning trigger signals and trust decay quantification basis for the hot start switch in the subsequent step S5.

[0057] In specific implementation, in S4-1, based on the target node identifier, the real-time node operation log and the real-time network jitter sequence are timestamp window aligned and tensor concatenated to obtain the fused state time series tensor. In the multi-entity computing network operating environment, the target node continuously generates two types of time series data during the execution of AI inference tasks. The first type is the real-time node operation log, which records multi-dimensional operating status information of the target node at a fixed sampling frequency, including CPU utilization, memory usage, firmware hash verification status, changes in trusted platform module metrics, temperature sensor readings, voltage sensor readings, etc. The second type is the real-time network jitter sequence, which records network communication quality indicators between the target node and the computing network scheduling center at a fixed sampling frequency, including round-trip latency, packet loss rate, bandwidth fluctuation rate, latency jitter variance, etc. The sampling frequencies of these two types of time series data may be different (e.g., the operation log is sampled once per second, and the network jitter sequence is sampled once every 500 milliseconds), and their respective timestamp markers may have slight offsets. Therefore, a timestamp window alignment operation must be performed before fusion. Specifically, the targeted data collection process based on the target node identifier is as follows: The network scheduling center uses the target node identifier as a filtering index to accurately extract the operation log entries that belong only to the target node from the global operation log data stream. At the same time, it accurately extracts the network jitter sampling records that belong only to the target node from the global network monitoring data stream, discarding the data of other nodes to ensure that subsequent processing is only for the status information of the target node.

[0058] In this process, firstly, a uniform time window length and time step are set, dividing the continuous time axis into equally spaced discrete time windows. For each time window, all runtime log sampling points falling within that window's time range are aggregated and statistically analyzed (e.g., by taking the average) to obtain the runtime log feature value within that window; similarly, all network jitter sampling points falling within that window's time range are aggregated and statistically analyzed to obtain the network jitter feature value within that window. Through this windowed aggregation operation, two types of time-series data that originally had different sampling frequencies are unified onto the same time scale, with each time window corresponding to one runtime log feature value and one network jitter feature value.

[0059] Furthermore, within each time window, the aggregated feature vector of the operation log (containing multiple dimensions such as CPU utilization, memory usage, temperature readings, voltage readings, and firmware hash verification status) and the aggregated feature vector of network jitter (containing multiple dimensions such as round-trip latency, packet loss rate, bandwidth fluctuation rate, and latency jitter variance) are horizontally concatenated. This involves connecting the two vectors end-to-end to merge them into a longer fused state vector. The first half of this vector represents the operation log features, and the second half represents the network jitter features. By vertically stacking the fused state vectors of all time windows in chronological order, a fused state time-series tensor can be constructed. The row dimension of this tensor corresponds to the time window sequence (time dimension), and the column dimension corresponds to all concatenated dimensions of the operation log features and network jitter features (feature dimensions). This tensor fully encodes the joint evolution trajectory of the target node's multidimensional operation state and network communication quality within the most recent few time windows, serving as direct input data for subsequent lightweight time-series prediction models.

[0060] In S4-2, the fused state temporal tensor is input into a lightweight temporal prediction model for forward inference and gradient calculation to obtain the trust decay gradient vector. The lightweight temporal prediction model is a temporal prediction neural network with a small number of parameters and low inference computational overhead, suitable for deployment in edge computing environments to perform real-time inference. This model takes the fused state temporal tensor as input and learns from the temporal dependencies and multi-dimensional feature interaction patterns in historical time-series data to predict the trust state value of the target node at several future time steps.

[0061] In this process, firstly, the fused state temporal tensor is fed into a lightweight temporal prediction model. Internally, the model performs layer-by-layer transformations and feature extraction on the input tensor according to its network structure (e.g., a lightweight recurrent neural network or a one-dimensional convolutional network), ultimately generating a prediction trust state matrix at the output layer. The row dimensions of this matrix correspond to several future prediction time steps, and the column dimensions correspond to multiple evaluation dimensions of the trust state (e.g., firmware integrity trust, network communication trust, operating environment trust, etc.). Each element in the matrix represents the trust state value of the target node predicted by the model at a certain future time step for a specific trust dimension.

[0062] Next, the gradient of the predicted trust state value with respect to time is calculated to quantify the rate and direction of change of the trust state in each dimension. Specifically, for each evaluation dimension of the trust state, the rate of change of the predicted trust state value along the time axis is calculated, i.e., the difference in trust state values ​​between adjacent prediction time steps. During execution, for a certain trust dimension, the trust state value in the next prediction time step is subtracted from the trust state value in the current prediction time step, and then divided by the time interval between adjacent prediction time steps to obtain an approximate instantaneous decay gradient value for that dimension at that time step. When the gradient value is negative, it indicates that the trust state is decaying in that dimension (i.e., the trust level decreases over time); when the gradient value is positive, it indicates that the trust state is recovering in that dimension; when the gradient value is close to zero, it indicates that the trust state remains stable in that dimension.

[0063] Furthermore, for a given trust dimension, the instantaneous decay gradient values ​​of that dimension at all adjacent time steps of the computable differences are summed, and then divided by the total number of time steps of the computable differences to obtain the average decay gradient of that dimension. This average decay gradient reflects the overall decay trend of the target node in that trust dimension. Combining the average decay gradients of all trust dimensions yields the trust decay gradient vector. Each component of this vector corresponds to the average decay rate of a trust dimension, and the vector as a whole completely characterizes the decay direction and intensity of the target node's trust state in each dimension.

[0064] In S4-3, the L2 norm of the trust decay gradient vector and the security degradation slope threshold are used for limit exceedance judgment, and the target node identifier that triggers the alarm is added to the early warning pool to obtain the high-risk node early warning list. In this process, since the trust decay gradient vector is a multi-dimensional vector, each component reflects the decay rate on different trust dimensions. Therefore, firstly, this multi-dimensional information is compressed into a scalar index that can reflect the overall decay intensity. In the embodiments of this application, the L2 norm (i.e., the Euclidean norm) is used as the measure of the overall decay intensity. Specifically, the decay gradient on each trust dimension is regarded as a vector in a multi-dimensional space, and the L2 norm is the length (modulus) of this vector, reflecting the comprehensive decay rate of the trust state across all dimensions. The larger the L2 norm, the faster the overall trust state of the target node is deteriorating; the smaller the L2 norm, the more stable the overall trust state or the slight fluctuations only exist in individual dimensions.

[0065] The security degradation slope threshold is a scalar constant pre-configured by the network management platform based on power business security standards. This threshold defines the safe upper limit of the trust decay rate—when the L2 norm of the trust decay gradient vector does not exceed this threshold, the trust decay of the target node is considered to be within the normal fluctuation range, and no warning needs to be triggered; when the L2 norm exceeds this threshold, the trust status of the target node is considered to be deteriorating at a rate exceeding the safe tolerance range, posing a high risk of falling below the trust failure threshold in a short period of time, requiring immediate warning. The L2 norm value of the trust decay gradient vector is compared with the security degradation slope threshold. If the L2 norm value is strictly greater than the security degradation slope threshold, the result is to trigger an alarm, indicating that the trust decay rate of the target node has exceeded the safe tolerance range; if the L2 norm value is less than or equal to the security degradation slope threshold, the result is normal, indicating that the trust decay of the target node is within the normal range, and no alarm is triggered.

[0066] When the limit violation assessment results in an alarm being triggered, the system appends the target node's identifier, along with its corresponding trust decay gradient vector and L2 norm value, to the early warning pool. The early warning pool is a dynamically maintained data structure that continuously accumulates information on all alarm-triggered nodes. All currently accumulated alarm node entries in the early warning pool constitute the high-risk node early warning list. Each entry in this list contains three pieces of information: the identifier of the target node that triggered the alarm, the trust decay gradient vector corresponding to the node, and the L2 norm value of that vector.

[0067] Specifically, in step S5, in response to the high-risk node warning list, a takeover node is matched from the privacy-compliant node pool based on the trust decay gradient vector, and the instance runtime context is silently cloned and injected into the takeover node to complete the hot start switch. It should be understood that in a real physical environment, the time window from a node being "damaged" to "network connection lost / system crash" can be extremely short. If tens or even hundreds of megabytes of AI application context and model weights are only packaged when the node is deemed untrustworthy, the original node, due to its damage, may have its computation process suspended or its I / O interface cut off, resulting in the inability to completely export the context data. Therefore, in the technical solution of this application, the remaining available time window of the high-risk node is accurately calculated based on the trust decay gradient vector, a suitable takeover node is matched from the privacy-compliant node pool, and the silent cloning injection of the instance runtime context and the redirection switch of data packet routing are completed before the remaining time window expires. This achieves a lossless hot start migration of AI inference tasks from high-risk nodes to takeover nodes, ensuring that business continuity is not affected by trust decay events.

[0068] In specific implementation, in S5-1, based on the trust decay gradient vector corresponding to the alarm nodes in the high-risk node warning list, the difference between the current trust state value and the failure threshold is extrapolated to obtain the remaining available time window. In this process, assuming that the trust state of the alarm node decays approximately linearly along the current decay gradient direction in the short term, the time required for the current trust state value to decay to the failure threshold can be estimated by dividing the difference between the current trust state value and the failure threshold by the decay rate. Specifically, the difference between the current trust state value and the failure threshold is extrapolated using the following formula: in, This represents the current trust status value. This is the failure threshold. For the trust decay gradient vector, Let L2 be the confidence decay gradient vector. This represents the remaining available time window.

[0069] In S5-2, a neighboring legitimate node is matched within the privacy-compliant node pool to obtain the takeover backup node identifier. It should be understood that in the edge computing architecture of the power IoT, the communication latency between nodes within the same substation or distribution area is extremely low (typically within milliseconds), while the communication latency between nodes across substations or regions increases significantly. Therefore, prioritizing legitimate nodes within the same local area network as the alarm node as the takeover backup node can effectively shorten context transmission time and increase the probability of a successful migration within the remaining available time window.

[0070] In this process, the alarm node itself is first excluded from the privacy-compliant node pool, resulting in a set of candidate takeover nodes. Then, for each node in the candidate takeover node set, the network topology distance between it and the alarm node is calculated. The network topology distance can be measured using metrics such as network hop count (i.e., the number of network switching devices a data packet needs to pass through from the alarm node to the candidate node) or measured round-trip latency. The system selects the node with the smallest network topology distance from the alarm node in the candidate takeover node set, designates it as the takeover backup node, and outputs its identifier as the takeover backup node identifier. If multiple candidate nodes have the same network topology distance as the alarm node, the overall fitness scores of these nodes are further compared, and the node with the highest score is selected as the takeover backup node.

[0071] In S5-3, a spatiotemporal constraint check is performed based on an indicator function to verify whether the sum of the instance runtime context transmission time and the takeover standby node warm-start time is less than the remaining available time window. If the check passes, an encrypted side channel is established to inject the instance runtime context into the memory address space of the takeover standby node to obtain a context injection completion status code. This step aims to prevent the system from blindly initiating migration operations when the remaining time is insufficient, which could lead to the alarm node's trust state falling below the failure threshold before the migration process is completed, resulting in both migration failure and service interruption. The spatiotemporal constraint check uses an indicator function to strictly compare the total migration operation time with the remaining available time window; migration is only allowed to start when the total time is strictly less than the remaining time window.

[0072] In this process, firstly, the time required to transmit the instance runtime context data from the alarm node to the takeover standby node depends on the size of the instance runtime context data and the network transmission bandwidth between the alarm node and the takeover standby node. Next, after receiving the instance runtime context data, the time required for the takeover standby node to load it into memory and restore the AI ​​model instance's runtime state depends on the takeover standby node's computing performance and the complexity of the instance runtime context data.

[0073] Then, the transmission time of the instance runtime context is added to the warm-start time of the takeover standby node to obtain the total migration operation time. An indicator function is then used to determine if this total time is strictly less than the remaining available time window. When the total time is strictly less than the remaining available time window, the indicator function is set to 1, the spatiotemporal constraint check passes, and the system can safely start the migration operation. When the total time is greater than or equal to the remaining available time window, the indicator function is set to 0, the spatiotemporal constraint check fails, and the system reverts to sub-step S5-2 to re-match a takeover standby node with a closer network topology or stronger computing performance to reduce transmission time or warm-start time, until the spatiotemporal constraint check passes.

[0074] Finally, after successful verification, the system establishes an encrypted side channel between the alarm node and the takeover standby node. The encrypted side channel is a dedicated encrypted transmission link independent of the normal business data channel. It employs an end-to-end encryption protocol to protect transmitted data, ensuring that the instance's runtime context is not eavesdropped on or tampered with during transmission. The side channel emphasizes that this transmission link is independent of the normal data flow of the AI ​​inference task; the context cloning injection operation is performed silently in the background, without interfering with the AI ​​inference task running on the alarm node. Through the encrypted side channel, the system serializes and packages the latest instance runtime context (including stack information snapshots and feature tensor snapshots) captured by the memory hook daemon on the alarm node and transmits it to the takeover standby node. Upon receiving the context data, the takeover standby node deserializes and restores the received data within its pre-allocated isolated memory region, loading the stack information and feature tensors into the corresponding memory address space to reconstruct the runtime state of the AI ​​model instance. Once the context data reception, deserialization, and memory loading are all completed, the takeover standby node returns a context injection completion status code to the computing network scheduling center, indicating that the instance runtime context has been successfully injected and the takeover standby node is now capable of continuing to execute AI inference tasks from the latest runtime status of the alarm node.

[0075] In S5-4, based on the context injection completion status code, the flow table update instruction of the software-defined network controller is triggered. This redirects the data packets from the original high-risk node to the takeover standby node and promotes the takeover standby node identifier to the migration target node identifier. During this process, the network scheduling center, upon receiving the context injection completion status code, first verifies its validity to confirm that the instance runtime context has been completely and correctly injected into the takeover standby node. After successful verification, the network scheduling center issues a flow table update instruction to the software-defined network controller. The software-defined network controller is a centralized control component in a multi-entity network responsible for managing network data plane forwarding rules. Its maintained flow tables define the forwarding paths of each data packet in the network. The flow table update instruction modifies all forwarding rules in the flow table that previously pointed to the original high-risk node to now point to the takeover standby node; that is, it replaces the network address of the original high-risk node with the network address of the takeover standby node.

[0076] Subsequently, after receiving the flow table update instruction, the software-defined network controller distributes the updated forwarding rules to each network switching device in the data plane. From the moment the flow table update takes effect, all data packets originally destined for the high-risk node (including input data for AI inference tasks, scheduling control messages, etc.) will be automatically redirected by the network switching devices to the takeover standby node. Since the takeover standby node has obtained the same AI model instance running state as the high-risk node through context injection in sub-step S5-3, the takeover standby node can seamlessly continue executing AI inference tasks from the latest running state of the high-risk node without having to reload the model from scratch or reprocess historical data, achieving a true hot start switchover.

[0077] Finally, after the flow table update is completed and the data packet routing is successfully redirected, the system promotes the takeover standby node identifier to the migration target node identifier. Specifically, in the task management record of the computing network scheduling center, the identifier of the current AI computing task's carrying node is updated from the original high-risk node identifier to the takeover standby node identifier, making the takeover standby node officially the new target node for the task. After the promotion, the continuous monitoring loop in step S4 will automatically switch the monitoring object, using the new migration target node identifier as an index to collect the operation logs and network jitter sequences of the takeover standby node, and continuously and proactively monitor its trust status, forming a complete closed-loop security mechanism. At the same time, the original high-risk node is removed from the scheduling view of the current task after the data packet routing redirection is completed, and its subsequent trust status restoration or offline processing is handled separately by the node operation and maintenance module of the computing network management platform.

[0078] In summary, the trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things, according to the embodiments of this application, is explained. It uses a lightweight time-series prediction model to fuse and model multi-dimensional features such as operation logs and network jitter, proactively identifying degradation trends and triggering migration plans in advance before the node's security status falls below a threshold. Based on this concept, it effectively avoids AI inference service interruptions and loss of runtime context caused by passive responses, achieves forward-looking assessment of the trust status of edge nodes and lossless hot migration of instance runtime contexts, and significantly improves the security continuity and resource utilization efficiency of computing task scheduling in a multi-entity computing network environment.

[0079] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things, characterized in that, include: S1, based on the trusted computing root certificate, performs integrity verification on the firmware hash value and platform register metric value reported by the edge node, and performs sliding window statistics on the acquired environmental telemetry data stream to obtain the trusted node matrix and node baseline feature set; S2, based on the power data hierarchical desensitization mapping rules, performs security protocol-level constraint transformation on the data privacy labels carried in the task request to obtain execution constraints, and performs matrix mask screening on the trusted node matrix based on the execution constraints to obtain a privacy-compliant node pool; S3, within the optimization space of the privacy-compliant node pool, performs multi-objective scheduling optimization on the node baseline feature set and execution constraints to lock the target node identifier, and instantiates and deploys on the target node to continuously generate instance runtime context; S4. Based on the target node identifier, collect the corresponding node's operation log and network jitter sequence, and use a lightweight time series prediction model to predict the trust decay trajectory and determine the high risk of the above operation log and network jitter sequence to obtain the trust decay gradient vector and the high-risk node warning list. S5 responds to the high-risk node warning list, matches the takeover node from the privacy-compliant node pool based on the trust decay gradient vector, and silently clones and injects the instance runtime context into the takeover node to complete the hot start switch.

2. The trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things according to claim 1, characterized in that, Step S1 includes: S1-1, based on the trusted computing root certificate, performs low-level integrity verification on the firmware hash value and platform register metric value reported by each edge node to obtain the node trust decision identifier set; S1-2, the trusted node matrix is ​​obtained by filtering the trusted node entities from the set of node trust decision identifiers; S1-3, based on the legitimate node identifiers in the trusted node matrix, accurately extracts the transparent telemetry data stream, and uses the sliding window algorithm to extract the mean and unbiased variance of the extracted temperature and voltage time series to obtain the node baseline feature set.

3. The trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things according to claim 1, characterized in that, Step S2 includes: S2-1, based on the preset privacy weight mapping matrix, performs affine transformation and truncation mapping on the data privacy labels carried in the power grid AI task request to obtain the execution constraints. S2-2, using execution constraints, performs a node privacy support metric evaluation on the underlying privacy computing capability vector of each node in the trusted node matrix to obtain the node compliance mask vector; S2-3, based on the node compliance mask vector, performs set derivation mask filtering on the trusted node matrix to obtain the privacy-compliant node pool.

4. The trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things according to claim 1, characterized in that, Step S3 includes: S3-1, using the legitimate node identifiers in the privacy-compliant node pool as indexes, performs matching and parsing of the node baseline feature set to obtain the candidate node comprehensive state matrix; S3-2, based on the comprehensive state matrix of candidate nodes, takes the minimum inference latency as the objective function, substitutes the execution constraint as a linear penalty term into the multi-objective optimization algorithm to calculate the fitness score and locks the node with the highest score to obtain the target node identifier. S3-3 sends deployment instruction messages to the corresponding physical devices based on the target node identifier, deserializes and loads the AI ​​model within the allocated isolated memory area, and starts a memory hook daemon process to capture stack information and feature tensors at high frequency, continuously generating instance runtime context.

5. The trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things according to claim 1, characterized in that, Step S4 includes: S4-1, based on the target node identifier, performs timestamp window alignment and tensor concatenation on the real-time node operation log and the real-time network jitter sequence to obtain the fused state time series tensor; S4-2, input the fused state temporal tensor into the lightweight temporal prediction model for forward inference and gradient solving to obtain the confidence decay gradient vector; S4-3, the L2 norm of the trust decay gradient vector and the security degradation slope threshold are used to determine if they exceed the limit, and the target node identifier that triggers the alarm is added to the early warning pool to obtain the high-risk node early warning list.

6. The trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things according to claim 1, characterized in that, Step S5 includes: S5-1, based on the trust decay gradient vector corresponding to the alarm node in the high-risk node warning list, extrapolate the difference between the current trust state value and the failure bottom line threshold to obtain the remaining available time window. S5-2, Matching nearby legitimate nodes in the privacy-compliant node pool to obtain the takeover backup node identifier; S5-3, based on the indicator function, perform a spatiotemporal constraint verification on whether the sum of the transmission time of the instance running context and the hot start time of the takeover standby node is less than the remaining available time window. After the verification is passed, establish an encrypted side channel to inject the instance running context into the memory address space of the takeover standby node to obtain the context injection completion status code. S5-4, based on the context injection completion status code, triggers the flow table update instruction of the software-defined network controller, redirects the data packets of the original high-risk node to the takeover standby node, and promotes the takeover standby node identifier to the migration target node identifier.

7. The trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things according to claim 6, characterized in that, S5-1 includes: extrapolating the difference between the current trust status value and the failure threshold using the following formula: in, This represents the current trust status value. This is the failure threshold. For the trust decay gradient vector, Let L2 be the norm of the trust decay gradient vector.

8. The trusted computing power scheduling method for edge intelligent AI applications in the power Internet of Things according to claim 3, characterized in that, S2-3 includes: Based on the node compliance mask vector, the verified edge node matrix is ​​hard-divided, where nodes with a mask value of 1 are assigned to the perfectly compliant node set, and nodes with a mask value of 0 are assigned to the node set to be compensated. To obtain the capacity gap tensor of damaged nodes, gap quantification is performed by calculating the dimension-by-dimensional difference between each node in the compensation node set and the task execution constraints. Based on the capability gap tensor of damaged nodes, for each damaged node in the set of compensating nodes, candidate compensating nodes are searched in the set of perfectly compliant nodes to obtain a set of virtual compliant collaborative nodes. The perfect compliance node set and the virtual compliance collaboration node set are expanded and merged to obtain a privacy compliance node pool.