A health perception and prediction system for computing power networks

By designing a health perception and prediction system in the computing power network, monitoring and analyzing the health of computing devices and network devices in real time, and using deep learning models to predict and resource allocation, the problem that traditional management cannot meet the needs of computing power network is solved, and efficient health management and optimization of the computing power network is achieved.

CN118631682BActive Publication Date: 2025-05-13BEIJING JIAOTONG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410879389.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-02
Publication Date
2025-05-13
Estimated Expiration
2044-07-02

AI Technical Summary

Technical Problem

Traditional computing device management cannot meet the comprehensive monitoring and management needs of computing power networks for network equipment health, affecting the smoothness of data transmission and communication.

Method used

A health awareness and prediction system for computing power networks is designed. Through the collaborative work of the computing resource layer, network resource layer and centralized control layer, the health parameters of computing devices and network devices are collected and analyzed in real time, and the equipment health status is predicted using deep learning models, and resources are uniformly allocated.

Benefits of technology

Real-time health monitoring and prediction of computing power network is realized, providing the optimal computing power network execution path, and improving the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118631682B_ABST
    Figure CN118631682B_ABST
Patent Text Reader

Abstract

The present invention provides a health perception and prediction system for a computing power network. The system includes: a computing resource layer: including a data center network and various computing devices, deploying service nodes of the computing power network, obtaining health parameter information of the computing devices, and sending it to a centralized control layer; a network resource layer: used to deploy network service nodes, perform routing and forwarding of data packets, receive and execute resource configuration strategies and computing tasks issued by the control layer, and send health parameter information of network devices to a centralized control layer; a centralized control layer: used to perceive the health status of the current computing devices and network devices based on the health parameter signals of the current computing devices and network devices, predict the health of the sequence signals of the next cycle, and uniformly allocate computing resources and network resources. The present invention takes network capabilities and computing capabilities into consideration, and combined with a health matrix with probability analysis, can provide the optimal computing power network execution path in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication network technology, and in particular to a health perception and prediction system for a computing power network. Background Art

[0002] In today's era of rapid digitalization and technological development, computing power networks have become one of the key factors in leading innovation and business development. The computing power network covers computing centers and data centers, and is the infrastructure that supports various applications, services, and businesses. With the continuous increase in business needs, the intelligent operation and maintenance system of the computing power network has become a core component to ensure its stable operation and efficient management. The computing power network emphasizes the automatic selection of computing power nodes and transmission paths that meet the needs based on the business flow requirements, comprehensive computing power and network dual-dimensional status information. In this context, higher requirements are placed on the intelligent operation and maintenance systems of data centers and computing centers, requiring more comprehensive and intelligent management and monitoring of computing equipment and network equipment.

[0003] Traditional computing equipment management can no longer meet the needs of today's computing power network, because the health of network equipment is crucial to ensure smooth data transmission and communication. Therefore, the intelligent operation and maintenance system not only focuses on the performance and health of computing equipment, but also needs to care about the status of network equipment to ensure the efficient operation of the entire computing power network system. In this context, advanced monitoring technology has become the key to achieving intelligent operation and maintenance. Through performance monitoring, health analysis, automated maintenance, security monitoring and predictive maintenance, the intelligent operation and maintenance system can respond to and solve problems in real time, improving the stability and reliability of the system. Summary of the invention

[0004] An embodiment of the present invention provides a health perception and prediction system for a computing power network to effectively provide an optimal computing power network execution path.

[0005] In order to achieve the above object, the present invention adopts the following technical scheme.

[0006] A computing network health perception and prediction system, comprising:

[0007] Computing resource layer: includes data center network and various computing devices, deploys service nodes of computing network, obtains health parameter information of computing devices through baseboard management controller or host operating system agent software, and periodically sends the health parameter information of computing devices to the centralized control layer;

[0008] Network resource layer: used to deploy network service nodes, route and forward data packets, receive and execute resource configuration strategies and computing tasks issued by the control layer, and complete resource reservation and occupation through the centralized controller; periodically send the health parameter information of network devices to the centralized control layer;

[0009] Centralized control layer: used to collect network resources and computing resources, perceive the health status of current computing devices and network devices based on the received health parameter signals of current computing devices and network devices, use the relationship between signals in the computing power network and the relationship between signals and health to predict the health of the sequence signals in the next cycle, and uniformly allocate computing resources and network resources.

[0010] Preferably, the computing resource layer includes a data center network and various computing devices, deploys service nodes of the computing power network, obtains health parameter information of the computing devices through the baseboard management controller or the agent software of the host operating system, and periodically sends the health parameter information of the computing devices to the centralized control layer.

[0011] Preferably, the network resource layer is used to cover the network transmission part of the computing power network, deploy network service nodes, provide specific network signals to the upper layer through information notification or information inquiry, receive and execute resource configuration strategies and computing tasks issued by the control layer, and complete resource reservation and occupation through the centralized controller according to the allocation strategy of the centralized control layer; and periodically send the health parameter information of the network equipment to the centralized control layer as feedback for the allocation strategy.

[0012] Preferably, the centralized control layer is composed of a controller, and a computing power scheduling engine is used to collect signals from the computing domain and the network domain. A computing power network health perception model and a computing power network health prediction model, as well as relationship formulas between signals and relationship formulas between signals and health in the computing power network are deployed in the computing power scheduling engine; the computing power network health perception model perceives the health status of the current computing devices and network devices based on the received health parameter signals of the current computing devices and network devices, and the computing power network health prediction model uses the relationship formulas between signals in the computing power network and the relationship formulas between signals and health to predict the health of the sequence signals of the next cycle, and learn the health relationship between signals.

[0013] Preferably, the computing power network health perception model is composed of an encoder module and a feature splicing module based on the Transformer model, wherein the Encoder module includes a multi-head attention module, a forward feedback module, an incremental connection and normalization module, a linear transformation and loss function module, and the Encoder model analyzes and mines the health parameter signals of computing devices and network devices collected by the computing resource layer and the network resource layer, and uses the feature splicing module to splice the computing signal and the communication signal, as well as the correlation or dependency between the signals, and adopts a multi-head attention mechanism to divide the spliced ​​joint signal into several equal parts, and output equal sequences of equal length.

[0014] Preferably, the multi-head attention module in the Encoder module is responsible for extracting effective feature information from the input signal; the forward feedback module uses linear transformation to map the input features to another space, and uses nonlinear activation to introduce nonlinear factors to further process and learn the extracted features; the incremental connection and normalization module is used to use incremental connection and normalization after the multi-head attention module and the forward feedback module, so that the network can learn deeper feature representations; the linear transformation and loss calculation module is used to classify and evaluate the processed signal to obtain the health status of the signal;

[0015] The multi-head attention module uses a multi-segment attention function to weight the input signal sequence, divides the input signal into segments of equal length, and obtains the attention features of each sub-segment. The calculation process of the multi-segment attention function is as follows:

[0016]

[0017] Among them, the dimension of the parameter matrix is ​​Ψ o , Concat(...) represents the concatenation operation of the attention weight results of each concatenated sub-segment, with a dimension of p*p;

[0018] The calculation process of the linear transformation and loss function layer is: the linear change maps the input multi-dimensional calculation and communication signal to one dimension for processing, and the multi-dimensional change is characterized as follows:

[0019] Among them, x represents the output vector of the previous incremental connection and normalization layer, Flatten() is a multidimensional conversion function, and the linearly transformed vector v is processed by the Softmax function to obtain the final health probability function. The Softmax function converts each element of the vector into a value representing probability, so that the sum of all elements of the output vector is 1.

[0020] Preferably, the computing power network health prediction model includes a computing and network signal correlation specification module and a health prediction module;

[0021] The calculation and network signal correlation specification module is used to capture the key characteristics and regularity of the equipment operation status by analyzing and modeling historical data, ensuring that the encoder and decoder can receive the sequence pairs with correlation, and realize the accurate prediction of the next operation cycle sequence signal in the model;

[0022] The health prediction module is used to adopt a Transformer model architecture, in which the encoder encodes the input sequence signal into a hidden representation, and the decoder decodes the encoded representation into an output sequence signal. By encoding and decoding the initial collected signal, the time relationship and dependency relationship between the signals are captured to achieve the prediction of the next period sequence signal.

[0023] Preferably, the calculation and network signal correlation specification module is used to divide the original data set into n equal-length sequence segments according to different sequence lengths, that is:

[0024] Data=(S1,S2,…,S (n-1) , S n )

[0025] Each segment contains data within a certain time window;

[0026] S i =(I1, I2, ..., I i )

[0027] Perform shift and truncation operations on the formed equal-length sequence fragments to obtain the input matrix I of the encoder and decoder encoder and I decoder ;

[0028]

[0029] Alignment operation: Align X and Y to obtain the input sequence pair of the CNHP model:

[0030]

[0031] Among them, S i Represents the input set of the CNHP model, and i represents the index of the sequence pair.

[0032] Preferably, the health prediction module is used to include an encoder-decoder using a self-attention neural network structure, wherein the encoder converts the original sequence {x1, ..., x n} is represented as a feature vector c, and the decoder generates the target sequence {y1, ..., y n};

[0033] The specific process of sequence prediction can be expressed as:

[0034] y=Decoder[Encoder(I encoder ), I decoder ],

[0035] Among them, Encoder(·) and Decoder(·) represent the encoding process and decoding process respectively. encoderrepresents the input vector of the encoder, and y represents the output vector of the decoder;

[0036] With the goal of maximizing the number of successful computing task scheduling and maximizing health, the objective function of the non-deterministic NP problem of resource allocation polynomial complexity is described as:

[0037]

[0038] Where φ(t i ) indicates whether the computing task is scheduled successfully, H(C n , C s ) represents the health of computing resources and network resources, α and β represent the corresponding weights;

[0039] The objective function constraints include: computing and communication capacity constraints, health constraints, reliability constraints, and transmission delay constraints:

[0040] The objective function is solved by using the end-to-end health scheduling transmission algorithm based on D3QN. The specific solution process includes:

[0041] A sensing unit is deployed in the computing engine to obtain the health status of network devices in the wide area network, computing devices in the data center, and network devices;

[0042] The sensing unit records the current computing power and network status, which is defined as:

[0043] S={S network , S′ network , S compute , S′ compute , Q n , Q c},

[0044] S network S′ is the health status of the network equipment calculated by the computing power network health perception model. network is the health status of the network device at the next moment calculated by the computing power network health prediction model, S compute S′ is the health status of the computing device calculated by the current computing power network health perception model. compute Q is the health status of the computing device at the next moment calculated by the computing power network health prediction model. n , Q c They represent the bandwidth requirement and computing resource requirement of the i-th service flow respectively;

[0045] The computing engine performs deterministic transmission scheduling action decision space A. The computing engine performs the decision definition of transmission scheduling and path optimization of business flows:

[0046] A={a0,a1,a2,...,a n}, n is the number of optional paths

[0047] a i represents the path of the i-th item, and A represents the action space;

[0048] The rewards scheme adopted is as follows:

[0049]

[0050] Where psize represents the size of the task, h(S cur ) indicates the comprehensive health of computing and network after computing tasks are assigned.

[0051] It can be seen from the technical solutions provided by the above embodiments of the present invention that the system of the present invention is oriented to the overall optimization goal of the computing network, takes network capacity and computing capacity into consideration, and combines the health matrix with probability analysis to provide the optimal computing network execution path in real time. Through real-time feedback, the network capacity and computing capacity are collaboratively optimized to improve the accuracy of the prediction.

[0052] Additional aspects and advantages of the present invention will be given in part in the following description, which will become obvious from the following description, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0054] Figure 1 A schematic diagram of the structure of a computing power network health perception and prediction system provided by an embodiment of the present invention;

[0055] Figure 2 A schematic diagram of the structure of a computing power network health perception model provided by an embodiment of the present invention;

[0056] Figure 3 A schematic diagram of the structure of a computing power network health prediction model provided by an embodiment of the present invention;

[0057] Figure 4 A schematic diagram of a two-dimensional health assessment and association system of a computing power network provided by an embodiment of the present invention;

[0058] Figure 5 A schematic diagram of a computing power network dynamically allocating resources to computing tasks provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0059] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention.

[0060] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The term "and / or" used herein includes any unit and all combinations of one or more associated listed items.

[0061] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless defined as herein.

[0062] To facilitate understanding of the embodiments of the present invention, several specific embodiments will be further explained below with reference to the accompanying drawings, and each embodiment does not constitute a limitation on the embodiments of the present invention.

[0063] The embodiment of the present invention provides a health perception and prediction system for a computing power network, which integrates network resources and computing resources, associates the health of network links with the health of computing power hardware facilities, and uses a multivariate regression model to establish a two-dimensional correlation model between computing power network signals and the health of computing power hardware facilities. According to the real-time operation results of the two-dimensional correlation model, the prediction results are automatically optimized and managed to improve self-management and self-control, and ensure the service quality of the computing power network.

[0064] The embodiment of the present invention proposes a scheduling problem related to health based on a two-dimensional association model, namely, maximizing the scheduling of computing tasks and improving the health of the computing power network to the greatest extent.

[0065] The structure of a computing power network health perception and prediction system provided by an embodiment of the present invention is as follows: Figure 1 As shown. The computing power network in the embodiment of the present invention is a network that combines network technology with computing power technology. In view of the different urgency of different signals (such as fault signals), the system adopts resource reservation and preemptive network strategies. Resource reservation guarantees bandwidth according to business priority, ensures that data transmission is completed within a specific time, and ensures that resource signals arrive at the dispatch center within a determined time. Preemptive network strategies provide more scalability time guarantees to meet future service needs. The computing power network is a new type of information infrastructure that integrates edge computing, cloud computing nodes, and network resources based on network technologies such as cloud-network integration and SDN (Software Defined Network).

[0066] The main features of the computing power network are: resource abstraction, business assurance, unified management and control, and flexible scheduling; the unified management and control requires that cloud computing, edge computing nodes, and network resources in the network are all managed in a unified manner, and the service will be uniformly scheduled according to specific business needs; the flexible scheduling requires real-time monitoring of business traffic, and dynamic adjustment of computing power resources according to traffic call conditions to ensure the reasonable allocation of various resources in the computing power network; the embodiment of the present invention is designed to monitor the network and computing power dynamic data in real time through a centralized computing power engine and distributed service nodes, and to achieve unified scheduling through a computing power engine. Whether it is to provide the deterministic service quality or to combine computing power with the network, the health perception and prediction system and working method of the computing power network designed in the embodiment of the present invention are required.

[0067] The computing power network health perception and prediction system of the embodiment of the present invention adopts a cloud-edge collaborative architecture to achieve real-time monitoring and data collection of the network and computing nodes in the computing power network, applies preemption and resource reservation technology to meet the transmission of key signals, and achieves real-time monitoring and data collection of the network and computing nodes in the computing power network.

[0068] Specifically, it includes: a computing resource layer composed of multiple data centers, a network resource layer composed of a wide area network, and a centralized control layer composed of a computing power scheduling engine.

[0069] The computing resource layer provides a variety of computing resources.

[0070] The network resource layer provides transmission capabilities and a variety of network transmission technologies. It is responsible for the routing and forwarding of data packets, ensuring that data packets are smoothly transmitted from the source to the destination node, and providing a variety of network transmission technologies.

[0071] The centralized control layer collects network resources and computing resources, perceives and predicts the health of computing resources and network resource-related equipment according to service requirements, allocates resources reasonably according to the availability and performance indicators of the network resource layer, and adjusts the allocation strategy according to real-time data to optimize performance. It provides end-to-end computing task scheduling capabilities, comprehensively considers computing and network resources, and allocates them across the entire network.

[0072] Computing resource layer: includes data center network and various computing devices (such as intelligent computing cabinets and general computing cabinets), deploys service nodes of computing network, and obtains health parameter information of computing devices through the baseboard management controller or the agent software of the host operating system. After the acquisition is completed, the health parameter information of the computing device is periodically sent to the centralized control layer as input information of the computing network health perception model in the centralized control layer to complete the measurement of computing health.

[0073] Network resource layer: used to cover the network transmission part of the computing network. Deploy network service nodes, which can use frame preemption mode to send fault signals as soon as possible, and provide specific network signals to the upper layer through information notification or information inquiry. The network resource layer receives and executes the resource configuration strategy and computing tasks issued by the control layer after the network resource allocation is completed using the end-to-end health scheduling transmission algorithm based on D3QN, and completes the reservation of resources according to the allocation strategy of the centralized control layer through the centralized controller. The network resource layer also periodically sends the health parameter information of the network equipment to the centralized control layer as feedback for the allocation strategy or input for the health model.

[0074] Centralized control layer: mainly composed of controllers, using computing power scheduling engines to collect signals from computing domains and network domains, deploying computing power network health perception models and computing power network health prediction models in the computing power scheduling engine, as well as the relationship between signals in the computing power network and the relationship between signals and health. The computing power network health perception model perceives the health status of the current computing devices and network devices based on the health parameter signals received from the current computing devices and network devices. The computing power network health prediction model uses the relationship between signals in the computing power network and the relationship between signals and health to predict the health of the sequence signals in the next cycle, learn the health relationship between signals, and provide feasible suggestions for the optimization of the computing power network.

[0075] The sequence signals include computing signals and communication signals. The computing signals include: hardware status (computing resource utilization, storage resource utilization) and software status (operating system software and other software status); the communication signals include: historical status of network protocols, and network device port related data (including throughput, queue depth, processing delay, etc.), network status (including delay-bandwidth product, network utilization, etc.).

[0076] The computing power network health perception model uses an encoder model architecture based on deep learning to analyze and mine the health parameter signals of computing devices and network devices collected by the computing resource layer and the network resource layer. The model can autonomously learn and summarize the rules and connections between signals based on key signals such as the operating status of computing power devices, network traffic conditions, and communication conditions between devices, so as to accurately perceive the health status of the computing power network. By learning and comparing historical data, the model can identify potential problems and issue alarms in a timely manner to help operation and maintenance personnel take appropriate measures.

[0077] The computing power network health prediction model uses the deep learning Encoder-Decode model and combines the spatiotemporal relationship between computing power and network sequence signals to build an efficient health prediction model. This model can accurately predict the health change trend of the computing power network in the future, providing important decision-making basis for operation and maintenance personnel.

[0078] Two-dimensional health assessment and correlation system of computing power network: Aims to establish a correlation model between computing power signals and network signals to deeply explore the relationship between them. By using methods such as linear regression, we can more accurately understand the correlation between computing power signals and network signals, and then evaluate the impact of signals on health, and effectively optimize the performance of computing power equipment and network equipment.

[0079] Dynamically allocate the communication resources of the wide area network and the computing resources of the data center to maximize the health of equipment along the way based on maximum flow scheduling.

[0080] The computing power network health perception and prediction system of the embodiment of the present invention applies preemption and resource reservation technology to meet the transmission of critical signals, provides a higher transmission priority, and the routers and switches in the data centers and network nodes of the computing resource layer and the network resource layer have a network resource reservation mechanism and provide frame preemption.

[0081] The routers and switches in the data center have a network resource reservation mechanism that can guarantee bandwidth according to business priorities and ensure that data transmission is completed within a specific time. Through the dynamic deployment of computing resources in the data center, the required computing time can be dynamically calculated according to the user's computing needs. By combining the transmission time of the routers with resource reservation capabilities in the data center and the planned computing time, the total time from the user data entering the data center to the computing results being transmitted out of the data center can be obtained. Therefore, the network inside the data center can ensure the certainty of the data.

[0082] Different signals have different priorities, and different priorities correspond to different resource reservation strategies and preemption strategies. Among these strategies, fault signals have the highest priority, common resource signals have lower priorities, and the spare priorities in the middle are used for other signals that may appear in the future. This priority setting ensures that the system can allocate and process resources appropriately according to their importance and urgency when processing various signals.

[0083] The structure of the computing power network health perception model provided by the embodiment of the present invention is as follows: Figure 2 The embodiment of the present invention deploys a computing power network health perception model in the computing power scheduling engine. The computing power network health perception model is based on the encoder model architecture of deep learning, analyzes and mines the signals collected by the network and computing power equipment, and implements health evaluation of the computing power network.

[0084] More specifically, the computing power network health perception model is composed of an encoder module and a feature splicing module based on the Transformer model. The main function of the feature splicing module is to splice together multiple signals such as computing signals and communication signals as correlation relationships to form a complete input tensor for model processing. The encoder consists of a multi-head attention module, a forward feedback module, an incremental connection and normalization module, a linear transformation and a loss function module.

[0085] More specifically, the feature concatenation module obtains the computational signal and the communication signal. In order not to affect the correlation or dependency between the signals, here we simply concatenate multiple signals in order, and the relationship between the signals still exists in the concatenated vector. Since the Encoder adopts a multi-head attention mechanism, the concatenated joint signal is equally divided into several parts and outputs equal-length equal sequences. The multi-head attention module in the Encoder module is responsible for extracting effective feature information from the input signal. These features can reflect the important characteristics and change trends of the signal, which helps the subsequent model to better understand and analyze the signal. The forward feedback module is used to further process and learn the extracted features to improve the model's ability to understand and represent the signal. The incremental connection and normalization modules can help the model better handle the correlation and variability between signals, and improve the stability and generalization ability of the model. The linear transformation and loss calculation modules are used to classify and evaluate the processed signals, derive the health status of the signal, and provide feedback and suggestions to the user or system.

[0086] More specifically, the principle and calculation process of the multi-head attention module are as follows: First, the input signal sequence is weighted using the attention function to obtain the attention features of each sub-segment. Multi-segment parallel computing is used to strengthen the association between computing power and long-distance sampling points in the network long signal scenario. In the attention weight calculation, multi-dimensional linear transformation is performed on the query, key, and value vectors. The gradient stability is improved by introducing a scaling factor for normalization.

[0087] More specifically, the present invention adopts a segmented attention mechanism. This mechanism divides the input signal into segments of equal length and speeds up the training speed by parallel processing, thereby strengthening the connection between segments and extracting the overall information. The calculation process of the multi-segment attention function is as follows:

[0088]

[0089] Among them, the dimension of the parameter matrix is ​​Ψ o , Concat(...) represents the concatenation operation of the attention weight results of each concatenated sub-segment, with a dimension of p*p, where the number of multi-segment self-attention functions is 16.

[0090] The core of the forward feedback module is linear transformation (weighted summation) and nonlinear activation. Linear transformation is responsible for mapping the input features to another space, while nonlinear activation introduces nonlinear factors, allowing the network to capture more complex feature relationships.

[0091] More specifically, the principle of the incremental connection and standardization module is: incremental connection and standardization are used after the multi-head attention module and the forward feedback module, mainly to solve the gradient vanishing and gradient exploding problems in the deep network, so that the network can learn deeper feature representations. At the same time, standardization can make the computing power and network representation more unified, accelerate the convergence speed of the model, and improve stability.

[0092] More specifically, the principle and calculation process of the linear transformation and loss function layer are as follows: Linear change mainly maps the input multi-dimensional computing and communication signals to one dimension for processing, which is conducive to calculating the corresponding loss function. The representation of multi-dimensional change is:

[0093] υ=Flatten(x),

[0094] Among them, x represents the output vector of the previous incremental connection and normalization layer, and Flatten() is a multidimensional conversion function. The linearly transformed vector v is processed by the Softmax function to obtain the final health probability function. The Softmax function is usually used for multi-classification problems. It can convert each element of the vector into a value representing probability, so that the sum of all elements of the output vector is 1. The state perception model is used to classify the health of the input signal, so the objective function of the model can be defined as the cross entropy loss function, which is specifically defined as follows:

[0095]

[0096] in, and Represents the input sequence signal and the corresponding label. The added part is the correction of the regularization term, and λ is the attenuation factor of the regularization term.

[0097] The structure of a computing power network health prediction model provided by an embodiment of the present invention is as follows: Figure 3 As shown. The computing power network health perception model based on Encoder-Decoder is deployed in the computing power scheduling engine. Combining the spatiotemporal relationship between computing power and network sequence signals, an efficient health prediction model is constructed. Specifically, the computing power network health prediction model includes a computing and network signal correlation specification module and a health prediction module. The correlation specification module is designed for the prediction of the equipment operation status sequence. By analyzing and modeling historical data, it captures the key characteristics and regularity of the equipment operation status, ensures that the encoder and decoder can receive sequence pairs with correlation, and realizes accurate prediction of the next operation cycle sequence signal in the model. The prediction module adopts the Transformer model architecture, in which the encoder is responsible for encoding the input sequence signal into a hidden representation, and the decoder is responsible for decoding the encoded representation into an output sequence signal. By encoding and decoding the initial acquisition signal, the model can effectively capture the temporal relationship and dependency between the signals, thereby realizing accurate prediction of the next cycle sequence signal. The entire prediction process is end-to-end, and the prediction mode of the signal can be learned directly from the original data without manual intervention or specifying specific rules.

[0098] Specifically, the sequence prediction data shaping module includes two key steps, namely "sequence pair" generation and sub-segment segmentation. The purpose of these two steps is to ensure that the encoder and decoder can receive relevant sequence pairs and achieve accurate prediction of the next running cycle sequence signal in the model. The following is a detailed description of these steps:

[0099] First, the original data set is divided into n equal-length sequence segments according to different sequence lengths, namely:

[0100] Data=(S1,S2,…,S (n-1) , S n )

[0101] Then, shift and truncate the formed equal-length sequence segments to form the input matrix I of the encoder and decoder encoder and I decoder .

[0102] 1. "Sequence pair" generation: In the process of sequence pair generation, the main consideration is to establish sequence pairs with interrelated characteristics. This is done through the following steps:

[0103] Dataset segmentation: The original dataset is segmented into equal-length sequence segments according to different sequence lengths. This ensures that each segment contains data within a certain time window.

[0104] S i =(I1, I2, ..., I i )

[0105] Shift and truncation: Perform shift and truncation operations on the formed equal-length sequence fragments to obtain the input matrix I of the encoder and decoder encoder and I decoder ;

[0106]

[0107] Alignment operation: Align X and Y to obtain the input form of the CNHP model.

[0108] Input sequence pairs:

[0109] Among them, S i represents the input set of the CNHP model, and i represents the index of the sequence pair. This ensures that the model can receive relevant sequence pairs generated based on the sequence signal of the previous run cycle.

[0110] 2. Sub-segment segmentation: This step involves segmenting the sequence segments. The specific process is similar to the segmentation process described previously and will not be repeated here.

[0111] Specifically, the function and process reasoning formula of the health prediction module include encoder-decoder: First, the encoder converts the original sequence {x1, ..., x n} is represented as a feature vector c; then, the decoder generates the target sequence {y1,…,y n The self-attention neural network is used as the model structure of the encoder and decoder, and the computational complexity of the model is effectively reduced through parallel computing, while significantly improving its computational efficiency.

[0112] The specific process of sequence prediction can be expressed as:

[0113] y=Decoder[Encoder(I encoder ), I decoder ],

[0114] Where Encoder(·) and Decoder(·) represent the derivation process described in Section 2.3.2, respectively. encoder represents the input vector of the encoder, and y represents the output vector of the decoder.

[0115] 2) Linear transformation: Root Mean Squared Error (RMSE) and Mean Absolute Error (MAE) are used to evaluate the prediction performance of the CNHP model. The specific definitions of RMSE and MAE are as follows:

[0116]

[0117] Among them, y i and They respectively represent the actual operating state and the predicted operating state corresponding to the sequence signal samples, and N represents the length of the sequence signal.

[0118] A schematic diagram of a two-dimensional health assessment and association system of a computing power network provided by an embodiment of the present invention is as follows Figure 4 As shown, it includes: taking the original signals including computing signals and communication signals, as well as the health of the current state and the health of the next state as input, using a linear regression model to establish the relationship between signals in the computing power network, and the relationship between signals and health; after training with the data set, adapting appropriate parameters, and establishing the relationship between the computing power in the computing power network and the signal of the network, as well as the relationship between the computing power and the health.

[0119] More specifically, the input signals include: computing signals including computing resource utilization, storage resource utilization and software status, communication signals including network protocol history status, throughput, queue depth, processing delay, delay-bandwidth product and network utilization, health including the health of the current state and the health of the next cycle. Output signals include the relationship between signals and the relationship between signals and health.

[0120] Figure 5 A schematic diagram of a computing power network dynamically allocating resources to computing tasks is provided in an embodiment of the present invention to maximize the health of the integrated equipment.

[0121] In order to jointly consider the dynamic resource allocation state and health in the computing network, the present invention proposes a resource allocation NP problem, namely maximizing the number of successful computing task scheduling and maximizing health. The objective function of the resource allocation NP (Non-deterministic Polynomial) problem of the present invention is described as:

[0122]

[0123] Where φ(t i ) indicates whether the computing task is scheduled successfully, H(C n , C s ) represents the health of computing resources and network resources. α, β represent the corresponding weights.

[0124] Constraints include:

[0125] Computing and communication capacity limitations: Link transmission resources are always less than link capacity, and computing resources are always less than computing capacity.

[0126] Health limit: The health of computing resources and network resource devices through which the transmission passes is always greater than 80 points.

[0127] Reliability constraints: The packet loss rate of each computing task is always less than 1%.

[0128] Transmission delay constraint: The acceptable delay of each computing task is always lower than the actual overall delay.

[0129] Specifically, in order to solve the objective function of the above resource allocation NP problem, the present invention proposes an end-to-end health scheduling transmission algorithm based on D3QN. The scheduling transmission algorithm model is deployed on a centralized controller, obtains device parameter information of the computing resource layer and the network resource layer, analyzes the requirements of the computing tasks, and requires maximizing the health of the overall computing resources and network resources on the basis of meeting the task requirements. The specific algorithm processing process is as follows:

[0130] (1) Deployment of perception units

[0131] Deploy perception units in the computing engine to obtain the health status of network devices in the wide area network, computing devices in the data center, and network devices.

[0132] (2) The sensing unit records the current computing power and network status, which is defined as:

[0133] S={S network , S′ network , S compute , S′ compute , Q n , Q c},

[0134] S network S′ is the health status of the network equipment calculated by the computing power network health perception model. network is the health status of the network device at the next moment calculated by the computing power network health prediction model, S compute S′ is the health status of the computing device calculated by the current computing power network health perception model. compute Q is the health status of the computing device at the next moment calculated by the computing power network health prediction model. n , Q c They represent the bandwidth requirement and computing resource requirement of the i-th business flow respectively.

[0135] (3) The computing engine performs deterministic transmission scheduling action decision space A. The computing engine performs the decision of transmission scheduling and path optimization of business flows and can be defined as:

[0136] A={a0,a1,a2,...,a n}, n is the number of optional paths

[0137] a i represents the path of the i-th item, and A represents the action space.

[0138] (4) Reward function R

[0139] The reward is used to evaluate the quality of the actions selected by the greedy algorithm, and to encourage the agent to explore the optimal resource allocation scheme with the number of successful scheduling calculation tasks and the maximum health as the guide. The reward scheme adopted in the embodiment of the present invention is as follows:

[0140]

[0141] Where psize represents the size of the task, h(S cur ) indicates the comprehensive health of computing and network after computing tasks are assigned.

[0142] In summary, an embodiment of the present invention discloses a health perception and prediction system for a computing power network, which adopts resource reservation and frame preemption technology to improve the timeliness of information. The system overcomes the problem of difficulty in manual judgment by introducing a health perception and prediction algorithm based on Encoder and Encoder-Decoder, reduces subjectivity and accuracy, and improves the accuracy and reliability of health judgment. By deploying network intelligent computing service nodes on network devices in the wide area network and computing power network service nodes in the data center, the device status information is obtained in real time, and the network capacity and computing power are collaboratively optimized through real-time feedback to improve the accuracy of the prediction. A two-dimensional correlation model of network and computing power health is established in the computing power scheduling engine, and the network capacity and computing power are comprehensively considered, which can provide the optimal computing power network execution path in real time.

[0143] The system of the present invention overcomes the problem of difficulty in manual judgment, reduces subjectivity and accuracy, and improves the accuracy and reliability of judgment. Aiming at the overall optimization goal of the computing power network, the network capacity and computing power are comprehensively considered, and combined with the health matrix with probability analysis, the optimal computing power network execution path can be provided in real time. Through real-time feedback, the network capacity and computing power are coordinated and optimized to improve the accuracy of the prediction.

[0144] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of an embodiment, and the modules or processes in the accompanying drawings are not necessarily required to implement the present invention.

[0145] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention or certain parts of the embodiments.

[0146] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0147] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A computing network health perception and prediction system, characterized in that: include: Computing resource layer: includes data center network and various computing devices, deploys service nodes of computing network, obtains health parameter information of computing devices through baseboard management controller or host operating system agent software, and periodically sends the health parameter information of computing devices to the centralized control layer; Network resource layer: used to deploy network service nodes, route and forward data packets, receive and execute resource configuration strategies and computing tasks issued by the control layer, and complete resource reservation and occupation through the centralized controller; periodically send the health parameter information of network devices to the centralized control layer; Centralized control layer: used to collect network resources and computing resources, perceive the health status of current computing devices and network devices according to the received health parameter signals of current computing devices and network devices, use the relationship between signals in the computing network and the relationship between signals and health to predict the health of the sequence signal in the next cycle, and uniformly allocate computing resources and network resources; The centralized control layer is composed of a controller, which uses a computing power scheduling engine to collect signals from the computing domain and the network domain, and deploys a computing power network health perception model and a computing power network health prediction model in the computing power scheduling engine, as well as a relationship between signals in the computing power network and a relationship between signals and health; The computing power network health perception model perceives the health status of the current computing devices and network devices based on the received health parameter signals of the current computing devices and network devices. The computing power network health prediction model uses the relationship between signals in the computing power network and the relationship between signals and health to predict the health of the sequence signals in the next cycle and learn the health relationship between signals. The computing power network health prediction model includes a computing and network signal correlation specification module and a health prediction module; The calculation and network signal correlation specification module is used to capture the key characteristics and regularity of the equipment operation status by analyzing and modeling historical data, ensuring that the encoder and decoder can receive the sequence pairs with correlation, and realize the accurate prediction of the next operation cycle sequence signal in the model; The health prediction module adopts the Transformer model architecture, in which the encoder encodes the input sequence signal into a hidden representation, and the decoder decodes the encoded representation into an output sequence signal. By encoding and decoding the initial acquisition signal, the time relationship and dependency relationship between the signals are captured, and the prediction of the next cycle sequence signal is realized; The health prediction module includes an encoder-decoder using a self-attention neural network structure. The encoder converts the original sequence {x1, ..., x n } is represented as a feature vector c, and the decoder generates the target sequence {y1, ..., y n }; The specific process of sequence prediction is expressed as: y=Decoder[Encoder(I encoder ),I decoder ] Among them, Encoder(·) and Decoder(·) represent the encoding process and decoding process respectively. encoder represents the input vector of the encoder, and y represents the output vector of the decoder; With the goal of maximizing the number of successful computing task scheduling and maximizing health, the objective function of the non-deterministic NP problem of resource allocation polynomial complexity is described as: Where φ(t i ) represents the computation task t i Whether the scheduling is successful, H(C n ,C s ) represents computing resources C n With network resources C s The health of , α, β represent the corresponding weights, and |T| represents the number of computing tasks; The objective function constraints include: computing and communication capacity constraints, health constraints, reliability constraints, and transmission delay constraints; The objective function is solved by using the end-to-end health scheduling transmission algorithm based on D3QN. The specific solution process includes: A sensing unit is deployed in the computing engine to obtain the health status of network devices in the wide area network, computing devices in the data center, and network devices; The sensing unit records the current computing power and network status, which is defined as: S={S network ,S′ network ,S compute ,S′ compute ,Q n ,Q c }, Among them, S network S′ is the health status of the network equipment calculated by the computing power network health perception model. network is the health status of the network device at the next moment calculated by the computing power network health prediction model, S compute S′ is the health status of the computing device calculated by the current computing power network health perception model. compute Q is the health status of the computing device at the next moment calculated by the computing power network health prediction model. n , Q c They represent the bandwidth requirement and computing resource requirement of the i-th service flow respectively; The computing engine performs deterministic transmission scheduling action decision space A. The computing engine performs the decision definition of transmission scheduling and path optimization of business flows: A={a0,a1,a2,…,a n }, n is the number of optional paths a i represents the path of the i-th item, and A represents the action space; The rewards scheme adopted is as follows: Where psize represents the size of the task, h(S cur ) indicates the comprehensive health of computing and network after computing tasks are assigned.

2. The system according to claim 1, characterized in that The network resource layer is used to cover the network transmission part of the computing network, deploy network service nodes, provide specific network signals to the upper layer through information notification or information inquiry, receive and execute resource configuration strategies and computing tasks issued by the control layer, and complete resource reservation and occupation through the centralized controller according to the allocation strategy of the centralized control layer; it also periodically sends the health parameter information of the network equipment to the centralized control layer as feedback for the allocation strategy.

3. The system according to claim 2, characterized in that The computing power network health perception model is composed of an Encoder module and a feature splicing module based on the Transformer module. The Encoder module includes a multi-head attention module, a forward feedback module, an incremental connection and normalization module, a linear transformation and loss function module. The Encoder module analyzes and mines the health parameter signals of computing devices and network devices collected by the computing resource layer and the network resource layer, and uses the feature splicing module to splice the computing signal and the communication signal, as well as the correlation or dependency between the signals. The multi-head attention mechanism is used to divide the spliced ​​joint signal into several equal parts, and output equal-length sequences.

4. The system according to claim 3, characterized in that The multi-head attention module in the encoder module is responsible for extracting effective feature information from the input signal; the forward feedback module uses linear transformation to map the input features to another space, and uses nonlinear activation to introduce nonlinear factors to further process and learn the extracted features; the incremental connection and normalization module is used to use incremental connection and normalization after the multi-head attention module and the forward feedback module, so that the network can learn deeper feature representations; the linear transformation and loss function module is used to classify and evaluate the processed signal to obtain the health status of the signal; The multi-head attention module uses a multi-segment attention function to weight the input signal sequence, divides the input signal into segments of equal length, and each segment pays attention to different aspects of information through the self-attention mechanism to obtain the weight label of the segment signal, obtains the attention feature of each sub-segment, and finally merges the attention feature outputs of all sub-segments. The calculation process of the multi-segment attention function is as follows: Among them, the dimension of the parameter matrix is ​​Ψ o ; The calculation process of the linear transformation and loss function module is as follows: the linear transformation maps the input multi-dimensional calculation and communication signal to one-dimensional processing, and the multi-dimensional transformation is characterized as follows: v=Flatten(x), Where x represents the output vector of the previous incremental connection and normalization layer, Flatten(·) is a multidimensional transformation function, and the linearly transformed vector v is processed by the Softmax function to obtain the final health probability function. The Softmax function converts each element of the vector into a value representing probability, so that the sum of all elements of the output vector is 1.

5. The system according to claim 1, characterized in that The calculation and network signal correlation specification module is used to divide the original data set into n equal-length sequence segments according to different sequence lengths, namely: Data=(S1,S2,…,S (n-1) ,S n ) Each segment contains data within a certain time window; S i =(I1,I2,…,I i ) Perform shift and truncation operations on the formed equal-length sequence fragments to obtain the input vector I of the encoder and decoder encoder and I decoer ; Alignment operation: Align X and Y to obtain the input sequence pair of the CNHP model: Among them, S i Represents the input set of the CNHP model, which includes the encoder input and decoder input composed of the original data, and i represents the index of the sequence pair.