Resource monitoring and predicting system for computing power network in power system
By deploying a multi-head self-attention mechanism model on node terminals and server sides in the power system and combining it with cloud training, the blind spots and lag problems of traditional monitoring methods are solved, and real-time, accurate monitoring and anomaly prediction of computing network resources are realized.
Patent Information
- Application Number
- CN202510799504.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional manual monitoring and static threshold alarm methods are insufficient to capture minute changes in computing resources in real time, resulting in blind spots and lags in the monitoring of computing networks in power systems, making it impossible to detect anomalies in a timely manner.
By deploying node terminals in the power system to collect resource data, and using a multi-head self-attention mechanism model on the server side for real-time monitoring and prediction, the model is centrally trained in the cloud and updated parameters are distributed to achieve real-time and accurate monitoring of computing network resources.
It enables real-time and accurate monitoring of computing network resources, timely detection of anomalies, improved prediction accuracy, and reduced power grid operation interruptions or failures caused by resource unavailability.
Smart Images

Figure CN120914970A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of resource management, and particularly relates to a resource monitoring and prediction system for a computing power network in a power system. BACKGROUND
[0002] With the acceleration of the digital transformation of the power system, the computing power network plays a key role in the stable operation, intelligent scheduling, fault analysis, etc. of the power system. The power system covers multiple links such as power generation, power transmission, power transformation, power distribution and power utilization, and the operation of equipment, data collection and analysis and decision-making of each link are highly dependent on computing power resources.
[0003] However, the computing power nodes in the power system are widely and complexly distributed, and the resource state dynamically changes, so that the traditional manual monitoring and simple resource management mode cannot meet the efficient and accurate operation and maintenance requirements. In the field of resource monitoring of the computing power network, a static threshold alarm is currently used, but this method has a monitoring blind area and hysteresis, and cannot capture the small changes of computing power resources in real time. SUMMARY
[0004] In view of the problem that the static threshold alarm in the prior art cannot capture the small changes of computing power resources in real time, the application provides a resource monitoring and prediction system for a computing power network in a power system, which can realize real-time and accurate monitoring of computing power network resources, and can timely discover abnormalities to provide protection for the stable operation of the power system. The specific technical scheme is as follows:
[0005] A resource monitoring and prediction system for a computing power network in a power system, comprising:
[0006] A plurality of node terminals are deployed on each computing power node of the power system, used to collect resource data of the computing power node, receive control instructions of operation configuration, and perform collection actions and collection tasks according to the control instructions;
[0007] A server end is deployed in a management and control platform of the power system, used to issue control instructions of operation configuration, receive resource data, monitor the resource state of each node terminal, construct a historical resource state information sequence according to the resource data, load the issued model parameters, run a resource prediction model based on a multi-head self-attention mechanism based on the historical resource state information sequence, output a predicted resource state, and monitor abnormal conditions of the predicted resource state;
[0008] A cloud end is used to centrally train a resource prediction model based on a multi-head self-attention mechanism, output updated model parameters, and issue the updated model parameters.
[0009] Preferably, the node terminal comprises:
[0010] A data collection unit is configured to collect resource data of the computing nodes, the resource data including CPU, memory, GPU, network I / O and disk I / O.
[0011] A monitoring management unit is configured to receive a control instruction of a running configuration issued by a server end, and update a collection action and a collection task of the data collection unit according to the control instruction.
[0012] Preferably, the monitoring management unit includes:
[0013] A collection action management module is configured to manage starting, stopping and restarting of the data collection unit according to the control instruction issued by the server end.
[0014] A collection task management module is configured to receive the control instruction issued by the server end, and update a collection index and a collection frequency of the data collection unit.
[0015] Preferably, the server end includes:
[0016] A data storage unit is configured to receive the resource data and store the resource data based on time sequence.
[0017] A resource prediction unit is configured to construct a historical resource state information sequence according to the resource data, load model parameters issued by a cloud end, run a resource prediction model based on a multi-head self-attention mechanism based on the historical resource state information sequence, and output a predicted resource state.
[0018] A resource monitoring unit is configured to monitor resource states of each node terminal according to the resource data, and monitor abnormal conditions of the predicted resource state.
[0019] A configuration management unit is configured to issue a control instruction of a running configuration to the node terminal according to a prediction result of the resource prediction unit and a monitoring result of the resource monitoring unit.
[0020] Preferably, the server end further includes:
[0021] A visualization unit is configured to display the prediction result and the monitoring result of the resource prediction unit and the resource monitoring unit.
[0022] Preferably, the server end further includes:
[0023] An alarm unit is configured to generate alarm information based on a preset alarm rule according to the prediction result of the resource prediction unit and the monitoring result of the resource monitoring unit.
[0024] Preferably, the data storage unit is connected to the cloud end, and is configured to back up the resource data stored based on time sequence to the cloud end.
[0025] Preferably, the cloud is specifically used for:
[0026] According to the resource data backed up by the data storage unit, the historical state information sequence of each resource of each computing power node is constructed;
[0027] According to the historical state information sequence of each resource, the first resource vector is obtained by attention aggregation of the sampling sequence of the same resource, and the second resource vector is obtained by aggregation of the first resource vectors of different resources under the same interval.
[0028] The second resource vector is input into a resource prediction model for training, and when the model training converges, the model parameters are updated.
[0029] Compared with the prior art, the beneficial effects of the present application are:
[0030] The resource monitoring and prediction system of the computing power network in the power system of the present application collects resource data by deploying node terminals at each computing power node of the power system, issues instructions from the server end of the management and control center, monitors resource states, constructs historical sequences, and predicts resource states using a multi-head self-attention mechanism model. The cloud centrally trains the model and issues updated parameters. The present application realizes real-time and accurate monitoring of the resources of the computing power network, and can timely detect abnormalities. The prediction model based on the multi-head self-attention mechanism fully excavates the time sequence and correlation characteristics of the resource data, and improves the prediction accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0031] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual proportions.
[0032] Figure 1 A resource monitoring and prediction system for a computing power network in a power system according to the present application.
[0033] Figure 2 A resource monitoring and prediction system for a computing power network in a power system according to the present application. DETAILED DESCRIPTION
[0034] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0035] It should be understood that the terms "include" and "comprise" as used in the specification, indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0036] It should also be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0037] It should be further understood that the term "and / or" used in the present application specification means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0038] The following examples are based on Figure 1 and Figure 2 .
[0039] The embodiments of the present application provide a resource monitoring and prediction system for a computing power network in a power system, comprising:
[0040] A plurality of node terminals are deployed on each computing power node of the power system, for collecting resource data of the computing power node, receiving control instructions of operation configuration, and performing collection actions and collection tasks according to the control instructions;
[0041] A server end is deployed in a management and control center of the power system, for issuing control instructions of operation configuration, receiving resource data, monitoring resource states of each node terminal, constructing a historical resource state information sequence according to the resource data, loading model parameters issued, running a resource prediction model based on a multi-head self-attention mechanism based on the historical resource state information sequence, outputting predicted resource states, and monitoring abnormal conditions of the predicted resource states;
[0042] A cloud end is used for centralized training of the resource prediction model based on the multi-head self-attention mechanism, outputting updated model parameters, and issuing.
[0043] The node terminals on each computing power node are connected to the server end deployed in the management and control center of the power system mainly through the communication network inside the power system. The node terminals and the server end are based on wired communication mode of Ethernet, such as transmitting data through optical fiber communication link. Since the resource data of the computing power node is large, reliable network connection is needed to realize efficient collection and instruction reception. For example, in a large power data center, each computing power node is connected to the server of the management and control center through a high-speed fiber channel, for transmitting operation configuration control instructions and resource data.
[0044] The server end and the cloud end are also connected by a wide area network (WAN) connection mode, through a dedicated Internet line or a wide area communication network built in the power system. This enables the cloud to remotely receive the data sent by the server end for training the resource prediction model based on the multi-head self-attention mechanism, and after the model training is completed in the cloud, the updated model parameters can be downloaded to the server end through the network. For example, the management and control center server of the power system is connected to the server cluster of the cloud through the dedicated line network provided by the telecom operator, realizing the uploading and downloading of model parameters.
[0045] In this embodiment, multiple node terminals are distributed on each computing power node of the power system, and the node terminals collect resource data of the computing power nodes according to a preset collection period or according to a specific operation configuration control instruction issued by the server end. The resource data includes but is not limited to CPU usage, memory occupation, network bandwidth usage, storage space usage, etc.
[0046] In this embodiment, the server end as the management and control center first issues a control instruction of operation configuration to the node terminal. The control instruction can be used to set the time interval of collection, the specific resource index type of collection, etc. After receiving the control instruction, the node terminal executes the corresponding collection task according to the instruction requirement, and sends the collected resource data back to the server end. The server end will receive the resource data sent by each node terminal in real time, and store and process it. At the same time, the server end will monitor the resource state of each node terminal, and construct a historical resource state information sequence according to the received resource data. This historical sequence contains the resource usage of each computing power node at different time points. The server end loads the parameters of the resource prediction model based on the multi-head self-attention mechanism issued by the cloud. Then, taking the constructed historical resource state information sequence as input, the resource prediction model is run. The multi-head self-attention mechanism can capture the long-term dependence relationship and complex patterns of resource data in time series, so as to more accurately predict the resource state of each computing power node in the future. After the server end outputs the predicted resource state, it will monitor the prediction result. When it is found that the predicted resource state is abnormal, such as the resource usage of a computing power node is expected to rise sharply in a short time and exceed its processing capacity, the server end will trigger the corresponding alarm mechanism.
[0047] The server end monitors the resource state of the node terminal in real time, and monitors the abnormal situation of the predicted resource state based on the output of the prediction model. Potential resource bottlenecks or hidden troubles can be predicted in advance. For example, if the prediction model finds that the memory usage of a certain computing power node will continue to rise rapidly in a short time in the future and exceed its maximum capacity, the system can issue an early warning to remind the operation and maintenance personnel to take measures, reducing the problems such as power grid operation interruption or power trading system failure caused by unavailable computing power resources.
[0048] In this embodiment, the cloud end collects the historical resource data and other information of each computing power node sent by the server end, and continuously trains and optimizes the resource prediction model based on the multi-head self-attention mechanism. After training, the cloud end outputs the updated model parameters, and distributes the updated model parameters to the server end to improve the accuracy of resource prediction of the server end.
[0049] Specifically, in a preferred embodiment of the present application, the node terminal comprises:
[0050] The data acquisition unit is configured to acquire resource data of the computing power node, and the resource data comprises CPU, memory, GPU, network I / O and disk I / O.
[0051] The monitoring management unit is configured to receive a control instruction of a running configuration distributed by the server end, and update the acquisition action and acquisition task of the data acquisition unit according to the control instruction.
[0052] The data acquisition unit and the monitoring management unit are connected through an internal communication interface. The interface can be a connection based on a hardware bus, or communication through an internal network. The data acquisition unit is responsible for acquiring resource data of the computing power node, including CPU, memory, GPU, network I / O and disk I / O, etc.
[0053] In specific implementation, the data acquisition unit is installed on each computing power node, and interacts with the operating system and hardware interface of the node to acquire resource data of the computing power node. For example, the CPU usage is obtained through the performance counter provided by the operating system; the memory usage is obtained through the memory management interface provided by the operating system to obtain the usage and remaining amount of the memory; the GPU usage is obtained through the interface provided by the GPU driver to obtain the usage of the GPU and the memory occupation; the network I / O is obtained through the network interface to obtain the sending and receiving rate of the network; and the disk I / O is obtained through the disk management interface to obtain the read / write rate and I / O waiting time of the disk.
[0054] Specifically, the monitoring management unit comprises:
[0055] The collection action management module is configured to manage the start, stop and restart of the data collection unit according to the control instruction issued by the server end.
[0056] The collection task management module is configured to receive the control instruction issued by the server end and update the collection index and collection frequency of the data collection unit.
[0057] The collection action management module and the collection task management module in the monitoring management unit communicate with each other through a software interface. The two modules jointly receive the control instruction from the server end and manage the start, stop, restart of the data collection unit and the update of the collection index and collection frequency according to the content of the control instruction from the server end.
[0058] In a specific implementation, the collection action management module receives the control instruction issued by the server end, parses the instruction content, starts, stops or restarts the data collection unit according to the instruction requirement. For example, if the server end issues an instruction to stop collection, the collection action management module sends a stop signal to the data collection unit, and the data collection unit stops the collection action after receiving the signal; or the collection action management module can also set the running state of the data collection unit according to the instruction and enter the sleep mode to save resources. The collection task management module receives the control instruction issued by the server end, parses the content of the instruction about the collection task. According to the instruction, the collection index and collection frequency of the data collection unit are updated. For example, if the server end requires to increase the collection index of GPU usage rate, the collection task management module will pass this requirement to the data collection unit, and the data collection unit will start collecting GPU usage rate data. Or the collection task management module can also adjust the collection frequency according to the instruction, for example, from collecting once every minute to collecting once every second, to adapt to different monitoring requirements.
[0059] Specifically, in a preferred embodiment of the present application, the server end comprises:
[0060] The data storage unit is configured to receive resource data and store the resource data based on time sequence.
[0061] Using a time sequence database can realize the storage and query of time sequence data. The time sequence database can index and manage data according to time stamp, and can query and analyze resource data according to time range; the data storage unit creates an independent time sequence data table for each computing power node to record the resource data of each node and the collection time.
[0062] The resource prediction unit is configured to construct a historical resource state information sequence according to the resource data, load model parameters issued by the cloud end, run a resource prediction model based on a multi-head self-attention mechanism based on the historical resource state information sequence, and output a predicted resource state.
[0063] The resource prediction unit receives parameters of a resource prediction model based on a multi-head self-attention mechanism issued by the cloud, takes a historical resource state information sequence as input, runs the resource prediction model, and outputs a predicted resource state, including a prediction of resource usage of each computing power node in a future period of time.
[0064] The resource monitoring unit is configured to monitor resource states of each node terminal according to resource data and monitor abnormal conditions of the predicted resource state.
[0065] The predicted results output by the resource prediction unit are compared and analyzed with actual resource data to monitor whether the predicted resource state is abnormal. For example, if the predicted CPU usage rate will remain within a reasonable range in a future period of time, but the actual monitoring finds that the CPU usage rate suddenly and sharply rises and exceeds the predicted range, it is determined that the predicted resource state is abnormal.
[0066] The configuration management unit is configured to issue a control instruction for running configuration of the node terminal according to the predicted results of the resource prediction unit and the monitoring results of the resource monitoring unit.
[0067] According to the prediction and monitoring results, a running configuration control instruction for the node terminal is generated. For example, if it is predicted that the memory occupation of a certain node will continuously increase and may exceed the capacity in a future period of time, the configuration management unit will generate an instruction to require the node terminal to reduce memory usage tasks or increase memory resources. The generated control instruction is issued to the monitoring management unit of the node terminal through a communication network to realize dynamic configuration and management of the node terminal.
[0068] The visualization unit is configured to display the predicted results and monitoring results of the resource prediction unit and the resource monitoring unit.
[0069] The visualization unit displays the predicted results of the resource prediction unit in the form of intuitive charts. For example, a line chart is used to display the predicted trend of the CPU usage rate, memory occupation and other resources of each computing power node in a future period of time. The visualization unit displays the monitoring results of the resource monitoring unit, including the current actual resource state of each node terminal and the abnormal conditions of the predicted resource state. For example, a dashboard is used to display the real-time usage rate of CPU, memory, GPU and other resources of each node, and different colors are used to mark normal and abnormal states.
[0070] In the system, the node terminal provides resource data to be collected, and through an open source data collection agent, it is responsible for pulling data from the computing power node and pushing data to the time series database. By querying the data of the time series database, the open source visualization unit generates charts, dashboards and other visualization results. The visualization process is as follows:
[0071] Step 1: Obtain computing power resource data;
[0072] The collection agent actively collects data from the computing power nodes, including through system commands, APIs, SNMP, etc.
[0073] Step 2, the collection agent forwards the data to the time series database;
[0074] The collection agent pushes the collected time series data to the time series database through HTTP and other protocols. The time series database stores data in "time series".
[0075] Step 3, the visualization unit initiates a data query;
[0076] The visualization unit initiates a query request to the time series database. Before that, the time series database is configured as a data source, and the query statement is defined, such as "query the memory usage of a certain computing power node in the past 24 hours".
[0077] Step 4, the time series database returns the query result;
[0078] The time series database executes the query and returns the time series data that meets the conditions to the visualization unit. After receiving the data, the visualization unit visualizes and displays it through dashboards, charts, etc.
[0079] The alarm unit is configured to generate alarm information based on a preset alarm rule according to the prediction result of the resource prediction unit and the monitoring result of the resource monitoring unit.
[0080] According to the preset alarm rule, it is judged whether the alarm condition is met. For example, the alarm is triggered when the predicted CPU usage rate exceeds the preset threshold for a certain time in the future, or when the actual monitored memory usage exceeds the preset threshold of the capacity. When the alarm condition is met, the alarm information is generated. The alarm information includes the alarm time, the alarm type, and the involved node terminal. At the same time, the alarm unit also sends the alarm information to the operation and maintenance personnel through various ways, such as email, SMS, system pop-up window, etc., so that the operation and maintenance personnel can know and handle the abnormal situation in time.
[0081] In a specific implementation, the alarm rule is configured in the visualization unit, the evaluation function (such as the index threshold, the trend judgment, etc.) is set based on the time series data of the time series database, the visualization unit service integration is added in the alarm unit, and the connection relationship between the two is established; the Webhook address is generated in the alarm unit, the alarm unit generates a unique Webhook address (used to receive the alarm notification of the visualization unit) for the visualization unit; in the contact point of the visualization unit, the Webhook address is configured, and the channel of “sending the alarm to the alarm unit” is defined; the notification strategy is set in the visualization unit, and the notification is sent through the contact point when the alarm is triggered; the alarm strategy is defined in the alarm unit, and the notification mode, the contact person, etc. are set after receiving the alarm of the visualization unit. According to the completed configuration, the process of the alarm mechanism is as follows:
[0082] First step; the time series database periodically evaluates the resource data;
[0083] The time series database provides real-time / historical data to the visualization unit according to the set period for alarm rule evaluation.
[0084] Second step, the data meets the alarm evaluation function;
[0085] The visualization unit checks whether the time series database data meets the evaluation function according to the alarm rule, and if the resource data meets the evaluation function for a period of time, it is determined as a valid alarm.
[0086] Third step, the visualization unit sends the alarm message;
[0087] The visualization unit triggers the alarm, generates an alarm message, and includes alarm content, trigger time, index, etc.
[0088] Fourth step, the visualization unit transmits the alarm signal to the alarm unit;
[0089] The alarm message is sent to the alarm unit through the configured contact point (Webhook).
[0090] Fifth step, the alarm unit forwards the message;
[0091] The alarm unit receives the Webhook message, processes it according to the alarm strategy, forwards the alarm through the internal channel, and notifies the pre-configured contact person through email, SMS, IM, etc. to complete the alarm closed loop.
[0092] Through the above steps, a complete alarm management mechanism is realized to ensure that system abnormalities are discovered and handled in a timely manner.
[0093] In a specific implementation, the data storage unit is connected to the resource prediction unit, the resource monitoring unit and the configuration management unit through an internal communication interface. The data storage unit is responsible for receiving and storing the resource data sent by the node terminal, and providing the data to other units for use. The resource prediction unit is connected to the cloud through a communication network, for receiving the model parameters issued by the cloud, and sending the prediction results to the visualization unit and the configuration management unit. The resource monitoring unit is connected to the node terminal through a communication network, for monitoring the resource status of the node terminal in real time, and sending the monitoring results to the visualization unit and the configuration management unit. The configuration management unit is connected to the node terminal through a communication network, for issuing control instructions for running configuration to the node terminal according to the prediction results and the monitoring results. The visualization unit is connected to the resource prediction unit and the resource monitoring unit through an internal communication interface, for receiving the prediction results and the monitoring results, and displaying the received information results. The alarm unit is connected to the resource prediction unit and the resource monitoring unit through an internal communication interface, for receiving the prediction results and the monitoring results, and generating alarm information according to the preset alarm rules.
[0094] Specifically, in a preferred embodiment of the application, the data storage unit is connected to the cloud, for backing up the resource data stored based on time sequence to the cloud.
[0095] Preferably, the cloud is specifically used for:
[0096] According to the resource data backed up by the data storage unit, the historical state information sequence of each resource of each computing power node is constructed.
[0097] Each resource of each computing power node provides data of T historical time steps, and the single resource sequence is s j ∈R T (jth resource, j = 1, 2,..., K), and the node-level input is K such sequences.
[0098] According to the historical state information sequence of each resource, the first resource vector is obtained by attention aggregation of the same resource sampling sequence, and the second resource vector is obtained by aggregation of the first resource vectors of different resources at the same interval.
[0099] First-level single resource time sequence attention aggregation:
[0100] In the embedding layer, the 1-dimensional value of each time step is mapped to a d-dimensional high-dimensional space. It is represented as:
[0101]
[0102] W emb ∈R T×d is an embedding matrix learned through training.
[0103] Let the number of heads be h, and the dimension of each head be d h = d / h, the query / key / value matrix is: where, is the parameter of the i-th head, and there are h heads in total.
[0104] Attention weights:
[0105]
[0106] Multi-head output:
[0107] O = Concat(O 1 ,…,O h )·W o ∈ R T×d ,O i = A i ·V i
[0108] The sequence in the time dimension is compressed into a fixed-length d-dimensional vector, and the formula is:
[0109]
[0110] where, 1,j is the first resource vector of the j-th resource
[0111] The second-level cross-resource association attention aggregation constructs a cross-resource matrix by concatenating the first resource vectors of the K resources into a matrix, and the formula is:
[0112]
[0113] Let the number of heads be h', and the dimension of each head be d h′ = d / h'):
[0114] The query / key / value matrix is: where, is the parameter of the i-th head, and there are h' heads in total.
[0115] Attention weights:
[0116]
[0117] Multi-head output (concatenated linear transformation, keeping the dimension Kxd):
[0118] O' = Concat(O 1 ,…,O h′ )·W o '∈ R K×d ,O i = A i ·Vi
[0119] The matrix of the resource dimension is compressed into a fixed-length d-dimensional vector (averaged for resource types), and the formula is:
[0120]
[0121] Wherein, v2 is the second resource vector, integrating the timing and correlation characteristics of K resources.
[0122] The full connection layer mapping maps the second resource vector to the prediction result of K resources in the future M time steps, and the formula is:
[0123] y = v2 · W pred + b e R M×K
[0124] Wherein, W pred e R d×(M×K) is the weight matrix, and b is the bias, which is learned by training.
[0125] The second resource vector is input into the resource prediction model for training, and when the model training converges, the model parameters are updated.
[0126] In the embodiment, the cloud is used to train the model, which is different from using the server side to train, because the cloud has higher computing resources. For the computationally intensive task of training a resource prediction model based on a multi-head self-attention mechanism, the cloud can provide higher computing power than the server side. The resources of the cloud can be flexibly expanded according to actual needs. When the model training task increases or the model size becomes larger, more computing resources can be quickly applied in the cloud, such as increasing the number of CPU cores, the size of memory or the number of GPU instances. The hardware resources of the server side are fixed, and when training large-scale models or simultaneously performing multiple model training tasks, resource shortages may occur, and the server hardware upgrade process is relatively cumbersome and requires downtime and other operations.
[0127] The cloud is used to update the model parameters. When the model is trained in the cloud and obtains updated model parameters, the updated parameters can be distributed to the server side and each node terminal. Ensure that the resource prediction model in the entire power system is based on the latest training results, and improve the resource prediction accuracy.
[0128] In addition, each functional unit in each embodiment of the application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0129] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0130] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and they should be covered in the scope of the specification of the present application.
Claims
1. A resource monitoring and prediction system for a computing power network in a power system, characterized in that, include: Multiple node terminals are deployed on various computing nodes of the power system to collect resource data of the computing nodes, receive control commands for operation configuration, and execute collection actions and tasks according to the control commands. On the server side, deployed in the power system's management and control platform, it is used to issue control commands for operation configuration, receive resource data, monitor the resource status of each node terminal, construct a historical resource status information sequence based on the resource data, load the issued model parameters, run a resource prediction model based on a multi-head self-attention mechanism based on the historical resource status information sequence, output the predicted resource status, and monitor abnormal situations of the predicted resource status. In the cloud, it is used to centrally train resource prediction models based on multi-head self-attention mechanisms, output updated model parameters, and distribute them.
2. The resource monitoring and prediction system for computing power network in a power system of claim 1, wherein, The node terminal includes: The data acquisition unit is used to collect resource data of the computing node, including CPU, memory, GPU, network I / O and disk I / O; The monitoring and management unit is used to receive control instructions for the operation configuration issued by the server, and update the acquisition actions and acquisition tasks of the data acquisition unit according to the control instructions.
3. The resource monitoring and prediction system for computing power network in a power system of claim 2, wherein, The monitoring and management unit includes: The data acquisition action management module is used to manage the start, stop, and restart of the data acquisition unit according to the control commands issued by the server. The data acquisition task management module is used to receive control commands issued by the server and update the acquisition indicators and acquisition frequency of the data acquisition unit.
4. The resource monitoring and prediction system for computing power network in a power system of claim 3, wherein, The server includes: A data storage unit is used to receive resource data and store the resource data based on time sequence; The resource prediction unit is used to construct a historical resource status information sequence based on resource data, load the model parameters sent from the cloud, run a resource prediction model based on a multi-head self-attention mechanism based on the historical resource status information sequence, and output the predicted resource status. The resource monitoring unit is used to monitor the resource status of each node terminal based on resource data, and to monitor and predict abnormal resource status conditions. The configuration management unit is used to issue control commands for running configuration to the node terminal based on the prediction results of the resource prediction unit and the monitoring results of the resource monitoring unit.
5. The resource monitoring and forecasting system for computing power network in a power system of claim 4, wherein, The server-side also includes: The visualization unit is used to display the prediction results and monitoring results of the resource prediction unit and the resource monitoring unit.
6. The resource monitoring and forecasting system for computing power network in a power system of claim 4, wherein, The server-side also includes: The alarm unit is used to generate alarm information based on the prediction results of the resource prediction unit and the monitoring results of the resource monitoring unit, according to preset alarm rules.
7. The resource monitoring and forecasting system for computing power network in a power system of claim 4, wherein, The data storage unit is connected to the cloud and is used to back up resource data stored based on time sequence to the cloud.
8. The resource monitoring and forecasting system for computing power network in a power system of claim 7, wherein, The cloud is specifically used for: Based on the resource data backed up by the data storage unit, construct a sequence of historical status information for each resource of each computing node; Based on the historical state information sequence of each resource, attention aggregation is performed on the sampling sequence of the same resource to obtain the first resource vector, and the first resource vectors of different resources at the same interval are aggregated to obtain the second resource vector. The second resource vector is input into a resource prediction model for training, and when the model training converges, the model parameters are updated.