Method and device for monitoring high availability of computing power resources of intelligent computing center
By adopting a multi-server architecture in the intelligent computing center, we ensure that monitoring data can still be provided when the storage module fails, solving the problem of high availability of computing power resource monitoring and realizing continuous and reliable computing power resource monitoring.
Patent Information
- Application Number
- CN202510472932.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-29
AI Technical Summary
In the intelligent computing center, when the monitoring data module storing computing resources fails, the relevant personnel cannot know the computing resources status, resulting in the computing resources monitoring service not having high availability.
The architecture of multiple first servers and storage servers is adopted. Each first server receives monitoring data from multiple second servers and sends it to the corresponding storage server; upon receiving a request from the query platform, the target first server or storage server sends response information to the query platform to ensure that monitoring data continues to be provided through the remaining server when a certain server fails.
The high availability of computing resource monitoring of intelligent computing centers is realized, ensuring continuous and reliable operation within a certain time range, and being able to timely understand the computing resource status.
Smart Images

Figure CN120386685A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to a method and device for highly available monitoring of computing power resources in an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training and model inference, etc.). The intelligent computing center covers facilities, hardware, software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] The "intelligent computing center" includes but is not limited to the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", which is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of the target result by processing information data, and a new type of productive force integrating information computing power, network carrying capacity and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0007] Currently, during the operation of an intelligent computing center, it is necessary to monitor computing power resources. However, in the case of a failure of the storage module that stores the monitoring data of the computing power resources, relevant personnel will not be able to know the status of the computing power resources, which results in the service of monitoring the computing power resources in the intelligent computing center not being highly available. Therefore, since the emergence of intelligent computing centers, how to achieve the high availability of computing power resource monitoring has been an urgent problem to be solved. Summary of the Invention
[0008] The present invention provides a method and device for highly available monitoring of computing power resources in an intelligent computing center, which is used to achieve the high availability of monitoring the computing power resources in the intelligent computing center.
[0009] In order to solve the above technical problems, the present invention is implemented as follows:
[0010] In a first aspect, the present invention provides a method for high-availability monitoring of computing resources in an intelligent computing center, wherein the intelligent computing center includes a query platform, multiple first servers, and multiple second servers, wherein the query platform is connected to the first servers, each of the first servers corresponds to a storage server, each of the first servers is connected to the multiple second servers, and the second servers are used to collect monitoring data of computing nodes in the intelligent computing center, wherein the method includes:
[0011] Step S1: Each of the first servers receives monitoring data sent by the plurality of second servers;
[0012] Step S2: Each of the first servers sends the monitoring data to the corresponding storage server;
[0013] Step S3: Upon receiving the query request sent by the query platform, the target first server or the target storage server sends a response message to the query platform, where the target first server is any normally operating server among the multiple first servers, and the target storage server is any normally operating server among the storage servers corresponding to the multiple first servers, and the response message includes the data obtained by the target first server or the target storage server based on the query request.
[0014] Optionally, the intelligent computing center further includes a node registration center. Before step S1, the method further includes:
[0015] Step S4: The second server determines the target node based on the node registration information sent by the node registration center;
[0016] Step S5: The second server collects the monitoring data monitored by the monitoring component of the target node;
[0017] Step S6: The plurality of second servers send the monitoring data collected by each of them to each of the first servers.
[0018] Optionally, step S5 includes:
[0019] Step S51: The second server receives a configuration file, where the configuration file is used to indicate a functional module and a data format corresponding to the second server;
[0020] Step S52: The second server collects monitoring data of the target node based on the configuration file, where the monitoring data is used to indicate the computing resource usage of the functional module indicated by the configuration file in the target node, and the format of the monitoring data is the same as the data format indicated by the configuration file.
[0021] Optionally, step S2 includes:
[0022] Step S21: Each of the first servers integrates the monitoring data sent by each of the multiple second servers according to the receiving time sequence to obtain the integrated monitoring data;
[0023] Step S22: Each of the first servers sends the integrated monitoring data to the corresponding storage server.
[0024] Optionally, the intelligent computing center further includes an alarm component. After step S4, the method further includes:
[0025] Step S7: The alarm component senses the monitoring data monitored by the monitoring component of the target node;
[0026] Step S8: When the monitoring data monitored by the monitoring component of the target node indicates abnormal use of computing resources, the alarm component issues an alarm message.
[0027] Optionally, after step S3, the method further includes:
[0028] Step S9: The query platform visually displays the monitoring data.
[0029] In a second aspect, the present invention provides a high-availability monitoring device for computing resources of an intelligent computing center. The intelligent computing center includes a query platform, multiple first servers, and multiple second servers. The query platform is connected to the first servers. Each of the first servers corresponds to a storage server. Each of the first servers is connected to the multiple second servers. The second servers are used to collect monitoring data of the computing nodes of the intelligent computing center. The device includes:
[0030] A first receiving module, configured to control each of the first servers to receive the monitoring data sent by the multiple second servers;
[0031] A first sending module, configured to control each of the first servers to send the monitoring data to the corresponding storage server;
[0032] A second sending module, configured to, when receiving a query request sent by the query platform, control the target first server or the target storage server to send a response message to the query platform. The target first server is any normally operating server among the multiple first servers, and the target storage server is any normally operating server among the storage servers corresponding to the multiple first servers. The response message includes the data queried by the target first server or the target storage server based on the query request.
[0033] In a third aspect, the present invention provides an electronic device comprising: a processor, a memory, and a program stored in the memory and runnable on the processor, wherein when the program is executed by the processor, the steps of the method for high-availability monitoring of computing resources of an intelligent computing center as described in the first aspect above are implemented.
[0034] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for high-availability monitoring of computing resources of an intelligent computing center as described in the first aspect above.
[0035] In a fifth aspect, the present invention provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps of the method for high-availability monitoring of computing resources of an intelligent computing center as described in the first aspect above.
[0036] In the present invention, each of the first servers receives the monitoring data sent by the multiple second servers; each of the first servers sends the monitoring data to the corresponding storage server; when receiving the query request sent by the query platform, the target first server or the target storage server sends a response message to the query platform, the target first server is any normally operating server among the multiple first servers, and the target storage server is any normally operating server among the storage servers corresponding to the multiple first servers. Since there are multiple first servers and storage servers with the same stored data, in the event that one of the servers fails, the complete monitoring data can still be sent to the query platform through the remaining servers, so that relevant personnel can know the status of computing resources and realize computing resource monitoring. Through the above method, the computing resource monitoring service of the intelligent computing center can have the ability to operate continuously and reliably within a certain time range, that is, to achieve high availability of computing resource monitoring of the intelligent computing center. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0038] Figure 1 A schematic diagram of the intelligent computing center architecture provided by an embodiment of the present invention;
[0039] Figure 2Schematic flowchart of a method for monitoring high availability of computing power resources in an intelligent computing center provided by an embodiment of the present invention;
[0040] Figure 3 Schematic structural diagram of a device for monitoring high availability of computing power resources in an intelligent computing center provided by an embodiment of the present invention;
[0041] Figure 4 Schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0042] The technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0043] First, the technical terms related to the present invention will be briefly described below.
[0044] The "computing power" referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to process information data and achieve the output of a target result, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.
[0045] The "computational power" (CP) referred to in the present invention means: the ability of a data center server to process data and achieve the output of results, a comprehensive index for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP_general + CP_intelligent + CP_super.
[0046] The "network power" (NP) referred to in the present invention means: the manifestation of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission inside and between data centers, and a comprehensive index for measuring network transmission scheduling ability.
[0047] The "Storage Power" (SP) described in the present invention refers to the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center and includes external storage devices such as storage arrays and built-in storage devices of servers. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0048] The "computing power infrastructure" described in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage power, and can realize the centralized computing, storage, transmission, and application of information.
[0049] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0050] The "computing power" described in the present invention includes general computing power, intelligent computing power, and super computing power.
[0051] The "general computing power" described in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0052] The "intelligent computing power" described in the present invention refers to a computing platform that is scaled for various artificial intelligence innovation applications based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing and machine vision.
[0053] The "super computing power" described in the present invention mainly refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.
[0054] The "Intelligent Computing Center" described in the present invention refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0055] The "Intelligent Computing Center" described in the present invention includes, but is not limited to, the "Intelligent Computing Center".
[0056] The "Intelligent Computing Center" described in the present invention, namely the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0057] The "Computing Power Center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, and having computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0058] The "Supercomputing Center" described in the present invention refers to, that is, a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0059] The "Computing Power Resources" described in the present invention refers to technologies and facilities required for the development of the digital society, having information computing, transmission, storage, and application capabilities, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0060] The "High-Availability Monitoring of Computing Power Resources" described in the present invention refers to: High availability means the ability of a system, service, or component to operate continuously and reliably within a certain time range without interruption or failure; the "High-Availability Monitoring of Computing Power Resources" refers to the service of monitoring the computing power resources of an intelligent computing center, which has the ability to operate continuously and reliably within a certain time range.
[0061] An embodiment of the present invention provides a method for high-availability monitoring of computing power resources of an intelligent computing center, wherein, as Figure 1 shown, the intelligent computing center includes a query platform (such as Figure 1 Grafana shown) and multiple first servers (such as Figure 1The Prometheus01, Prometheus02, Prometheus03 shown) and multiple second servers (such as Figure 1 Prometheus-A, Prometheus-B, Prometheus-C shown), the query platform is connected to the first server, each of the first servers corresponds to a storage server, each of the first servers is connected to the multiple second servers, and the second servers are used to collect monitoring data of computing nodes of the intelligent computing center.
[0062] Figure 2 This is a high-availability monitoring method for computing power resources of an intelligent computing center provided by an embodiment of the present invention, including the following steps:
[0063] Step S1: Each of the first servers receives the monitoring data sent by the multiple second servers.
[0064] In this step, the second server can also be called the lower-layer server, which is responsible for locally collecting monitoring data. The first server can be called the upper-layer server, and it collects monitoring data from multiple lower-layer servers through a federation mechanism. The upper-layer server can act as an aggregation layer to summarize and perform high-level monitoring and analysis on the data.
[0065] It should be noted that the content received by each first server is the same.
[0066] Step S2: Each of the first servers sends the monitoring data to the corresponding storage server.
[0067] Among them, each first server cooperates with a data transmission component (for example Figure 1 Thanos Sidecar shown) to be responsible for uploading the data of the first server to the storage server. The storage server is a persistent storage server corresponding to the data transmission component, and in this embodiment, it is the Thanos storage layer corresponding to Thanos Sidecar, for example Figure 1 MINIO shown.
[0068] It should be noted that the data that the first server can store by itself is relatively limited. The first server sends data to the storage server at preset intervals to save more monitoring data, which is convenient for subsequent querying of historical monitoring data.
[0069] Step S3: When receiving the query request sent by the query platform, the target first server or the target storage server sends response information to the query platform. The target first server is any normally operating server among the multiple first servers, and the target storage server is any normally operating server among the storage servers corresponding to the multiple first servers. The response information includes the data queried by the target first server or the target storage server based on the query request.
[0070] In this step, relevant personnel send a query request through the query platform to obtain the target monitoring data they want to query. According to the requirements corresponding to the query request, the target monitoring data can be the monitoring data within a week or the monitoring data half a year ago. Depending on the time period corresponding to the target monitoring data to be queried, the target monitoring data can be sent by the first server or the storage server. It can be understood that when the time period corresponding to the target monitoring data is within the time period that the first server can store, the first server sends the target monitoring data; when the time period corresponding to the target monitoring data exceeds the time period that the first server can store, the storage server sends the target monitoring data.
[0071] The query platform queries the data in the first server and the storage server through a query component (such as Figure 1 the Thanos query shown). The Thanos Query backend configures a VIP (100.64.1.100) through HAProxy and Keepalived to ensure load balancing and failover among multiple Thanos Query instances. Through the above configuration, in this embodiment, when receiving a query request, a normally operating first server is randomly selected from the multiple first servers as the target first server, or a normally operating storage server is randomly selected from the multiple storage servers as the target storage server, and the monitoring data is sent to the query platform. In the case where a certain server among the multiple first servers or multiple storage servers has a running failure, it is still possible to select a target server from the remaining first servers or storage servers to send the monitoring data to the query platform, thereby realizing the monitoring of computing power resources.
[0072] In the method provided in this embodiment, each of the first servers receives the monitoring data sent by the multiple second servers; each of the first servers sends the monitoring data to the corresponding storage server; in the case of receiving a query request sent by the query platform, the target first server or the target storage server sends response information to the query platform, where the target first server is any normally operating server among the multiple first servers, and the target storage server is any normally operating server among the storage servers corresponding to the multiple first servers. Since there are multiple first servers and storage servers with the same stored data, in the case where a certain server fails, the complete monitoring data can still be sent to the query platform through the remaining servers, enabling relevant personnel to know the computing power resource status and realizing the monitoring of computing power resources. Through the above method, the service of monitoring the computing power resources of the intelligent computing center can be made to have the ability to continuously and reliably operate within a certain time range, that is, to achieve high availability of the monitoring of the computing power resources of the intelligent computing center.
[0073] Optionally, the intelligent computing center further includes a node registration center. Before step S1, the method further includes:
[0074] Step S4: The second server determines the target node based on the node registration information sent by the node registration center;
[0075] Step S5: The second server collects the monitoring data monitored by the monitoring components of the target node;
[0076] Step S6: The multiple second servers send the monitoring data collected by each of them to each of the first servers.
[0077] In this embodiment, the intelligent computing center further includes a Figure 1 node registration center consul as shown. All computing nodes of the intelligent computing center will be registered in the node registration center. For example, when adding or deleting nodes, the node information can be updated in a timely manner.
[0078] A node registration center is deployed in each second server, which is used to send node registration information to the second server, enabling the second server to identify the target node. In subsequent steps, the second server does not need to collect the monitoring data of each associated node, but only needs to collect the monitoring data monitored by the monitoring components of the target node, thereby improving the execution efficiency of the second server.
[0079] Optionally, step S5 includes:
[0080] Step S51: The second server receives a configuration file, which is used to indicate the function module and data format corresponding to the second server;
[0081] Step S52: The second server collects the monitoring data of the target node based on the configuration file. The monitoring data is used to indicate the computing power resource usage of the functional modules indicated in the configuration file in the target node, and the format of the monitoring data is the same as the data format indicated in the configuration file.
[0082] In this embodiment, the monitoring data corresponding to the target node is sent to multiple second servers. Each server among the multiple second servers corresponds to different functional modules according to the configuration file, and the data format corresponding to each configuration file is the same. Through the configuration file, each second server only collects the monitoring data of its corresponding functional module, which realizes the classification of the monitoring data and facilitates subsequent data analysis. In addition, when the monitoring data comes from the target node, due to the differences in the actual physical devices of the target node, the data format may be diverse. The second server can unify the data format of the monitoring data to facilitate subsequent data analysis, thereby improving the analysis efficiency.
[0083] Optionally, step S2 includes:
[0084] Step S21: Each first server integrates the monitoring data sent by each of the multiple second servers according to the receiving time sequence to obtain the integrated monitoring data;
[0085] Step S22: Each first server sends the integrated monitoring data to the corresponding storage server.
[0086] In this step, after the first server summarizes and integrates the monitoring data corresponding to different functional modules sent by multiple second servers, it is sent to the storage server, so that the storage server can directly send the integrated monitoring data to the query platform later, reducing the integration process and improving the execution efficiency of the method.
[0087] Optionally, the intelligent computing center further includes an alarm component. After step S4, the method further includes:
[0088] Step S7: The alarm component senses the monitoring data monitored by the monitoring component of the target node;
[0089] Step S8: When the monitoring data monitored by the monitoring component of the target node indicates abnormal computing power resource usage, the alarm component issues an alarm message.
[0090] In this embodiment, the intelligent computing center further includes Figure 1The alarm component nightingale shown is used to sense the monitoring data of the target node and can issue an alarm for abnormal monitoring data before the query platform receives the monitoring data, so that the operation and maintenance personnel can handle it in time to reduce losses.
[0091] Optionally, after step S3, the method further includes:
[0092] Step S9: The query platform displays the monitoring data in a visual manner.
[0093] In this embodiment, the query platform can visualize the monitoring data through various methods such as charts, tables, and geographic maps, making the monitoring data more intuitive, facilitating data analysis, and improving the efficiency of data analysis.
[0094] like Figure 3 As shown, an embodiment of the present invention further provides a high-availability monitoring device 300 for computing resources of an intelligent computing center, wherein the intelligent computing center includes a query platform, multiple first servers, and multiple second servers. The query platform is connected to the first servers, each of the first servers corresponds to a storage server, each of the first servers is connected to the multiple second servers, and the second servers are used to collect monitoring data of computing nodes of the intelligent computing center. The device 300 includes:
[0095] A first receiving module 301, configured to control each of the first servers to receive the monitoring data sent by the plurality of second servers;
[0096] A first sending module 302, configured to control each of the first servers to send the monitoring data to the corresponding storage server;
[0097] The second sending module 303 is used to control the target first server or the target storage server to send response information to the query platform when receiving the query request sent by the query platform. The target first server is any normally operating server among the multiple first servers, and the target storage server is any normally operating server among the storage servers corresponding to the multiple first servers. The response information includes the data obtained by the target first server or the target storage server based on the query request.
[0098] Optionally, the intelligent computing center further includes a node registration center, and the apparatus 300 further includes:
[0099] A first determining module, configured to control the second server to determine a target node based on the node registration information sent by the node registration center;
[0100] The first collection module is used to control the second server to collect the monitoring data monitored by the monitoring components of the target node;
[0101] The third sending module is used to control multiple second servers to send the collected monitoring data to each first server.
[0102] Optionally, the first collection module is further used to:
[0103] Control the second server to receive a configuration file, where the configuration file is used to indicate the function modules and data formats corresponding to the second server;
[0104] Control the second server to collect the monitoring data of the target node based on the configuration file, where the monitoring data is used to indicate the computing power resource usage of the function modules indicated by the configuration file in the target node, and the format of the monitoring data is the same as the data format indicated by the configuration file.
[0105] Optionally, the first sending module 302 is further used to:
[0106] Control each first server to integrate the monitoring data sent by each of the multiple second servers according to the receiving time sequence to obtain the integrated monitoring data;
[0107] Control each first server to send the integrated monitoring data to the corresponding storage server.
[0108] Optionally, the intelligent computing center further includes an alarm component, and the device 300 further includes:
[0109] A sensing module, which is used to control the alarm component to sense the monitoring data monitored by the monitoring components of the target node;
[0110] An alarm module, which is used to control the alarm component to send an alarm message when the monitoring data monitored by the monitoring components of the target node indicates abnormal computing power resource usage.
[0111] Optionally, the device 300 further includes:
[0112] A display module, which is used to control the query platform to visually display the monitoring data.
[0113] The computing power resource highly available monitoring device 300 of the intelligent computing center provided by the present invention can implement each process of the above-mentioned embodiments of the computing power resource highly available monitoring method of the intelligent computing center, and the technical features correspond one by one and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0114] Please refer to Figure 4, the present invention also provides an electronic device, including a processor 401, a memory 402, and a computer program stored on the memory 402 and executable on the processor 401. When the computer program is executed by the processor 401, it implements each process of the foregoing embodiments of the method for monitoring high availability of computing power resources in the intelligent computing center, and can achieve the same technical effects. To avoid repetition, details are not described herein again.
[0115] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the foregoing embodiments of the method for monitoring high availability of computing power resources in the intelligent computing center, and can achieve the same technical effects. To avoid repetition, details are not described herein again. Among them, the computer-readable storage medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0116] An embodiment of the present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement each process of the foregoing Figure 2 embodiments of the method for monitoring high availability of computing power resources in the intelligent computing center shown, and can achieve the same technical effects. To avoid repetition, details are not described herein again.
[0117] It should be noted that in this document, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0119] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.
Claims
1. A method for monitoring the high availability of computing power resources in an intelligent computing center, characterized in that, The intelligent computing center includes a query platform, a plurality of first servers, and a plurality of second servers. The query platform is connected to the first servers. Each of the first servers corresponds to a storage server, and each of the first servers is connected to the plurality of second servers. The second servers are used to collect monitoring data of the computing nodes of the intelligent computing center. The method includes: Step S1: Each of the first servers receives the monitoring data sent by the plurality of second servers; Step S2: Each of the first servers sends the monitoring data to the corresponding storage server; Step S3: When receiving a query request sent by the query platform, the target first server or the target storage server sends response information to the query platform. The target first server is any normally operating server among the plurality of first servers, and the target storage server is any normally operating server among the storage servers corresponding to the plurality of first servers. The response information includes the data obtained by the target first server or the target storage server based on the query request.
2. The method according to claim 1, wherein The intelligent computing center further includes a node registration center. Before step S1, the method further includes: Step S4: The second server determines a target node based on the node registration information sent by the node registration center; Step S5: The second server collects the monitoring data monitored by the monitoring components of the target node; Step S6: The plurality of second servers send the monitoring data collected by each of them to each of the first servers.
3. The method according to claim 2, wherein Step S5 includes: Step S51: The second server receives a configuration file, which is used to indicate the function modules and data formats corresponding to the second server; Step S52: The second server collects the monitoring data of the target node based on the configuration file. The monitoring data is used to indicate the computing power resource usage of the function modules indicated by the configuration file in the target node, and the format of the monitoring data is the same as the data format indicated by the configuration file.
4. The method according to claim 2, wherein Step S2 includes: Step S21: Each of the first servers integrates the monitoring data sent by each of the plurality of second servers according to the receiving time sequence to obtain the integrated monitoring data; Step S22: Each of the first servers sends the integrated monitoring data to the corresponding storage server.
5. The method according to claim 2, wherein The intelligent computing center further includes an alarm component. After step S4, the method further includes: Step S7: The alarm component senses the monitoring data monitored by the monitoring components of the target node; Step S8: When the monitoring data monitored by the monitoring components of the target node indicates abnormal computing power resource usage, the alarm component issues an alarm message.
6. The method according to any one of claims 1 to 5, characterized in that After step S3, the method further includes: Step S9: The query platform visually displays the monitoring data.
7. An apparatus for monitoring the high availability of computing power resources in an intelligent computing center, characterized in that, The intelligent computing center includes a query platform, a plurality of first servers, and a plurality of second servers. The query platform is connected to the first servers. Each of the first servers corresponds to a storage server, and each of the first servers is connected to the plurality of second servers. The second servers are used to collect monitoring data of the computing nodes of the intelligent computing center. The device includes: A first receiving module, configured to control each of the first servers to receive the monitoring data sent by the plurality of second servers; A first sending module, configured to control each of the first servers to send the monitoring data to the corresponding storage server; A second sending module, configured to, when receiving a query request sent by the query platform, control the target first server or the target storage server to send response information to the query platform. The target first server is any normally operating server among the plurality of first servers, and the target storage server is any normally operating server among the storage servers corresponding to the plurality of first servers. The response information includes the data obtained by the target first server or the target storage server based on the query request.
8. An electronic device, characterized in that, Comprising: A processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, the steps of the method for highly available monitoring of computing power resources of the intelligent computing center according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the method for highly available monitoring of computing power resources of the intelligent computing center according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that, Including computer instructions. When the computer instructions are executed by a processor, the steps of the method for highly available monitoring of computing power resources of the intelligent computing center according to any one of claims 1 to 6 are implemented.