Method and device for monitoring computing power operation task of intelligent computing center

By creating a data link on the Index service node and returning the running status of the DiskANN object in real time, the problem of poor timeliness of client acquisition status information in the existing technology is solved, and real-time monitoring of the computing power resource operation tasks of the intelligent computing center is realized.

CN120216290APending Publication Date: 2025-06-27DATACANVAS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510302321.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the existing DiskANN object running status monitoring scheme for running tasks, the client needs to send status query instructions, resulting in poor timeliness of status information.

Method used

Receive the client's monitoring instructions on the Index service node, create a data link between the DiskANN service node and the client, and return the running status of the DiskANN object in real time.

Benefits of technology

Real-time monitoring of the running status of DiskANN objects that are calling the computing power resources of the Intelligent Computing Center is realized, and the timeliness of the client to obtain the running status is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216290A_ABST
    Figure CN120216290A_ABST
Patent Text Reader

Abstract

The invention provides a monitoring method and device for computing power operation tasks of an intelligent computing center, the monitoring method is applied to an Index service node, and the monitoring method comprises the following steps: S1, receiving a monitoring instruction sent by a client, the monitoring instruction being used for indicating the operation state of a DiskANN object monitoring the operation tasks of the computing power resources of the intelligent computing center at present; and S2, in response to the monitoring instruction, creating a data link between a DiskANN service node and the client, and returning the running state of the DiskANN object obtained from the DiskANN service node to the client in real time through the data link. According to the method and the system, the running state of the DiskANN object which is calling the running task of the computing power resource of the intelligent computing center can be monitored in real time, and the timeliness of obtaining the running state of the DiskANN object by the client is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure technologies, and in particular to a monitoring method and device for the operation tasks of the computing power of an intelligent computing center. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training and model inference, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from the underlying computing power to the top-level application enabling.

[0004] The "intelligent computing center" includes but is not limited to the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", which is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of the target result by processing information data, and a new type of productive force integrating information computing power, network carrying capacity and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0007] In the existing monitoring scheme for the running status of the DiskANN object of the running task, the client sends a status query instruction (Status RPC) to the Index service node, the Index service node forwards the status query instruction (Status RPC) to the DiskANN service node, the DiskANN service node obtains the status information of the DiskANN object specified by the status query instruction (Status RPC), and then sends the status information to the client through the Index service node. That is to say, each time the client needs to send a status query instruction (Status RPC), and the obtained status information actually only represents the status of the DiskANN object at the moment when the DiskANN service node obtains the status information, and the timeliness of the status information of the DiskANN object obtained by the client is poor. Summary of the Invention

[0008] The present invention provides a method and device for monitoring the computing power operation tasks of an intelligent computing center, so as to solve the problem of poor timeliness of the status information of the DiskANN object obtained by the client in the existing monitoring solution for the running status of the DiskANN object of the running task. In the existing solution, the client sends a status query instruction (StatusRPC) to the Index service node, the Index service node forwards the status query instruction (Status RPC) to the DiskANN service node, the DiskANN service node obtains the status information of the DiskANN object specified by the status query instruction (Status RPC), and then sends the status information to the client through the Index service node.

[0009] To solve the above technical problems, the present invention is implemented as follows:

[0010] In a first aspect, the present invention provides a method for monitoring the computing power operation tasks of an intelligent computing center, which is applied to an Index service node and includes:

[0011] Step S1: Receive a monitoring instruction sent by the client, where the monitoring instruction is used to indicate monitoring the running status of the DiskANN object of the running task that is currently invoking the computing power resources of the intelligent computing center;

[0012] Step S2: In response to the monitoring instruction, create a data link between the DiskANN service node and the client, and through the data link, return the running status of the DiskANN object obtained from the DiskANN service node to the client in real time.

[0013] Optionally, the task includes: a task of importing data;

[0014] The running status includes the amount of data imported.

[0015] Optionally, the task includes: a task of creating a DiskANN object, a task of loading a DiskANN object;

[0016] Returning the running status to the client in real time includes:

[0017] Returning the running status to the client in real time by using an asynchronous return method.

[0018] Optionally, the task has multiple task types;

[0019] The step S2 includes:

[0020] Step S21: Divide the data link into multiple queues corresponding to the task types;

[0021] Step S22: According to the task type, return the running status to the client through the corresponding queue.

[0022] Optionally, the monitoring instruction is further used to indicate monitoring whether an error occurs when the DiskANN object runs a task;

[0023] After the step S1, it includes:

[0024] Step S3: In response to the monitoring instruction, monitor whether an error occurs when the DiskANN object runs a task;

[0025] Step S4: In the case of monitoring that an error occurs to the DiskANN object, record the running error information of the DiskANN object, and return the running error information to the client through the data link.

[0026] Optionally, the running status further includes at least one of the following:

[0027] The maximum number of asynchronous input / output operations AIO, memory usage, disk usage, and central processing unit CPU load.

[0028] In a second aspect, the present invention provides a monitoring device for the computing power operation task of an intelligent computing center, which is applied to an Index service node and includes:

[0029] A receiving module, configured to receive a monitoring instruction sent by a client, where the monitoring instruction is used to indicate monitoring the running status of a DiskANN object of a running task that is currently invoking the computing power resources of the intelligent computing center;

[0030] An execution module, configured to, in response to the monitoring instruction, create a data link between the DiskANN service node and the client, and return the running status of the DiskANN object obtained from the DiskANN service node to the client in real time through the data link.

[0031] In a third aspect, the present invention provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements the steps in the monitoring method for the computing power operation task of the intelligent computing center as described in any item of the first aspect.

[0032] In a fourth aspect, the present invention provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements the steps in the monitoring method for the computing power operation task of the intelligent computing center as described in any item of the first aspect.

[0033] Fifth aspect, the present invention provides a computer program product, including computer instructions, which when executed by a processor, implement the steps of the monitoring method for the computing power operation task of the intelligent computing center as described in any one of the first aspect.

[0034] In the present invention, through step S1: receiving a monitoring instruction sent by a client, the monitoring instruction is used to indicate monitoring the running status of a DiskANN object of a running task that is currently invoking the computing power resources of the intelligent computing center; step S2: in response to the monitoring instruction, creating a data link between the DiskANN service node and the client, and through the data link, returning the running status of the DiskANN object obtained from the DiskANN service node to the client in real time. The present invention can monitor the running status of the DiskANN object of the running task that is currently invoking the computing power resources of the intelligent computing center in real time, and the timeliness for the client to obtain the running status of the DiskANN object is high. Description of the Drawings

[0035] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0036] Figure 1 is a schematic diagram of the overall process of data interaction based on DiskANN;

[0037] Figure 2 is a schematic flowchart of the monitoring method for the computing power operation task of the intelligent computing center of the present invention;

[0038] Figure 3 is a principle block diagram of the monitoring device for the computing power operation task of the intelligent computing center;

[0039] Figure 4 is a principle block diagram of the electronic device of the present invention. Detailed Embodiments

[0040] The technical solutions in the present invention will be clearly and completely described below with reference to the drawings in the present invention. Obviously, the described embodiments are some of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0041] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same category, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "or" in the present invention means at least one of the connected objects. For example, "A or B" covers three scenarios, namely, Scenario 1: including A and not including B; Scenario 2: including B and not including A; Scenario 3: including both A and B. The character " / " generally indicates that the related objects before and after are in an "or" relationship.

[0042] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0043] First, a brief description of the technical terms related to the present invention will be given below.

[0044] The "computing power" referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of a target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.

[0045] The "computational power" (Computational Power, CP) referred to in the present invention means: the ability of a data center server to process data and achieve result output, a comprehensive indicator for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP_general + CP_intelligent + CP_super.

[0046] The "carrying capacity" (Network Power, NP) referred to in the present invention means: the performance of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission inside and between data centers, and a comprehensive indicator for measuring network transmission scheduling ability.

[0047] The "Storage Power" (SP) described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, which is a comprehensive indicator for measuring the data storage capacity of a data center and includes external storage devices such as storage arrays and built-in storage devices of servers. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0048] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage power, which can realize the centralized computing, storage, transmission, and application of information, presenting characteristics such as multi-source ubiquitous, intelligent and agile, secure and reliable, and green and low-carbon, and is of great significance for boosting industrial transformation and upgrading, empowering China's scientific and technological innovation, meeting people's beautiful life, and realizing high-efficiency social governance.

[0049] The "new type of information infrastructure" described in the present invention mainly refers to: network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing. With the emergence and popularization of new general technologies, the form of the new type of information infrastructure will be more diverse.

[0050] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.

[0051] The "general computing power" described in the present invention refers to: the computing ability provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0052] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform based on the large-scale deployment of dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, and so on.

[0053] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0054] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0055] The "intelligent computing center" described in the present invention includes but is not limited to intelligent computing centers.

[0056] The "intelligent computing center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.

[0057] The "computing power center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, transportation power, and storage power, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0058] The "supercomputing center" described in the present invention refers to: that is, the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0059] The "computing power resources" described in the present invention refers to: technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0060] The "DiskANN" described in the present invention is a vector retrieval engine based on distributed storage, capable of storing and retrieving vector data at the billion level on a single computer. Compared with traditional vector retrieval algorithms, DiskANN has higher storage efficiency and faster retrieval speed.

[0061] In the present invention, the overall process of data interaction based on DiskANN can be referred to Figure 1 , and this overall process mainly includes the following tasks: creating a DiskANN object (abbreviated as Create), importing vector data into the DiskANN object (abbreviated as ImportData), importing the vector data into the DiskANN service node (abbreviated as Push Data), building a graph index for the imported vector data (abbreviated as Bulid), loading the graph index into memory (abbreviated as Load), querying vector data based on the graph index (abbreviated as Search), closing the DiskANN object (abbreviated as Close), and destroying the DiskANN object (abbreviated as Destroy), etc.

[0062] Among them, creating a DiskANN object means that the DiskANN service node creates a DiskANN object according to the client's Create request. When the DiskANN object is created, there is no data in it, and it is in a data-free state and cannot provide any services to the outside. The client's Create request is sent to the DiskANN service node through the Index service node. After the DiskANN service node creates the DiskANN object, it notifies the client through the Index service node.

[0063] In the task of importing vector data into the DiskANN object, the client sends an ImportData request containing vector data to the Index service node, and the Index service node temporarily stores the vector data in the ImportData request in the database of the Index service node. After the Index service node finishes importing the data, it notifies the client.

[0064] After the data import task, the client can send a build (Bulid) request to the Index service node. According to the build (Bulid) request, the Index service node executes the Push Data task, that is, the Index service node imports the vector data stored in the database of the Index service node into the DiskANN service node, and the DiskANN service node notifies the Index service node after the vector data is written to disk to form a file.

[0065] The build (Bulid) task means that the DiskANN service node processes the vector data of the imported DiskANN object based on the build (Bulid) request sent by the Index service node, obtains the graph index of the vector data, and saves the vector data and the graph index. After the DiskANN service node completes the build (Bulid) task, it notifies the client through the Index service node.

[0066] The Load task means that the client sends a load request to the DiskANN service node through the Index service node, and the DiskANN service node loads at least part of the graph index of the vector data of the DiskANN object into memory based on the load request. After the loading is completed, the DiskANN service node notifies the client through the Index service node.

[0067] The Search task means that the client sends a vector data query request to the DiskANN service node through the Index service node, and the DiskANN service node retrieves the vector data most similar to the vector data to be queried based on the graph index, and returns the retrieved most similar vector data to the client through the Index service node.

[0068] The Close task means that the client sends a close request for the DiskANN object to the DiskANN service node through the Index service node, and the DiskANN service node closes the DiskANN object based on the close request. After the closing is completed, the DiskANN service node notifies the client through the Index service node.

[0069] The Destroy task means that the client sends a destroy request for the DiskANN object to the DiskANN service node through the Index service node, and the DiskANN service node destroys the DiskANN object based on the destroy request. After the destruction is completed, the DiskANN service node notifies the client through the Index service node.

[0070] The above process includes three execution entities: the client, the Index service node, and the DiskANN service node. Among them, the purpose of using the Index service node is to make the DiskANN service node as lightweight as possible. The DiskANN service node only processes the core business of DiskANN, and outsources other services to the Index service node to ensure the stability and reliability of the overall service. Before executing the above process, the Index service node and the DiskANN service node need to perform a handshake connection in advance to provide services for subsequent operations. The above Index service node and DiskANN service node can be deployed on different physical machines or on the same physical machine, and the present invention does not limit this.

[0071] In the existing monitoring solution for the running status of DiskANN objects of running tasks, the client sends a status query instruction (Status RPC) to the Index service node, and the Index service node forwards the status query instruction (Status RPC) to the DiskANN service node. The DiskANN service node obtains the status information of the DiskANN object specified by the status query instruction (Status RPC), and then sends the status information to the client through the Index service node. That is to say, each time the client needs to send a status query instruction (Status RPC), and the obtained status information actually only represents the status of the DiskANN object at the moment when the DiskANN service node obtains this status information, so the timeliness of the status information of the DiskANN object obtained by the client is poor.

[0072] Based on this, the present invention provides a monitoring method for the computing power operation tasks of an intelligent computing center, which is applied to the Index service node. See Figure 2 as shown in Figure 2 which is a schematic flow chart of the monitoring method for the computing power operation tasks of the intelligent computing center of the present invention, and includes:

[0073] Step S1: Receive a monitoring instruction sent by the client, where the monitoring instruction is used to indicate monitoring the running status of a DiskANN object of a running task that is currently invoking the computing power resources of the intelligent computing center;

[0074] Step S2: In response to the monitoring instruction, create a data link between the DiskANN service node and the client, and through the data link, return the running status of the DiskANN object obtained from the DiskANN service node to the client in real time.

[0075] In the present invention, there can be multiple DiskANN objects on the DiskANN service node, and the multiple DiskANN objects can be in states of running different tasks that invoke the computing power resources of the intelligent computing center.

[0076] It should be noted that the DiskANN object calls the computing power resources of the intelligent computing center to run tasks. Based on the abundant computing power resources of the intelligent computing center, the efficiency of the DiskANN object running tasks is improved. And improving the efficiency of the DiskANN object running tasks also means that the change rate of the running state of the DiskANN object is faster, resulting in worse timeliness of the status information obtained by the existing monitoring method implemented through the status query instruction (Status RPC). The monitoring implemented by the embodiments of the present invention through steps S1 and S2 can monitor the running state of the DiskANN object that is currently calling the computing power resources of the intelligent computing center in real time, and the client has high timeliness in obtaining the running state of the DiskANN object.

[0077] In some optional embodiments, the data link is a data transmission path between the DiskANN service node and the client relayed by the Index service node. Different from Figure 1 the data interaction process shown, in the data link, the Index service node completely acts as a data relay and does not participate in any data processing and processing to ensure the real-time nature of the client receiving the running state. In some optional embodiments, based on multiple existing communication channels between the Index service node and the client and between the Index service node and the DiskANN service node, the two communication channels with the fastest transmission rates (one with the fastest transmission rate between the Index service node and the client, and the other with the fastest transmission rate between the Index service node and the DiskANN service node) can be determined to create the data link to ensure the high transmission rate of the data link, reduce the latency of the running state returned to the client, and ensure the real-time nature of the client receiving the running state.

[0078] In the present invention, through step S1: receiving a monitoring instruction sent by the client, the monitoring instruction is used to indicate monitoring the running state of the DiskANN object that is currently calling the computing power resources of the intelligent computing center; step S2: in response to the monitoring instruction, creating a data link between the DiskANN service node and the client, and through the data link, the running state of the DiskANN object obtained from the DiskANN service node is returned to the client in real time. The present invention can monitor the running state of the DiskANN object that is currently calling the computing power resources of the intelligent computing center in real time, and the client has high timeliness in obtaining the running state of the DiskANN object.

[0079] In some embodiments of the present invention, optionally, the task includes: a task of importing data; the running state includes the amount of data imported.

[0080] It should be noted that the task of importing data, that is, the task of importing vector data into the DiskANN object (abbreviated as ImportData) described above. When the running task is the ImportData task, the running status includes the amount of data imported, that is to say, the running status includes the real-time amount of vector data imported into the DiskANN object.

[0081] In some embodiments of the present invention, optionally, the tasks include: the task of creating a DiskANN object, the task of loading a DiskANN object;

[0082] Return the running status to the client in real time, including:

[0083] Use an asynchronous return method to return the running status to the client in real time.

[0084] It should be noted that the task of creating a DiskANN object is the Build task described above. The task of loading a DiskANN object is the Load task described above.

[0085] Asynchronous return, that is, asynchronous data transmission. In this method, the sending and receiving of data do not need to be carried out under strict time synchronization. Different from synchronous data transmission, asynchronous transmission uses specific control signals (usually start bits and stop bits) at the beginning and end of data transmission so that the receiving device can correctly identify the boundaries of the data.

[0086] In practical applications, both the Build task and the Load task are heavy operations, that is, they will occupy a large amount of computing power resources and consume a large amount of time, thus interfering with the data link and causing a delay in returning the running status data to the client. In such a case, an asynchronous return method is adopted to improve the flexibility of returning the running status data, reduce the interference of heavy operations on the data link, thereby reducing the data transmission delay, and ensuring that the client can receive the running status data representing the running situation of the DiskANN object in a timely manner.

[0087] In some embodiments of the present invention, optionally, the task has multiple task types;

[0088] Step S2 includes:

[0089] Step S21: Divide the data link into multiple queues corresponding to the task types;

[0090] Step S22: Return the running status to the client through the corresponding queue according to the task type.

[0091] The task types in the present invention refer to: creating a DiskANN object (abbreviated as Create), importing vector data into the DiskANN object (abbreviated as ImportData), importing the vector data into the DiskANN service node (abbreviated as PushData), building a graph index for the imported vector data (abbreviated as Bulid), loading the graph index into memory (abbreviated as Load), querying vector data based on the graph index (abbreviated as Search), closing the DiskANN object (abbreviated as Close), and destroying the DiskANN object (abbreviated as Destroy), etc.

[0092] Through step S21, the data link is divided into multiple queues corresponding to the task types, and through step S22, the running status is returned to the client through the corresponding queues according to the task types, which can avoid the interference of the running status of the DiskANN object when running tasks of different task types and ensure that the client can receive the running status in a timely manner.

[0093] In some embodiments of the present invention, optionally, the monitoring instruction is further used to indicate monitoring whether an error occurs when the DiskANN object is running a task;

[0094] After step S1, it includes:

[0095] Step S3: In response to the monitoring instruction, monitor whether an error occurs when the DiskANN object is running a task;

[0096] Step S4: In the case where it is monitored that the DiskANN object has an error, record the running error information of the DiskANN object and return the running error information to the client through the data link.

[0097] In some optional embodiments, the running status and the running error information are returned to the client through the same data packet.

[0098] In some optional embodiments, the error is an error generated when running tasks of various task types. It can be understood that in some embodiments, the error determination rule can be defined by the user according to their own needs. For example, the user sets a data import rate threshold, and when the import rate calculated based on the amount of real-time imported data is less than the data import rate threshold, it is determined that the DiskANN object has an error. In this example, the running error information may include a schematic diagram of the change trend of the import rate within a preset time period, which can be the state parameters of the DiskANN object that has an error within the preset time period, so as to facilitate the user to analyze the cause of the error based on the running error information.

[0099] In some alternative embodiments, the running error information may include the state parameters of the DiskANN object where the error occurred within a preset time period. The state parameters may include at least one of the following: dataset information (e.g., the number of data points, the dimension of each data point), index structure (e.g., the index type used (e.g., graph-based index or tree structure), the construction parameters of the index (such as the number of neighbors, distance metric, etc.)), query parameters (e.g., the number of queries, the type of queries), performance metrics (e.g., query time, memory usage, recall rate, and precision rate (if there is labeled data)), status flags (e.g., whether it has been initialized, whether data has been loaded, whether a query is being executed), and log information (e.g., the timestamp of the operation, error and warning messages).

[0100] In some embodiments of the present invention, optionally, the running state further includes at least one of the following:

[0101] The maximum number of asynchronous input / output operations AIO, memory usage, disk usage, and central processing unit CPU load.

[0102] The present invention provides a monitoring device for the computing power operation tasks of an intelligent computing center, which is applied to an Index service node. Refer to Figure 3 as shown Figure 3 is a schematic block diagram of the monitoring device for the computing power operation tasks of the intelligent computing center. The monitoring device 30 for the computing power operation tasks of the intelligent computing center includes:

[0103] A receiving module 31, configured to receive a monitoring instruction sent by a client, where the monitoring instruction is used to indicate monitoring the running state of a DiskANN object of a running task that is currently invoking the computing power resources of the intelligent computing center;

[0104] An execution module 32, configured to create a data link between the DiskANN service node and the client in response to the monitoring instruction, and return the running state of the DiskANN object obtained from the DiskANN service node to the client in real time through the data link.

[0105] In some embodiments of the present invention, optionally, the task includes: a task of importing data;

[0106] The running state includes the amount of data imported.

[0107] In some embodiments of the present invention, optionally, the task includes: a task of creating a DiskANN object, a task of loading a DiskANN object;

[0108] The execution module 32 is further configured to return the running state to the client in real time in an asynchronous return manner.

[0109] In some embodiments of the present invention, optionally, the task has multiple task types;

[0110] The execution module 32 is further configured to divide the data link into multiple queues corresponding to the task types;

[0111] The execution module 32 is further configured to return the running state to the client through the corresponding queue according to the task type.

[0112] In some embodiments of the present invention, optionally, the monitoring instruction is further used to instruct to monitor whether an error occurs when the DiskANN object runs a task;

[0113] The execution module 32 is further configured to monitor whether an error occurs when the DiskANN object runs a task in response to the monitoring instruction;

[0114] The execution module 32 is further configured to record the running error information of the DiskANN object and return the running error information to the client through the data link when it is monitored that the DiskANN object has an error.

[0115] In some embodiments of the present invention, optionally, the running state further includes at least one of the following:

[0116] The maximum number of asynchronous input / output operations AIO, memory usage, disk usage, and central processing unit CPU load.

[0117] The monitoring device for the computing power operation task of the intelligent computing center provided by the present invention can implement Figures 1 to 2 each process implemented by the method embodiments and achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0118] The present invention provides an electronic device 40. Refer to Figure 4 as shown, Figure 4 which is a schematic block diagram of the electronic device 40 of the present invention, including a processor 41, a memory 42, and a program or instruction stored on the memory 42 and executable on the processor 41. When the program or instruction is executed by the processor, the steps in any one of the monitoring methods for the computing power operation task of the intelligent computing center of the present invention are implemented.

[0119] The present invention provides a readable storage medium with a program or instruction stored thereon. When the program or instruction is executed by a processor, each process of the embodiment of the monitoring method for the computing power operation task of the intelligent computing center as described above is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0120] Among them, the readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc. In some examples, the readable storage medium may be a non-transitory readable storage medium.

[0121] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, each process of the monitoring method embodiment of the computing power operation task of the intelligent computing center described in any one of the above is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0122] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.

[0123] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0124] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose and scope protected by the claims of the present invention, and all of them belong to the protection scope of the present invention.

Claims

1. A method for monitoring computing power operation tasks of an intelligent computing center, characterized in that: Applied to Index service nodes, including: Step S1: receiving a monitoring instruction sent by a client, wherein the monitoring instruction is used to instruct to monitor the running status of the DiskANN object of the running task that is currently calling the computing power resources of the intelligent computing center; Step S2: In response to the monitoring instruction, a data link is created between the DiskANN service node and the client, and the running status of the DiskANN object obtained from the DiskANN service node is returned to the client in real time through the data link.

2. The method for monitoring computing power operation tasks of an intelligent computing center according to claim 1, characterized in that: The tasks include: a task of importing data; The running status includes the data volume of the imported data.

3. The method for monitoring computing power operation tasks of an intelligent computing center according to claim 1, characterized in that: The tasks include: creating a DiskANN object and loading a DiskANN object; Returning the running status to the client in real time includes: The running status is returned to the client in real time in an asynchronous return manner.

4. The method for monitoring computing power operation tasks of an intelligent computing center according to claim 1, characterized in that: The tasks have multiple task types; The step S2 comprises: Step S21: dividing the data link into a plurality of queues corresponding to the task types; Step S22: according to the task type, the running status is returned to the client through the corresponding queue.

5. The method for monitoring the computing power operation tasks of the intelligent computing center according to claim 1 is characterized in that: The monitoring instruction is also used to instruct to monitor whether an error occurs when the DiskANN object runs a task; The step S1 then comprises: Step S3: In response to the monitoring instruction, monitoring whether an error occurs when the DiskANN object runs the task; Step S4: When an error occurs in the DiskANN object, the operation error information of the DiskANN object is recorded, and the operation error information is returned to the client through the data link.

6. The method for monitoring computing power operation tasks of an intelligent computing center according to any one of claims 1 to 5, characterized in that: The operating state also includes at least one of the following: Maximum number of asynchronous input / output operations (AIO), memory usage, disk usage, and CPU load.

7. A monitoring device for computing power operation tasks of an intelligent computing center, characterized in that: Applied to Index service nodes, including: A receiving module, used to receive a monitoring instruction sent by a client, wherein the monitoring instruction is used to instruct the running status of the DiskANN object that monitors the running task that is currently calling the computing power resources of the intelligent computing center; An execution module is used to create a data link between the DiskANN service node and the client in response to the monitoring instruction, and return the running status of the DiskANN object obtained from the DiskANN service node to the client in real time through the data link.

8. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps in the method for monitoring the computing power operation tasks of the intelligent computing center as described in any one of claims 1 to 6 are implemented.

9. A readable storage medium, characterized in that: The readable storage medium stores programs or instructions, and when the programs or instructions are executed by the processor, the steps in the monitoring method of the computing power operation task of the intelligent computing center as described in any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the method for monitoring the computing power operation tasks of an intelligent computing center as described in any one of claims 1 to 6.