Computing power operation task fault tolerance method and device of intelligent computing center

By verifying the computing power resources of parallel tasks in the intelligent computing center and aborting the task inadequate resources, the system crash problem is solved and the task operation efficiency and security is improved.

CN120295840APending Publication Date: 2025-07-11DATACANVAS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510458109.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When the Intelligent Computing Center runs multiple tasks of DiskANN objects in parallel, insufficient computing resources lead to a crash in the system. After recovery, it needs to re-run the completed tasks, which is inefficient.

Method used

Verify the computing power resources before running tasks in parallel. If there is insufficient, some tasks will be aborted and error information will be generated, resources that have completed tasks will be released, and idle state resources will be dynamically updated to ensure safe and stable operation.

Benefits of technology

It effectively avoids system crashes, improves task operation efficiency, and achieves high fault tolerance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295840A_ABST
    Figure CN120295840A_ABST
Patent Text Reader

Abstract

The invention provides a computing power operation task fault tolerance method and device for an intelligent computing center, and the method comprises the steps: S1, verifying whether the intelligent computing center can provide enough computing power resources for the parallel operation of a plurality of tasks or not before the parallel operation of the plurality of tasks for a DiskANN object; s2, if the verification result indicates that the intelligent computing center cannot provide enough computing power resources for parallel operation of multiple tasks at present, calling the computing power resources of the intelligent computing center to operate a part of tasks which can be supported by the current computing power resources to operate, stopping operation of the other part of tasks, generating error reporting information, and sending the error reporting information to the intelligent computing center; sending the error report information to a client associated with the user; and S3, if the verification result indicates that enough computing power resources can be provided, calling the computing power resources of the intelligent computing center to run the plurality of tasks in parallel. According to the invention, safe and stable operation of the intelligent computing center can be ensured, and high fault-tolerant capability is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure technologies, and in particular to a fault tolerance method and device for computing power operation tasks of an intelligent computing center. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training and model inference, etc.) by using large-scale heterogeneous computing power resources, including general computing power and intelligent computing power. The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".

[0005] An "intelligent computing center" is an artificial intelligence computing center, which is a type of computing power infrastructure that provides computing power services, data services and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", which is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing information data, and a new type of productive force integrating information computing power, network carrying capacity and data storage capacity, and mainly provides services to the society through computing power infrastructure.

[0007] Since the emergence of intelligent computing centers, the problem of low fault tolerance of computing power operation tasks in intelligent computing centers is a problem to be solved. When the intelligent computing center runs multiple tasks in parallel for DiskANN objects, it needs to occupy a large amount of computing power resources. If the computing power resources are insufficient during the execution process, it will cause the system of the intelligent computing center to crash. Users need to spend a lot of time and energy to achieve recovery, and often need to re-run the tasks that have been completed before the crash after recovery, and the efficiency of running tasks is very low. Summary of the Invention

[0008] The present invention provides a method and device for fault tolerance of computing power operation tasks in an intelligent computing center to solve the problem of low fault tolerance of computing power operation tasks in the intelligent computing center since the emergence of the intelligent computing center. When the intelligent computing center runs multiple tasks for DiskANN objects in parallel, a large amount of computing power resources are required. If the computing power resources are insufficient during the execution process, the system of the intelligent computing center will crash, and users need to spend a lot of time and energy to recover. Moreover, often after recovery, the tasks that have been completed before the crash need to be run again, and the efficiency of running tasks is very low.

[0009] To solve the above technical problems, the present invention is implemented as follows:

[0010] In a first aspect, the present invention provides a method for fault tolerance of computing power operation tasks in an intelligent computing center, including:

[0011] Step S1: Before running multiple tasks for a DiskANN object in parallel, check whether the intelligent computing center can currently provide sufficient computing power resources for parallel running of the multiple tasks to obtain a check result;

[0012] Step S2: If the check result indicates that the intelligent computing center cannot currently provide sufficient computing power resources for parallel running of the multiple tasks, call the computing power resources of the intelligent computing center to run a part of the tasks that the current computing power resources can support, abort the running of the other part of the tasks, generate an error message, and send the error message to the client associated with the user;

[0013] Step S3: If the check result indicates that the intelligent computing center can currently provide sufficient computing power resources for parallel running of the multiple tasks, call the computing power resources of the intelligent computing center to run the multiple tasks in parallel.

[0014] Optionally, the computing power resources include: asynchronous input / output AIO resources;

[0015] The step S1 includes:

[0016] Step S11: Determine whether the number of currently idle AIOs on the intelligent computing center is greater than or equal to the target number of AIOs required for parallel running of the multiple tasks;

[0017] Step S12: If the number of currently idle AIOs on the intelligent computing center is greater than or equal to the target number of AIOs, determine that the check result is that the intelligent computing center can currently provide sufficient AIO resources for parallel running of the multiple tasks;

[0018] Step S13: If the number of AIOs in the intelligent computing center that are currently in an idle state is less than the target number of AIOs, determine that the verification result is that the intelligent computing center currently cannot provide sufficient AIO resources for parallel running of the multiple tasks.

[0019] Optionally, it includes:

[0020] Upon completion of each of the tasks, control the intelligent computing center to release the AIO resources occupied by the completed task during operation, and update the number of idle AIOs that can be used for running tasks.

[0021] The step S13 includes:

[0022] Step S131: Every preset waiting time, return to the step S11 until the number of AIOs in the intelligent computing center that are currently in an idle state is greater than or equal to the target number of AIOs.

[0023] Optionally, the step S11 includes:

[0024] Step S01: Multiply the number of tasks of the multiple tasks by the preset number of AIOs required to run a single task to obtain the target number of AIOs.

[0025] Optionally, the task includes at least one of the following: creation task, data import task, construction task, loading task, query task, shutdown task, destruction task.

[0026] Optionally, the computing power resources include at least one of the following: memory resources, AIO resources.

[0027] In a second aspect, the present invention provides a computing power operation task fault tolerance device for an intelligent computing center, including:

[0028] A verification module, configured to verify whether the intelligent computing center can currently provide sufficient computing power resources for parallel running of the multiple tasks before parallel running of the multiple tasks for the DiskANN object, to obtain a verification result;

[0029] An execution module, configured to, if the verification result indicates that the intelligent computing center currently cannot provide sufficient computing power resources for parallel running of the multiple tasks, call the computing power resources of the intelligent computing center to run a part of the tasks that the current computing power resources can support, abort the running of the other part of the tasks, generate an error message, and send the error message to the client associated with the user.

[0030] The execution module is further configured to, if the verification result indicates that the intelligent computing center can currently provide sufficient computing power resources for parallel execution of the multiple tasks, call the computing power resources of the intelligent computing center to parallelly execute the multiple tasks.

[0031] In a third aspect, the present invention provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps in the method for fault tolerance of computing power operation tasks of the intelligent computing center according to any one of the first aspects are implemented.

[0032] In a fourth aspect, the present invention provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps in the method for fault tolerance of computing power operation tasks of the intelligent computing center according to any one of the first aspects are implemented.

[0033] In a fifth aspect, the present invention provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the steps in the method for fault tolerance of computing power operation tasks of the intelligent computing center according to any one of the first aspects are implemented.

[0034] In the present invention, through step S1: before parallelly executing multiple tasks for a DiskANN object, verifying whether the intelligent computing center can currently provide sufficient computing power resources for parallel execution of multiple tasks to obtain a verification result; step S2: if the verification result indicates that the intelligent computing center cannot currently provide sufficient computing power resources for parallel execution of multiple tasks, calling the computing power resources of the intelligent computing center to execute a part of the tasks that the current computing power resources can support, aborting the execution of the other part of the tasks, generating an error message, and sending the error message to the client associated with the user; step S3: if the verification result indicates that the intelligent computing center can currently provide sufficient computing power resources for parallel execution of multiple tasks, calling the computing power resources of the intelligent computing center to parallelly execute multiple tasks, the present invention can effectively avoid system crashes caused by insufficient computing power resources during parallel execution of tasks, ensure the safe and stable operation of the intelligent computing center, improve the efficiency of executing tasks, and achieve high fault tolerance for computing power operation tasks of the intelligent computing center. Description of the Drawings

[0035] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0036] Figure 1 It is a schematic diagram of the overall process of data interaction for DiskANN;

[0037] Figure 2 It is a schematic flow chart of the fault tolerance method for the computing power operation task of the intelligent computing center of the present invention;

[0038] Figure 3 It is a schematic diagram of multiple loading tasks running in parallel;

[0039] Figure 4 It is a principle block diagram of the fault tolerance device for the computing power operation task of the intelligent computing center of the present invention;

[0040] Figure 5 It is a principle block diagram of the electronic device of the present invention. Detailed implementation manners

[0041] Next, the technical solutions in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0042] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "or" in the present invention means at least one of the connected objects. For example, "A or B" covers three scenarios, namely, Scenario 1: including A and not including B; Scenario 2: including B and not including A; Scenario 3: including both A and B. The character " / " generally indicates an "or" relationship between the associated objects before and after.

[0043] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0044] Next, a brief description will be given first to the technical terms involved in the present invention.

[0045] The "computing power" described in the present invention refers to: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to output a target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.

[0046] The "Computational Power (CP)" described in the present invention refers to: the ability of a data center server to process data and output results, which is a comprehensive indicator for measuring the computing power of a data center and includes general computing power, supercomputing power, and intelligent computing power. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP_general + CP_intelligent + CP_super.

[0047] The "Network Power (NP)" described in the present invention refers to: the performance of the data transmission capacity of computing power facilities, which is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., and involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling ability.

[0048] The "Storage Power (SP)" described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, which is a comprehensive indicator for measuring the data storage ability of a data center and includes external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0049] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure integrating information computing power, network carrying power, and data storage power, which can realize the centralized computing, storage, transmission, and application of information.

[0050] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing. With the emergence and popularization of new general technologies, the form of the new type of information infrastructure will be more diverse.

[0051] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and supercomputing power.

[0052] The "general computing power" described in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0053] The "intelligent computing power" described in the present invention refers to a computing platform deployed on a large scale for various artificial intelligence innovation applications based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, and so on.

[0054] The "super computing power" described in the present invention mainly refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.

[0055] The "intelligent computing center" described in the present invention refers to a facility that mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training, and model inference scenarios) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0056] The "intelligent computing center" described in the present invention includes, but is not limited to, the "intelligent computing center".

[0057] The "intelligent computing center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.

[0058] The "computing power center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, and having computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0059] The "supercomputing center" described in the present invention, that is, the supercomputing data center, is a data center based on supercomputers or large-scale computing clusters, which can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0060] The "computing power resources" described in the present invention refer to technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0061] The "DiskANN" described in the present invention is a vector retrieval engine based on distributed storage, capable of storing and retrieving vector data at the billion level on a single computer. Compared with traditional vector retrieval algorithms, DiskANN has higher storage efficiency and faster retrieval speed.

[0062] The "computing power operation task" described in the present invention refers to specific workloads or jobs that are executed on computing power resources and require a certain amount of computing power support, usually involving scenarios such as complex data processing, numerical calculation, model training, or simulation.

[0063] In the present invention, the overall process of data interaction based on DiskANN can be referred to Figure 1 , and this overall process mainly includes the following tasks: creating a DiskANN object (abbreviated as Create), importing vector data into the DiskANN object (abbreviated as ImportData), importing the vector data into the DiskANN service node (abbreviated as Push Data), building a graph index for the imported vector data (abbreviated as Bulid), loading the graph index into memory (abbreviated as Load), querying vector data based on the graph index (abbreviated as Search), closing the DiskANN object (abbreviated as Close), and destroying the DiskANN object (abbreviated as Destroy), etc.

[0064] Among them, creating a DiskANN object means that the DiskANN service node creates a DiskANN object according to the client's Create request. When the DiskANN object is created, there is no data in it, and it is in a data-free state and cannot provide any services to the outside. The client's Create request is sent to the DiskANN service node through the Index service node. After the DiskANN service node creates the DiskANN object, it notifies the client through the Index service node.

[0065] In the task of importing vector data for the DiskANN object, the client sends an ImportData request containing vector data to the Index service node, and the Index service node temporarily stores the vector data in the ImportData request in the database of the Index service node. After the Index service node finishes importing the data, it notifies the client.

[0066] After the data import task, the client can send a Build request to the Index service node. Based on the Build request, the Index service node executes the Push Data task, that is, the Index service node imports the vector data stored in the database of the Index service node into the DiskANN service node, and the DiskANN service node notifies the Index service node after the vector data is written to disk to form a file.

[0067] The Build task means that the DiskANN service node processes the vector data of the imported DiskANN object based on the Build request sent by the Index service node, obtains the graph index of the vector data, and saves the vector data and the graph index. After the DiskANN service node finishes the Build task, it notifies the client through the Index service node.

[0068] The Load task means that the client sends a load request to the DiskANN service node through the Index service node, and the DiskANN service node loads at least part of the graph index of the vector data of the DiskANN object into memory based on the load request. After the loading is completed, the DiskANN service node notifies the client through the Index service node.

[0069] The Search task means that the client sends a vector data query request to the DiskANN service node through the Index service node, and the DiskANN service node retrieves the vector data most similar to the vector data to be queried based on the graph index, and returns the retrieved most similar vector data to the client through the Index service node.

[0070] The Close task means that the client sends a close request for the DiskANN object to the DiskANN service node through the Index service node, and the DiskANN service node closes the DiskANN object based on the close request. After the closing is completed, the DiskANN service node notifies the client through the Index service node.

[0071] The "Destroy" task means that the client sends a destruction request of the DiskANN object to the DiskANN service node through the Index service node, and the DiskANN service node destroys the DiskANN object based on this destruction request. After the destruction is completed, the DiskANN service node notifies the client through the Index service node.

[0072] The above process involves three execution entities: the client, the Index service node, and the DiskANN service node. Among them, the purpose of using the Index service node is to make the DiskANN service node as lightweight as possible. The DiskANN service node only processes the core DiskANN business, and outsources other services to the Index service node to ensure the stability and reliability of the overall service. Before executing the above process, the Index service node and the DiskANN service node need to establish a handshake connection in advance to provide services for subsequent operations. The above Index service node and DiskANN service node can be deployed on different physical machines or on the same physical machine, and the present invention does not limit this.

[0073] When multiple tasks for DiskANN objects are running in parallel, a large amount of computing power resources are required. If there is insufficient computing power resources during the execution process, it will cause the system of the intelligent computing center to crash. Users need to spend a lot of time and effort to recover, and often need to re-run the tasks that have been completed before the crash after recovery, resulting in very low operating efficiency.

[0074] The present invention provides a method for fault tolerance of computing power operation tasks in an intelligent computing center. See Figure 2 as shown Figure 2 is a schematic flowchart of the method for fault tolerance of computing power operation tasks in the intelligent computing center of the present invention, including:

[0075] Step S1: Before running multiple tasks for DiskANN objects in parallel, check whether the intelligent computing center can currently provide sufficient computing power resources for running multiple tasks in parallel to obtain a check result;

[0076] Step S2: If the check result indicates that the intelligent computing center currently cannot provide sufficient computing power resources for running multiple tasks in parallel, call the computing power resources of the intelligent computing center to run a part of the tasks that the current computing power resources can support, abort the running of the other part of the tasks, generate an error message, and send the error message to the client associated with the user;

[0077] Step S3: If the check result indicates that the intelligent computing center can currently provide sufficient computing power resources for running multiple tasks in parallel, call the computing power resources of the intelligent computing center to run multiple tasks in parallel.

[0078] In the present invention, as shown in Figure 1 below, the task may include at least one of the following: creating a DiskANN object (abbreviated as Create), importing vector data into the DiskANN object (abbreviated as ImportData), importing vector data into the DiskANN service node (abbreviated as Push Data), building a graph index for the imported vector data (abbreviated as Bulid), loading the graph index into memory (abbreviated as Load), querying vector data based on the graph index (abbreviated as Search), closing the DiskANN object (abbreviated as Close), and destroying the DiskANN object (abbreviated as Destroy).

[0079] In step S2 of the present invention, if the verification result indicates that the intelligent computing center currently cannot provide sufficient computing power resources to run multiple tasks in parallel, first run a part of the multiple tasks whose running requirements can be met by the current computing power resources, and the remaining part of the multiple tasks will be aborted. This can effectively avoid the risk of system crash caused by running multiple tasks in parallel, and further avoid the situation where tasks that have been completed before the crash need to be re-run after recovery, improving the running efficiency. It can be understood that when the intelligent computing center currently cannot provide sufficient computing power resources to run multiple tasks in parallel, instead of completely aborting the running of all tasks, first run a part of the multiple tasks whose running requirements can be met by the current computing power resources, which improves the running efficiency.

[0080] It should be noted that for the selection of the part of the tasks to be run first, the multiple tasks that need to be run in parallel can be converted into multiple tasks to be run serially. That is to say, sort the multiple tasks and verify one by one in the sorted order whether the computing power resources can meet the running requirements until the part of the tasks to be run first is determined. For example, if there are 5 tasks that need to be run in parallel, including: Task 1, Task 2, Task 3, Task 4, Task 5, first verify whether the current computing power resources can meet the running requirements of Task 1. If so, call the computing power resources of the intelligent computing center to run Task 1. Then, verify whether the current computing power resources (at this time, the current computing power resources are the computing power resources before running Task 1 minus the computing power resources occupied by running Task 1) can meet the running requirements of Task 2. If so, call the computing power resources of the intelligent computing center to run Task 2, and so on. It should be noted that there is also a situation. For example, before verifying whether the current computing power resources can meet the running requirements of Task 5, Task 1 has been completed and the occupied computing power resources have been released. Then, the current computing power resources used to verify Task 5 are the computing power resources before running Task 4 minus the computing power resources occupied by running Task 4, plus the computing power resources occupied by running Task 1. That is to say, the current computing power resources are dynamically changing.

[0081] In some alternative embodiments, for the tasks of another part that have been suspended from running, they can be run after the tasks of the first - run part are completed and the computing power resources are released.

[0082] It should be noted that the error message can include the reason for the error. For example, "The intelligent computing center currently cannot provide sufficient computing power resources to run multiple tasks in parallel." The error message can also include pre - set error suggestions for the error reason, such as a suggestion to increase computing power resources. In the present invention, by generating the error message in step S2 and sending the error message to the client associated with the user, the user can clearly know the reason for the error, laying a solid foundation for solving the error problem and realizing the quick restart of parallel - running tasks.

[0083] In the present invention, through step S1: Before parallel - running multiple tasks for the DiskANN object, check whether the intelligent computing center can currently provide sufficient computing power resources to run multiple tasks in parallel to obtain a check result; step S2: If the check result indicates that the intelligent computing center currently cannot provide sufficient computing power resources to run multiple tasks in parallel, call the computing power resources of the intelligent computing center to run a part of the tasks that the current computing power resources can support, suspend the running of the other part of the tasks, generate an error message, and send the error message to the client associated with the user; step S3: If the check result indicates that the intelligent computing center can currently provide sufficient computing power resources to run multiple tasks in parallel, call the computing power resources of the intelligent computing center to run multiple tasks in parallel. The present invention can effectively avoid system crashes caused by insufficient computing power resources when running tasks in parallel, ensure the safe and stable operation of the intelligent computing center, improve the efficiency of running tasks, and achieve high fault - tolerance capabilities for the computing power operation tasks of the intelligent computing center.

[0084] In some embodiments of the present invention, optionally, the computing power resources include: Asynchronous Input / Output (AIO) resources;

[0085] Step S1 includes:

[0086] Step S11: Determine whether the number of AIOs that are currently in an idle state on the intelligent computing center is greater than or equal to the target number of AIOs required to run multiple tasks in parallel;

[0087] Step S12: If the number of AIOs that are currently in an idle state on the intelligent computing center is greater than or equal to the target number of AIOs, determine that the check result is that the intelligent computing center can currently provide sufficient AIO resources to run multiple tasks in parallel;

[0088] Step S13: If the number of AIOs that are currently in an idle state on the intelligent computing center is less than the target number of AIOs, determine that the check result is that the intelligent computing center currently cannot provide sufficient AIO resources to run multiple tasks in parallel.

[0089] In the present invention, the AIO resource, namely Linux AIO. AIO (Asynchronous Input / Output) is an interface that allows an application to submit multiple I / O requests in parallel without incurring the thread overhead for each request. The main difference between it and synchronous I / O is that synchronous I / O must wait for the kernel to complete the I / O operation before returning, while asynchronous I / O does not have to wait for the I / O operation to complete. Instead, it initiates an I / O operation to the kernel and immediately returns. When the kernel completes the I / O operation, it notifies the application through signals or callbacks.

[0090] Linux AIO is a native asynchronous I / O interface provided by the Linux kernel and became a standard feature in the Linux 2.6 kernel version. Currently, the AIO interface is most suitable for directly accessing raw block devices such as disks, flash drives, or storage arrays using "O_DIRECT". When using Linux AIO, the following functions are generally used:

[0091] io_setup: Open an I / O context to submit and retrieve I / O requests.

[0092] io_submit: Submit requests to the I / O context, which sends them to the device driver for processing on the device.

[0093] io_getevents: Retrieve completed I / O requests from the I / O context in the form of event completion objects.

[0094] io_destroy: Destroy the I / O context.

[0095] Although Linux AIO is also asynchronous, it may still block and its behavior in some cases is unpredictable. Moreover, it only supports storage files in direct I / O mode.

[0096] It is mainly used in the specific field of databases. In contrast, io_uring, first introduced in the Linux 5.1 kernel in 2019, is a more advanced and flexible asynchronous I / O framework. It supports both storage files and network files, and also supports more asynchronous system calls. It is truly asynchronous I / O in design.

[0097] In the present invention, the AIO in the idle state, namely the AIO that can be used to run tasks.

[0098] It should be noted that the number of AIOs that are currently in an idle state on the intelligent computing center is greater than or equal to the target number of AIOs, indicating that the number of AIOs that can be used to run tasks is greater than or equal to the target number of AIOs. The intelligent computing center can currently provide sufficient AIO resources for parallel running of multiple tasks. If the number of AIOs that are currently in an idle state on the intelligent computing center is less than the target number of AIOs, it means that the number of AIOs that can be used to run tasks is less than the target number of AIOs, and the intelligent computing center currently cannot provide sufficient AIO resources for parallel running of multiple tasks.

[0099] In the present invention, through step S11: determining whether the number of AIOs that are currently in an idle state on the intelligent computing center is greater than or equal to the target number of AIOs required for parallel running of multiple tasks; step S12: if the number of AIOs that are currently in an idle state on the intelligent computing center is greater than or equal to the target number of AIOs, determining that the verification result is that the intelligent computing center can currently provide sufficient AIO resources for parallel running of multiple tasks; step S13: if the number of AIOs that are currently in an idle state on the intelligent computing center is less than the target number of AIOs, determining that the verification result is that the intelligent computing center currently cannot provide sufficient AIO resources for parallel running of multiple tasks. Before parallel running multiple tasks for the DiskANN object, the judgment on whether the intelligent computing center can currently provide sufficient AIO resources for parallel running of multiple tasks is realized, which can effectively avoid system crashes caused by insufficient AIO resources during parallel running of tasks, ensure the safe and stable operation of the intelligent computing center, and improve the running efficiency of tasks.

[0100] In some embodiments of the present invention, optionally, it includes:

[0101] After each task is completed, control the intelligent computing center to release the AIO resources occupied by the completed task during operation, and update the number of idle AIOs that can be used to run tasks;

[0102] Step S13 includes:

[0103] Step S131: Every preset waiting time, return to step S11 until the number of AIOs that are currently in an idle state on the intelligent computing center is greater than or equal to the target number of AIOs.

[0104] It should be noted that the present invention also realizes a dynamic update mechanism for AIOs, that is, after each task is completed, control the intelligent computing center to release the AIO resources occupied by the completed task during operation, and update the number of idle AIOs that can be used to run tasks, improving the utilization efficiency of AIO resources, thereby being able to accelerate the replenishment of the number of idle AIOs to be greater than or equal to the target number of AIOs, shortening the time for aborting the parallel running of multiple tasks, and improving the running efficiency of tasks.

[0105] The present invention also returns to step S11 every preset waiting time through step S131 until the number of AIOs in the idle state on the intelligent computing center is greater than or equal to the target number of AIOs, realizing the automatic judgment of the number of AIOs in the idle state, ensuring that the processes of the other part of the tasks that have been aborted can be restored to operation in a timely manner, and improving the efficiency of running tasks.

[0106] In some embodiments of the present invention, optionally, step S11 includes:

[0107] Step S01: Multiply the number of tasks of multiple tasks by the preset number of AIOs required to run a single task to obtain the target number of AIOs.

[0108] In the present invention, the number of AIOs required for a single task is preset, which improves the efficiency of determining the target number of AIOs for each process running multiple tasks in parallel, thereby promoting the efficient execution of steps S1 to S3 of the present invention and improving the efficiency of running tasks.

[0109] The specific presetting method can be preset in the configuration file. For the number of AIOs required for a single task, the number of AIOs required for different tasks can be set separately. For example, as shown in Figure 3 As shown, multiple tasks running in parallel are all load tasks. For load tasks, the number of AIOs required for a single task can be 64 * 1024 = 7705, that is, the number of AIOs consumed by each load task is 77056.

[0110] In some embodiments of the present invention, optionally, the tasks include at least one of the following: creation task, data import task, construction task, load task, query task, close task, and destruction task.

[0111] In the present invention, as shown in Figure 1 As shown, the creation task is to create a DiskANN object (abbreviated as Create). The data import task includes vector data import for the DiskANN object (abbreviated as ImportData), and pushing the vector data into the DiskANN service node (abbreviated as Push Data). The construction task is to build a graph index for the imported vector data (abbreviated as Bulid). The load task is to load the graph index into memory (abbreviated as Load). The query task is to query vector data based on the graph index (abbreviated as Search). The close task is to close the DiskANN object (abbreviated as Close); the destruction task is to destroy the DiskANN object (abbreviated as Destroy).

[0112] In some embodiments of the present invention, optionally, the computing power resources include at least one of the following: memory resources, AIO resources.

[0113] It should be noted that when the computing power resources include memory resources and AIO resources, to verify whether the intelligent computing center can currently provide sufficient computing power resources for parallel running of multiple tasks, it is necessary to verify whether both the memory resources and the AIO resources meet the requirements for parallel running of multiple tasks. In the case where both are satisfied, proceed to step S3; in other cases except where both are satisfied, proceed to step S2.

[0114] The present invention provides a fault-tolerant device for computing power operation tasks of an intelligent computing center. Refer to Figure 4 as shown in Figure 4 which is a principle block diagram of the fault-tolerant device for computing power operation tasks of the intelligent computing center of the present invention. The fault-tolerant device 30 for computing power operation tasks of the intelligent computing center includes:

[0115] A verification module 31, configured to verify whether the intelligent computing center can currently provide sufficient computing power resources for parallel running of the multiple tasks before parallel running of multiple tasks for a DiskANN object, and obtain a verification result;

[0116] An execution module 32, configured to, if the verification result indicates that the intelligent computing center cannot currently provide sufficient computing power resources for parallel running of the multiple tasks, call the computing power resources of the intelligent computing center to run a part of the tasks that the current computing power resources can support, abort the running of the other part of the tasks, generate an error message, and send the error message to the client associated with the user;

[0117] The execution module 32 is further configured to, if the verification result indicates that the intelligent computing center can currently provide sufficient computing power resources for parallel running of the multiple tasks, call the computing power resources of the intelligent computing center to parallel run the multiple tasks.

[0118] In some embodiments of the present invention, optionally, the computing power resources include: asynchronous input / output AIO resources;

[0119] The verification module 31 is further configured to determine whether the number of currently idle AIOs on the intelligent computing center is greater than or equal to the target number of AIOs required for parallel running of the multiple tasks;

[0120] The verification module 31 is further configured to, if the number of currently idle AIOs on the intelligent computing center is greater than or equal to the target number of AIOs, determine that the verification result is that the intelligent computing center can currently provide sufficient AIO resources for parallel running of the multiple tasks;

[0121] The verification module 31 is further configured to determine that the verification result is that the intelligent computing center currently cannot provide sufficient AIO resources for parallel running of the multiple tasks if the number of AIOs in the idle state on the intelligent computing center is less than the target number of AIOs.

[0122] In some embodiments of the present invention, optionally, upon completion of each of the tasks, the intelligent computing center is controlled to release the AIO resources occupied during the running of the completed tasks, and the number of AIOs in the idle state that can be used for running tasks is updated;

[0123] The execution module 32 is further configured to return to step S11 every preset waiting time until the number of AIOs in the idle state on the intelligent computing center is greater than or equal to the target number of AIOs.

[0124] In some embodiments of the present invention, optionally, the verification module 31 is further configured to multiply the number of tasks of the multiple tasks by a preset number of AIOs required for running a single task to obtain the target number of AIOs.

[0125] In some embodiments of the present invention, optionally, the tasks include at least one of the following: creation task, data import task, construction task, loading task, query task, shutdown task, destruction task.

[0126] In some embodiments of the present invention, optionally, the computing power resources include at least one of the following: memory resources, AIO resources.

[0127] The computing power operation task fault tolerance device of the intelligent computing center provided by the present invention can implement Figures 1 to 2 each process implemented by the method embodiments and achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0128] The present invention provides an electronic device 40. Refer to Figure 5 as shown. Figure 5 is a schematic block diagram of the electronic device 40 of the present invention, including a processor 41, a memory 42, and a program or instruction stored in the memory 42 and executable on the processor 41. When the program or instruction is executed by the processor, the steps in any one of the computing power operation task fault tolerance methods of the intelligent computing center of the present invention are implemented.

[0129] The present invention provides a readable storage medium with a program or instruction stored thereon. When the program or instruction is executed by a processor, each process of the embodiments of the computing power operation task fault tolerance method of the intelligent computing center as described above is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0130] Among them, the readable storage medium includes, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc. In some examples, the readable storage medium may be a non-transitory readable storage medium.

[0131] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, each process of the fault tolerance method for the computing power operation task of the intelligent computing center in any one of the above embodiments is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.

[0132] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including that element.

[0133] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0134] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention's purpose and claims, and all of them belong to the protection scope of the present invention.

Claims

1. A fault tolerance method for computing power operation tasks in an intelligent computing center, characterized in that, Including: Step S1: Before running multiple tasks for the DiskANN object in parallel, check whether the intelligent computing center can currently provide sufficient computing power resources for parallel running of the multiple tasks to obtain a check result; Step S2: If the check result indicates that the intelligent computing center cannot currently provide sufficient computing power resources for parallel running of the multiple tasks, call the computing power resources of the intelligent computing center to run a part of the tasks that the current computing power resources can support, abort running the other part of the tasks, generate an error message, and send the error message to the client associated with the user; Step S3: If the check result indicates that the intelligent computing center can currently provide sufficient computing power resources for parallel running of the multiple tasks, call the computing power resources of the intelligent computing center to run the multiple tasks in parallel.

2. The method for fault tolerance of computing power operation tasks of an intelligent computing center according to claim 1, characterized in that The computing power resources include: asynchronous input / output AIO resources; The step S1 includes: Step S11: Determine whether the number of currently idle AIOs on the intelligent computing center is greater than or equal to the target number of AIOs required for parallel running of the multiple tasks; Step S12: If the number of currently idle AIOs on the intelligent computing center is greater than or equal to the target number of AIOs, determine that the check result is that the intelligent computing center can currently provide sufficient AIO resources for parallel running of the multiple tasks; Step S13: If the number of currently idle AIOs on the intelligent computing center is less than the target number of AIOs, determine that the check result is that the intelligent computing center cannot currently provide sufficient AIO resources for parallel running of the multiple tasks.

3. The fault tolerance method for the computing power operation task of the intelligent computing center according to claim 2, wherein Including: For each completed task, control the intelligent computing center to release the AIO resources occupied during the running of the completed task, and update the number of idle AIOs that can be used for running tasks; The step S13 includes: Step S131: Return to step S11 every preset waiting time until the number of currently idle AIOs on the intelligent computing center is greater than or equal to the target number of AIOs.

4. The fault tolerance method for the computing power operation task of the intelligent computing center according to claim 2, wherein, The step S11 includes: Step S01: Multiply the number of tasks of the multiple tasks by the preset number of AIOs required for running a single task to obtain the target number of AIOs.

5. The method for fault tolerance of computing power operation tasks of an intelligent computing center according to claim 1, characterized in that The tasks include at least one of the following: create task, data import task, construction task, loading task, query task, close task, destroy task.

6. The method for fault tolerance of computing power operation tasks of an intelligent computing center according to claim 1, characterized in that The computing power resources include at least one of the following: memory resources, AIO resources.

7. A fault tolerance device for computing power operation tasks of an intelligent computing center, characterized in that, Including: A verification module, configured to verify whether the intelligent computing center can currently provide sufficient computing power resources for parallel running of multiple tasks for a DiskANN object before parallel running the multiple tasks, and obtain a verification result; An execution module, configured to, if the verification result indicates that the intelligent computing center cannot currently provide sufficient computing power resources for parallel running of the multiple tasks, call the computing power resources of the intelligent computing center to run a part of the tasks that the current computing power resources can support, abort running another part of the tasks, generate an error message, and send the error message to a client associated with the user; The execution module is further configured to, if the verification result indicates that the intelligent computing center can currently provide sufficient computing power resources for parallel running of the multiple tasks, call the computing power resources of the intelligent computing center to parallel run the multiple tasks.

8. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps in the method for fault tolerance of computing power operation tasks of the intelligent computing center according to any one of claims 1 to 6 are implemented.

9. A readable storage medium, characterized in that: A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, the steps in the method for fault tolerance of computing power operation tasks of the intelligent computing center according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that, It includes computer instructions. When the computer instructions are executed by a processor, the steps in the method for fault tolerance of computing power operation tasks of the intelligent computing center according to any one of claims 1 to 6 are implemented.