Computing power operation task fault repairing method and device of intelligent computing center
By determining the target file of the task data record of the DiskANN object failure in the intelligent computing center and starting to restore the task from the file, the problem of low task recovery rate in the existing technology is solved, and efficient task recovery is achieved.
Patent Information
- Application Number
- CN202510348173.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
AI Technical Summary
Existing DiskANN object running task failure repair methods require restarting the task from scratch, resulting in low recovery rates and inefficiency.
Avoid reruns from scratch by determining the target file recorded by the failed task data on the DiskANN object and starting to restore the task from the target file after the failure is fixed.
The task recovery rate is accelerated, the efficiency of running tasks is improved, and the rapid recovery of DiskANN object running tasks after fault repair is achieved.
Smart Images

Figure CN120216244A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure technologies, and particularly relates to a method and device for repairing faults in computing power operation tasks of an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power and intelligent computing power. An intelligent computing center encompasses facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] An "intelligent computing center" includes, but is not limited to, an "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of a target result by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.
[0007] In the existing method for repairing faults in DiskANN object running tasks, after task fault repair, the DiskANN object needs to restart the task interrupted by the fault, that is, the DiskANN object starts running the entire task interrupted by the fault from the beginning. Coupled with the fact that the DiskANN object running task requires a large amount of computing power resources (especially memory), the rate of the DiskANN object resuming the running task is low, resulting in very low efficiency of the running task. Summary of the Invention
[0008] The present invention provides a method and apparatus for repairing faults in the computing power operation tasks of an intelligent computing center, so as to solve the problem that in the existing method for repairing faults in the operation tasks of DiskANN objects, after the task fault is repaired, the DiskANN object needs to restart the task interrupted by the fault, that is, the DiskANN object runs the entire task interrupted by the fault from the beginning. Coupled with the fact that the operation tasks of the DiskANN object require a large amount of computing power resources (especially memory), the speed of the DiskANN object to resume running the task is low, resulting in very low efficiency of the running task.
[0009] In order to solve the above technical problems, the present invention is implemented as follows:
[0010] In a first aspect, the present invention provides a method for repairing faults in the computing power operation tasks of an intelligent computing center, including:
[0011] Step S1: In response to a running fault of the DiskANN object calling the computing power resources of the intelligent computing center to run a task, determine the target file in which the task data record interrupted by the running fault is located on the DiskANN object;
[0012] Step S2: In response to the repair of the running fault, start from the target file and resume the DiskANN object calling the computing power resources of the intelligent computing center to run the task.
[0013] Optionally, before the step S1, it includes:
[0014] Step S01: Determine the DiskANN object required for the running task;
[0015] Step S02: Load the DiskANN object on the DiskANN service node. After the loading is completed, instruct the DiskANN object to call the computing power resources of the intelligent computing center to run the task.
[0016] Optionally, a directory structure divided according to the task running stage is preset on the DiskANN object; during the running of the task, the DiskANN object records the task data corresponding to the task running stage into the directory structure to form a file;
[0017] The step S1 includes:
[0018] Step S11: Determine the first task running stage of the DiskANN object running task interrupted by the running fault;
[0019] Step S12: Query the directory structure according to the first task running stage to determine the target file.
[0020] Optionally, the step S2 includes:
[0021] Step S21: Determine whether the running failure causes the data in the target file to be untrusted, and obtain a determination result;
[0022] Step S22: If the determination result indicates that the running failure causes the data in the target file to be untrusted, re-run the tasks in the first task running stage, and record the task data during the re-run into the directory structure to form a first file, and replace the target file with the first file;
[0023] Step S23: If the determination result indicates that the running failure does not cause the data in the target file to be untrusted, run the tasks in the next stage after the first task running stage.
[0024] Optionally, the directory structure includes at least one of the following directories:
[0025] A directory for recording task data in the building stage, a directory for recording task data in the completed building stage, a directory for recording task data in the deletion stage, and a directory for recording object names without data.
[0026] In a second aspect, the present invention provides a device for repairing a computing power running task failure in an intelligent computing center, including:
[0027] A determination module, configured to determine a target file in which task data records are interrupted by the running failure on the DiskANN object in response to a running failure of a computing power resource running task called by the DiskANN object in the intelligent computing center;
[0028] An execution module, configured to resume running the tasks of the DiskANN object calling the computing power resources of the intelligent computing center starting from the target file in response to the running failure being repaired.
[0029] Optionally, it further includes:
[0030] A loading module, configured to determine the DiskANN object required for the running task;
[0031] The loading module is further configured to load the DiskANN object on the DiskANN service node, and after the loading is completed, instruct the DiskANN object to call the computing power resources of the intelligent computing center to run tasks.
[0032] In a third aspect, the present invention provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor, and when the program or instruction is executed by the processor, it implements the steps in the method for repairing a computing power running task failure in an intelligent computing center as described in any item of the first aspect.
[0033] In a fourth aspect, the present invention provides a readable storage medium, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps in the method for repairing a computing power operation task failure of the intelligent computing center as described in any one of the first aspects are implemented.
[0034] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the steps of the method for repairing a computing power operation task failure of the intelligent computing center as described in any one of the first aspects are implemented.
[0035] In the present invention, through step S1: in response to a running failure of a computing power resource operation task of the intelligent computing center called by a DiskANN object, determining a target file in which task data records interrupted by the running failure are recorded on the DiskANN object; step S2: in response to the running failure being repaired, starting from the target file, resuming the computing power resource operation task of the intelligent computing center called by the DiskANN object, avoiding the DiskANN object from restarting the entire task interrupted by the failure from the beginning, and accelerating the task recovery rate; and, the DiskANN object calls the abundant computing power resources of the intelligent computing center to run the task, accelerating the task recovery rate. The present invention can achieve a rapid recovery of the DiskANN object running the task after a failure repair, and improve the efficiency of the running task. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0037] Figure 1 is a schematic diagram of the overall process of data interaction based on DiskANN;
[0038] Figure 2 is a schematic diagram of the process of the method for repairing a computing power operation task failure of the intelligent computing center of the present invention;
[0039] Figure 3 is a schematic diagram of a directory structure;
[0040] Figure 4 is a schematic block diagram of the device for repairing a computing power operation task failure of the intelligent computing center of the present invention;
[0041] Figure 5 is a schematic block diagram of the electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0043] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "or" in the present invention means at least one of the connected objects. For example, "A or B" covers three scenarios, namely, Scenario 1: including A and not including B; Scenario 2: including B and not including A; Scenario 3: including both A and B. The character " / " generally indicates an "or" relationship between the associated objects before and after.
[0044] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0045] First, the technical terms related to the present invention will be briefly described below.
[0046] The "computing power" referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of a target result through processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.
[0047] The "computational power" (Computational Power, CP) referred to in the present invention means: the ability of a data center server to process data and achieve result output, a comprehensive index for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP_general + CP_intelligent + CP_super.
[0048] The "Network Power (NP)" described in the present invention refers to: It is an indication of the data transmission capacity of computing power facilities, including comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc., involving network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.
[0049] The "Storage Power (SP)" described in the present invention refers to: It is the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server-internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0050] The "computing power infrastructure" described in the present invention refers to: A new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize centralized computing, storage, transmission, and application of information.
[0051] The "new type of information infrastructure" described in the present invention refers to: mainly including network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0052] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.
[0053] The "general computing power" described in the present invention refers to: The computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0054] The "intelligent computing power" described in the present invention refers to: For various artificial intelligence innovation applications, a computing platform is deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, etc.
[0055] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0056] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0057] The "intelligent computing center" described in the present invention includes but is not limited to the intelligent computing center.
[0058] The "intelligent computing center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0059] The "computing power center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0060] The "supercomputing center" described in the present invention refers to: that is, the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0061] The "computing power resources" described in the present invention refers to: technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0062] The "DiskANN" described in the present invention is a vector retrieval engine based on distributed storage, capable of storing and retrieving vector data at the billion level on a single computer. Compared with traditional vector retrieval algorithms, DiskANN has higher storage efficiency and faster retrieval speed.
[0063] The "computing power operation task" described in the present invention refers to: specific workloads or jobs that are executed on computing power resources and require a certain amount of computing power support, usually involving scenarios such as complex data processing, numerical calculations, model training, or simulation.
[0064] In the present invention, the overall process of data interaction based on DiskANN can be referred to Figure 1 , and this overall process mainly includes the following tasks: creating a DiskANN object (abbreviated as Create), importing vector data into the DiskANN object (abbreviated as ImportData), importing the vector data into the DiskANN service node (abbreviated as Push Data), building a graph index for the imported vector data (abbreviated as Bulid), loading the graph index into memory (abbreviated as Load), querying vector data based on the graph index (abbreviated as Search), closing the DiskANN object (abbreviated as Close), and destroying the DiskANN object (abbreviated as Destroy), etc.
[0065] Among them, creating a DiskANN object means that the DiskANN service node creates a DiskANN object according to the client's Create request. When the DiskANN object is created, there is no data in it, which is in a data-free state and cannot provide any services to the outside. The client's Create request is sent to the DiskANN service node through the Index service node. After the DiskANN service node finishes creating the DiskANN object, it notifies the client through the Index service node.
[0066] In the task of importing vector data into the DiskANN object, the client sends an ImportData request containing vector data to the Index service node, and the Index service node temporarily stores the vector data in the ImportData request in the database of the Index service node. After the Index service node finishes importing the data, it notifies the client.
[0067] After the data import task, the client can send a build (Bulid) request to the Index service node. According to the build (Bulid) request, the Index service node executes the Push Data task, that is, the Index service node imports the vector data stored in the database of the Index service node into the DiskANN service node, and the DiskANN service node notifies the Index service node after the vector data is written to disk to form a file.
[0068] The Build task refers to the DiskANN service node processing the vector data of the imported DiskANN object based on the Build request sent by the Index service node, obtaining the graph index of the vector data, and saving the vector data and the graph index. After the DiskANN service node completes the Build task, it notifies the client through the Index service node.
[0069] The Load task refers to the client sending a load request to the DiskANN service node through the Index service node, and the DiskANN service node loading at least part of the graph index of the vector data of the DiskANN object into memory based on the load request. After the loading is completed, the DiskANN service node notifies the client through the Index service node.
[0070] The Search task refers to the client sending a vector data query request to the DiskANN service node through the Index service node, and the DiskANN service node retrieving the vector data most similar to the vector data to be queried based on the graph index, and returning the retrieved most similar vector data to the client through the Index service node.
[0071] The Close task refers to the client sending a close request for the DiskANN object to the DiskANN service node through the Index service node, and the DiskANN service node closing the DiskANN object based on the close request. After the closing is completed, the DiskANN service node notifies the client through the Index service node.
[0072] The Destroy task refers to the client sending a destroy request for the DiskANN object to the DiskANN service node through the Index service node, and the DiskANN service node destroying the DiskANN object based on the destroy request. After the destruction is completed, the DiskANN service node notifies the client through the Index service node.
[0073] The above process includes three executing entities: the client, the Index service node, and the DiskANN service node. Among them, the purpose of using the Index service node is to make the DiskANN service node as lightweight as possible. The DiskANN service node only processes the core DiskANN business, and outsources other businesses to the Index service node to ensure the stability and reliability of the overall service. Before executing the above process, the Index service node and the DiskANN service node need to perform a handshake connection in advance to provide services for subsequent operations. The above Index service node and DiskANN service node can be deployed on different physical machines or on the same physical machine, and the present invention does not limit this.
[0074] The present invention provides a method for repairing a computing power operation task failure in an intelligent computing center. Refer to Figure 2 as shown in Figure 2 which is a schematic flow diagram of the method for repairing a computing power operation task failure in the intelligent computing center of the present invention, and includes:
[0075] Step S1: In response to a running failure of a computing power resource operation task called by a DiskANN object, determine a target file on the DiskANN object where the task data recording is interrupted by the running failure;
[0076] Step S2: In response to the running failure being repaired, start from the target file and resume the computing power resource operation task called by the DiskANN object.
[0077] In practical applications, during the running of a task by a DiskANN object, the DiskANN object will record task data to form files. It should be noted that a large amount of task data will be generated during the running of a task, and multiple files will be formed.
[0078] In the present invention, the running failure causes the interruption of the running task, which also causes the interruption of the process of the DiskANN object recording task data. The file where the task data recording is interrupted is the target file on the DiskANN object where the task data recording is interrupted by the running failure in the present invention.
[0079] In step S1 of the present invention, determining the target file on the DiskANN object where the task data recording is interrupted by the running failure may specifically include: determining, from multiple files already formed during the running of the task by the DiskANN object, the file where the task data recording is interrupted by the running failure as the target file.
[0080] In the present invention, the DiskANN object calls the abundant computing power resources of the intelligent computing center to run tasks, which speeds up the task running rate and thus improves the task execution efficiency. Moreover, after the fault is repaired, based on the abundant computing power resources of the intelligent computing center, the task recovery rate can also be accelerated.
[0081] In the present invention, as shown in Figure 1 the tasks may include at least one of the following: creating a DiskANN object (abbreviated as Create), importing vector data into the DiskANN object (abbreviated as ImportData), importing vector data into the DiskANN service node (abbreviated as Push Data), building a graph index for the imported vector data (abbreviated as Bulid), loading the graph index into memory (abbreviated as Load), querying vector data based on the graph index (abbreviated as Search), closing the DiskANN object (abbreviated as Close), and destroying the DiskANN object (abbreviated as Destroy).
[0082] It should be noted that in combination with the introduction of each task in this article, the operation of the task requires the cooperation of the Index service node and the DiskANN service node to be completed. Therefore, the operation failure of the DiskANN object calling the computing power resources of the intelligent computing center to run tasks described in the present invention may actually be the failure of the DiskANN object running the task on the DiskANN service node, or the failure of the Index service node. Correspondingly, the task failure is repaired, which can be achieved by repairing the failures of the DiskANN service node and / or the Index service node to ensure that the task can be correctly run.
[0083] When restoring the DiskANN object to call the computing power resources of the intelligent computing center to run tasks, in combination with specific practices, if the task is interrupted and the task is running in the stage of independent operation of the DiskANN object, then restoring the DiskANN object to call the computing power resources of the intelligent computing center to run tasks can be: the DiskANN object calls the computing power resources of the intelligent computing center to continue running the task. For example: in a query task, the Index service node has sent a vector data query request to the DiskANN service node, and the task has been running in the stage of the DiskANN object on the DiskANN service node retrieving the vector data most similar to the vector data to be queried based on the graph index. In this case, if a running failure occurs and the task is interrupted, then after the task failure is repaired, it is only necessary to restore the DiskANN object to retrieve the vector data most similar to the vector data to be queried based on the graph index, without the intervention of the Index service node.
[0084] If a task is interrupted and the task is running in a stage where the DiskANN object does not run independently, the computing power resources of the intelligent computing center called by the DiskANN object can be restored to run the task. This can be achieved by instructing the Index service node to continue running the task to resume running the task on the DiskANN service node, thereby restoring the computing power resources of the intelligent computing center called by the DiskANN object to run the task. For example, during the execution of the Push Data task (the Push Data task means that the Index service node imports vector data in the database with the Index service node into the DiskANN service node), there is always a process in which the Index service node imports data into the DiskANN object on the DiskANN service node. In this case, if a running failure occurs and the task is interrupted, after the task failure is repaired, restoring the computing power resources of the intelligent computing center called by the DiskANN object to run the task can specifically include: determining the data that was interrupted for import based on the target file, and instructing the Index service node to start from the interrupted import data and continue to import the subsequent data into the DiskANN object, so as to restore the computing power resources of the intelligent computing center called by the DiskANN object to continue running the task.
[0085] In the present invention, through step S1: in response to a running failure of the DiskANN object calling the computing power resources of the intelligent computing center to run a task, determining a target file in which the task data record on the DiskANN object is interrupted by the running failure; step S2: in response to the running failure being repaired, starting from the target file, restoring the DiskANN object to call the computing power resources of the intelligent computing center to run the task, avoiding the DiskANN object from restarting the entire task interrupted by the failure from the beginning, and accelerating the task recovery rate; and, the DiskANN object calls the abundant computing power resources of the intelligent computing center to run the task, accelerating the task recovery rate. The present invention can achieve a rapid recovery of the DiskANN object running the task after a failure is repaired, improving the efficiency of running the task.
[0086] In some embodiments of the present invention, optionally, before step S1, it includes:
[0087] Step S01: determining the DiskANN object required for running the task;
[0088] Step S02: loading the DiskANN object on the DiskANN service node, and after the loading is completed, instructing the DiskANN object to call the computing power resources of the intelligent computing center to run the task.
[0089] In the present invention, before step S1, only the DiskANN objects required for running the task are loaded on the DiskANN service node, rather than loading the entire DiskANN service node completely (i.e., loading all DiskANN objects), realizing the on-demand loading of DiskANN objects, reducing the occupation of computing power resources (especially memory), and improving the utilization efficiency of computing power resources. It can be understood that avoiding loading the entire DiskANN service node completely (i.e., loading all DiskANN objects) also improves the efficiency of running the task.
[0090] In some embodiments of the present invention, optionally, a directory structure divided according to the task running stage is preset on the DiskANN object; during the process of running the task, the DiskANN object records the task data corresponding to the task running stage into the directory structure to form files.
[0091] Step S1 includes:
[0092] Step S11: Determine the first task running stage of the DiskANN object whose running fails and is interrupted.
[0093] Step S12: Query the directory structure according to the first task running stage to determine the target file.
[0094] In the present invention, a directory structure divided according to the task running stage is preset on the DiskANN object; during the process of running the task, the DiskANN object records the task data corresponding to the task running stage into the directory structure to form files, which facilitates the search for task data.
[0095] Based on this, by step S11: Determine the first task running stage of the DiskANN object whose running fails and is interrupted; step S12: Query the directory structure according to the first task running stage to determine the target file, the efficiency of determining the target file is improved, which is beneficial to quickly and accurately restoring the DiskANN object to call the computing power resources of the intelligent computing center to run the task after the fault is repaired.
[0096] In some embodiments of the present invention, optionally, step S2 includes:
[0097] Step S21: Determine whether the data in the target file is untrustworthy due to the running failure, and obtain a determination result.
[0098] Step S22: If the determination result indicates that the data in the target file is untrustworthy due to the running failure, re-run the task in the first task running stage, and record the task data during the re-run into the directory structure to form a first file, and replace the target file with the first file.
[0099] Step S23: If the determination result indicates that the running failure does not render the data in the target file untrustworthy, run the task in the next stage after the first task running stage.
[0100] Understandably, a task failure may cause distortion of task data. In the present invention, through Step S21: Determine whether the running failure renders the data in the target file untrustworthy to obtain a determination result; Step S22: If the determination result indicates that the running failure renders the data in the target file untrustworthy, re-run the task in the first task running stage, and correspondingly record the task data during the re-run into the directory structure to form a first file, and replace the target file with the first file; Step S23: If the determination result indicates that the running failure does not render the data in the target file untrustworthy, run the task in the next stage after the first task running stage. The present invention adds a determination of the credibility of the task data in the target file, avoids the situation of resuming the running task based on a target file with untrustworthy data, improves the quality of the running task, and reduces the probability of failures occurring in subsequent tasks.
[0101] In some specific embodiments, refer to Figure 3 as shown in Figure 3 is a schematic diagram of the directory structure. The directory structure includes: a directory (tmp directory) for recording task data in the building stage, a directory (normal directory) for recording task data in the completed building stage, a directory (destroyed directory) for recording task data in the deletion stage, and a directory (Blackhole directory) for recording the object names (IDs) without data. Each directory may include multiple files. For example, the destroyed directory includes files named "data.bin", "id.bin", "disk.index", "mem.index.data", etc. If the target file is a file in the tmp directory or the destroyed directory, the failure will cause distortion of the task data in the file, and the data in the target file is untrustworthy; if the target file is a file in the normal directory, the failure will not cause distortion of the task data in the file, and the data in the target file is trustworthy.
[0102] In some embodiments of the present invention, optionally, the directory structure includes at least one of the following directories:
[0103] A directory for recording task data in the building stage, a directory for recording task data in the completed building stage, a directory for recording task data in the deletion stage, and a directory for recording the object names without data.
[0104] In some specific embodiments, refer to Figure 3 as shown in Figure 3It is a schematic diagram of the directory structure, and the directory structure includes: a directory (tmp directory) for recording task data in the building stage, a directory (normal directory) for recording task data in the completed building stage, a directory (destroyed directory) for recording task data in the deletion stage, and a directory (Blackhole directory) for recording the object names (IDs) without data. Each directory can include multiple files. For example, the destroyed directory includes files named "data.bin", "id.bin", "_disk.index", "_mem.index.data", etc.
[0105] In the actual production process, the data is in the normal directory, and the client can use it by initiating a load rpc request. If it is in the tmp directory, only re-PushData and re-build are possible.
[0106] The present invention provides a device for repairing computing power operation task failures in an intelligent computing center. Refer to Figure 4 as shown in Figure 4 It is a principle block diagram of the device for repairing computing power operation task failures in the intelligent computing center of the present invention. The device 40 for repairing computing power operation task failures in the intelligent computing center includes:
[0107] A determination module 41, configured to determine a target file in which task data recorded by the DiskANN object is interrupted by the operation failure in response to an operation failure of the DiskANN object calling the computing power resource of the intelligent computing center to run a task;
[0108] An execution module 42, configured to resume the DiskANN object calling the computing power resource of the intelligent computing center to run a task starting from the target file in response to the repair of the operation failure.
[0109] In some embodiments of the present invention, optionally, it further includes:
[0110] A loading module, configured to determine a DiskANN object required for running a task;
[0111] The loading module is further configured to load the DiskANN object on the DiskANN service node. After the loading is completed, it instructs the DiskANN object to call the computing power resource of the intelligent computing center to run a task.
[0112] In some embodiments of the present invention, optionally, a directory structure divided according to the task running stage is preset on the DiskANN object; during the process of running a task, the DiskANN object records task data corresponding to the task running stage into the directory structure to form a file;
[0113] The determining module 41 is further configured to determine a first task running stage of the DiskANN object running task interrupted by the running failure;
[0114] The determining module 41 is further configured to query the directory structure according to the first task running stage to determine the target file.
[0115] In some embodiments of the present invention, optionally,
[0116] The executing module 42 is further configured to determine whether the running failure causes the data in the target file to be untrusted, and obtain a determination result;
[0117] The executing module 42 is further configured to, if the determination result indicates that the running failure causes the data in the target file to be untrusted, re-run the task in the first task running stage, and record the task data during the re-run into the directory structure to form a first file, and replace the target file with the first file;
[0118] The executing module 42 is further configured to, if the determination result indicates that the running failure does not cause the data in the target file to be untrusted, run the task in the next stage after the first task running stage.
[0119] In some embodiments of the present invention, optionally, the directory structure includes at least one of the following directories:
[0120] A directory for recording task data in the building stage, a directory for recording task data in the built stage, a directory for recording task data in the deletion stage, and a directory for recording object names without data.
[0121] The device for repairing the computing power running task failure of the intelligent computing center provided by the present invention can implement Figures 1 to 3 each process implemented by the method embodiments, and achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0122] The present invention provides an electronic device 50. Refer to Figure 5 as shown, Figure 5 is a schematic block diagram of the electronic device 50 of the present invention, including a processor 51, a memory 52, and a program or instruction stored in the memory 52 and executable on the processor 51. When the program or instruction is executed by the processor, the steps in any one of the methods for repairing the computing power running task failure of the intelligent computing center of the present invention are implemented.
[0123] The present invention provides a readable storage medium, on which a program or instructions are stored. When the program or instructions are executed by a processor, the various processes of the embodiment of the method for repairing the computing power operation task failure of the intelligent computing center as described in any one of the above are implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.
[0124] Among them, the readable storage medium is, for example, a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk or an optical disc, etc. In some examples, the readable storage medium may be a non-transitory readable storage medium.
[0125] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the various processes of the embodiment of the method for repairing the computing power operation task failure of the intelligent computing center as described in any one of the above are implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.
[0126] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment method can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0128] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention, and all of them belong to the protection scope of the present invention.
Claims
1. A method for repairing a computing power operation task failure in an intelligent computing center, characterized in that: include: Step S1: in response to a running failure of a DiskANN object calling a computing resource of an intelligent computing center to run a task, determining a target file on the DiskANN object where task data recorded by the running failure is interrupted; Step S2: In response to the operation failure being repaired, starting from the target file, the DiskANN object is restored to call the computing power resources of the intelligent computing center to run the task.
2. The method for repairing computing power operation task failures in an intelligent computing center according to claim 1, characterized in that: The step S1 previously comprises: Step S01: determine the DiskANN object required to run the task; Step S02: Load the DiskANN object on the DiskANN service node. After loading is complete, instruct the DiskANN object to call the computing resources of the intelligent computing center to run the task.
3. The method for repairing computing power operation task failures in an intelligent computing center according to claim 1, characterized in that: The DiskANN object is preset with a directory structure divided according to the task running stage; In the process of running the task, the DiskANN object records the task data into the directory structure to form a file according to the task running stage; The step S1 comprises: Step S11: determining the first task running phase of the DiskANN object running task interrupted by the running fault; Step S12: querying the directory structure according to the first task running phase to determine the target file.
4. The method for repairing a computing power operation task failure in an intelligent computing center according to claim 3 is characterized in that: The step S2 comprises: Step S21: determining whether the operation failure causes the data in the target file to be unreliable, and obtaining a determination result; Step S22: if the determination result indicates that the operation failure causes the data in the target file to be unreliable, re-run the task in the first task operation phase, and record the task data during the re-run into the directory structure to form a first file, and replace the target file with the first file; Step S23: If the determination result indicates that the operation failure does not cause the data in the target file to be unreliable, execute the task of the next stage after the first task execution stage.
5. The method for repairing a computing power operation task failure in an intelligent computing center according to claim 3 is characterized in that: The directory structure includes at least one of the following directories: The directory that records the task data in the construction phase, the directory that records the task data in the completed construction phase, the directory that records the task data in the deletion phase, and the directory that records the names of objects that have no data.
6. A computing power operation task fault repair device for an intelligent computing center, characterized in that: include: A determination module, configured to determine, in response to a running failure of a DiskANN object invoking a computing resource of an intelligent computing center to run a task, a target file on the DiskANN object where task data recorded by the running failure is interrupted; An execution module is used to restore the DiskANN object to call the computing power resources of the intelligent computing center to run the task starting from the target file in response to the running fault being repaired.
7. The computing power operation task fault repair device of the intelligent computing center according to claim 6 is characterized in that: Also includes: Loading module, which is used to determine the DiskANN objects required to run the task; The loading module is also used to load the DiskANN object on the DiskANN service node, and after loading is completed, instruct the DiskANN object to call the computing power resources of the intelligent computing center to run the task.
8. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps in the method for repairing the computing power operation task fault of the intelligent computing center as described in any one of claims 1 to 5 are implemented.
9. A readable storage medium, characterized in that: The readable storage medium stores programs or instructions, and when the programs or instructions are executed by the processor, the steps in the method for repairing computing power operation task faults of the intelligent computing center as described in any one of claims 1 to 5 are implemented.
10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the method for repairing computing power operation task faults of an intelligent computing center as described in any one of claims 1 to 5.