Computing power operation task fault defense method and device of intelligent computing center
By combining and querying the status of the region collection to be created in the intelligent computing center, ensuring that the data recovery task is run in parallel after the region is created, the system failure problem caused by the failure to create the region in multi-threaded operation is solved, and safe and efficient data recovery is achieved.
Patent Information
- Application Number
- CN202510557151.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-05
AI Technical Summary
In the intelligent computing center, multiple threads have a lack of verification perception of region creation before running data recovery tasks, which leads to running data recovery tasks when region is not created and triggers a risk of system failure.
Before multiple worker threads run data recovery tasks in parallel, combine all regions that are not currently created for running the data recovery task to form a collection of regions to be created, and query the region creation status at a preset time until all regions are created, and then run the data recovery task in parallel.
It avoids the risk of system failure caused by the failure to create the region, ensures that multiple worker threads run data recovery tasks safely and efficiently with sufficient computing power, and improves the operation efficiency of the tasks.
Smart Images

Figure CN120429083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent computing centers, smart computing centers and computing power infrastructure, and in particular to a method and device for preventing computing power operation task failures in intelligent computing centers. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged.
[0003] An "Intelligent Computing Center" is a facility that uses large-scale heterogeneous computing resources, including general-purpose and intelligent computing power, to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). The Intelligent Computing Center encompasses facilities, hardware, and software, and provides a full stack of capabilities, from bottom-level computing power to top-level application enablement.
[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.
[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.
[0007] Since the advent of intelligent computing centers, insufficient fault protection during computing task execution has been a challenge. Regions must be created before running data recovery tasks. Existing multi-threaded data recovery tasks lack verification of region creation before running them, making it easy for recovery tasks to run before region creation is complete, leading to system failure risks. Summary of the Invention
[0008] This invention provides a method and device for preventing computing task failures in intelligent computing centers. This approach addresses the issue of insufficient fault prevention during computing task execution, a problem that has plagued the emergence of intelligent computing centers. Existing multi-threaded data recovery tasks lack verification of region creation before executing them. This makes it easy for data recovery tasks to be executed before regions are fully created, leading to the risk of system failure.
[0009] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:
[0010] In a first aspect, the present invention provides a method for preventing computing power operation task failures in an intelligent computing center, comprising:
[0011] Step S1: Before multiple worker threads run the data recovery task in parallel, all currently uncreated regions required for running the data recovery task are combined to obtain a set of regions to be created;
[0012] Step S2: every preset waiting time, query whether the to-be-created region in the to-be-created region set has been created, and obtain the query result;
[0013] Step S3: If the query result indicates that there are regions to be created that have not been created, return to step S2 until the query result indicates that all regions to be created have been created;
[0014] Step S4: If the query result indicates that all the regions to be created have been created, control the multiple working threads to call the computing resources of the intelligent computing center to run the data recovery tasks in parallel.
[0015] Optionally, step S2 includes:
[0016] Step S21: before entering the step of querying whether the to-be-created region in the to-be-created region set has been created, verifying whether the number of currently completed queries exceeds a preset query number threshold, and obtaining a verification result;
[0017] Step S22: If the verification result indicates that the number of completed queries is less than or equal to the query number threshold, proceed to the step of querying whether the to-be-created region in the to-be-created region set has been created;
[0018] Step S23: If the verification result indicates that the number of completed queries is greater than the query number threshold, the step of querying whether the to-be-created region in the to-be-created region set has been created is not entered, an error message is generated, and the error message is sent to the interactive terminal associated with the user.
[0019] Optionally, after step S1 and before step S2, the following steps are included:
[0020] Step S01: determining whether the number of regions to be created in the region set to be created is 0, and obtaining a determination result;
[0021] Step S02: If the judgment result indicates that the number of regions to be created is 0, the process terminates and proceeds to step S2, controlling the multiple worker threads to call the computing resources of the intelligent computing center to run the data recovery task in parallel;
[0022] Step S03: If the judgment result indicates that the number of regions to be created is not 0, proceed to step S2.
[0023] Optionally, the multiple worker threads are created by the main thread;
[0024] During the parallel execution of the data recovery task, each of the worker threads executes at least one computing power execution subtask of the data recovery task;
[0025] After step S4, the method further includes:
[0026] Step S5: During the parallel execution of the data recovery task, the main thread is instructed to monitor the fault mark of each of the working threads; if a fault mark is detected in any of the working threads, the main thread is controlled to exit, and the working threads are controlled to exit the execution of the computing power operation subtask; if a fault mark is not detected in any of the working threads, the main thread is instructed to control the multiple working threads to exit the execution of the computing power operation subtask after all the computing power operation subtasks are completed.
[0027] Optionally, step S5 includes:
[0028] Step S51: determining whether the monitoring time of the main thread monitoring the fault mark of each worker thread exceeds a preset monitoring time threshold;
[0029] Step S52: If no fault mark is detected on the working thread after the monitoring time threshold is exceeded, and the computing power running subtask in the working thread has not been fully executed, control the main thread to exit, and control the working thread to exit executing the computing power running subtask.
[0030] In a second aspect, the present invention provides a device for preventing computing power operation task failures in an intelligent computing center, comprising:
[0031] A combining module is used to combine and run all currently uncreated regions required for the data recovery task before multiple worker threads run the data recovery task in parallel to obtain a set of regions to be created;
[0032] A query module, configured to query whether a region to be created in the set of regions to be created has been created every preset waiting time, and obtain a query result;
[0033] a first execution module configured to, if the query result indicates that any of the regions to be created have not been created, return the query module to cause the query module to query whether the regions to be created in the set of regions to be created have been created every preset waiting time, and obtain query results until the query result indicates that all of the regions to be created have been created;
[0034] The second execution module is used to control the multiple working threads to call the computing resources of the intelligent computing center to run the data recovery task in parallel if the query result indicates that all the regions to be created have been created.
[0035] Optionally, the query module is further configured to verify whether the number of currently completed queries exceeds a preset query number threshold before entering the step of querying whether the to-be-created region in the to-be-created region set has been created, and obtain a verification result;
[0036] The query module is further configured to, if the verification result indicates that the number of currently completed queries is less than or equal to the query number threshold, proceed to the step of querying whether the to-be-created region in the to-be-created region set has been created;
[0037] The query module is further configured to, if the verification result indicates that the number of currently completed queries is greater than the query number threshold, not proceed to the step of querying whether the to-be-created region in the to-be-created region set has been created, generate an error message, and send the error message to an interactive terminal associated with the user.
[0038] In a third aspect, the present invention provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps in the method for preventing computing power operation task failures in an intelligent computing center as described in any one of the first aspects are implemented.
[0039] In a fourth aspect, the present invention provides a readable storage medium storing a program or instruction, which, when executed by a processor, implements the steps of the method for preventing computing power operation task failures in an intelligent computing center as described in any one of the first aspects.
[0040] In a fifth aspect, the present invention provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps of the method for preventing computing power operation task failures in an intelligent computing center as described in any one of the first aspects.
[0041] In the present invention, through step S1: before multiple working threads run the data recovery task in parallel, all the currently uncreated regions required for running the data recovery task are combined to obtain a set of regions to be created; step S2: every preset waiting time, query whether the regions to be created in the set of regions to be created have been created, and obtain the query result; step S3: if the query result indicates that there are regions to be created that have not been created, return to step S2 until the query result indicates that all the regions to be created have been created; step S4: if the query result indicates that all the regions to be created have been created, control multiple working threads to call the computing power resources of the intelligent computing center to run the data recovery task in parallel. The present invention can avoid the system failure risk brought by running the data recovery task when there are regions to be created that have not been created, realize computing power running task failure defense, and ensure that multiple working threads run the data recovery task safely and efficiently with the sufficient computing power support of the intelligent computing center. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0043] Figure 1 A flowchart of a method for preventing computing power operation task failures in an intelligent computing center according to the present invention;
[0044] Figure 2 This is a schematic diagram of the region creation process;
[0045] Figure 3 This is a structural diagram of the thread pool;
[0046] Figure 4 This is a principle block diagram of the computing power operation task failure prevention device of the intelligent computing center of the present invention;
[0047] Figure 5 This is a principle block diagram of the electronic device of the present invention. DETAILED DESCRIPTION
[0048] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0049] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable under appropriate circumstances, so that the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "or" in the present invention represents at least one of the connected objects. For example, "A or B" covers three options, namely, option one: including A but not including B; option two: including B but not including A; option three: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.
[0050] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0051] First, the technical terms involved in the present invention are briefly explained below.
[0052] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0053] The "computing power" (CP) mentioned in the present invention refers to: the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, supercomputing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP=CP general + CP intelligent + CP super.
[0054] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capability of the computing power facility, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.
[0055] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in terms of data storage capacity, performance, security and reliability, and environmental friendliness. It is a comprehensive indicator for measuring a data center's data storage capacity, encompassing both external storage devices such as storage arrays and internal server storage. Storage capacity is commonly measured in exabytes (EB, 1EB = 2^60 bytes), while performance is commonly measured in IOPS / TB (Input / Output Operations Per Second / TB). Disaster recovery ratio is a key indicator of security and reliability.
[0056] The "computing power infrastructure" mentioned in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized calculation, storage, transmission and application of information.
[0057] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0058] The "computing power" mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.
[0059] The "general computing power" mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0060] The "intelligent computing power" mentioned in this invention refers to: a computing platform based on specialized chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various innovative artificial intelligence applications, such as natural language processing and machine vision.
[0061] The "supercomputing power" mentioned in the present invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, genetic analysis, etc.
[0062] The "intelligent computing center" described in this article refers to a facility that provides the computing power, data, and algorithms required for artificial intelligence applications (such as AI deep learning model development, model training, and model inference) by utilizing large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center encompasses facilities, hardware, and software, and can provide a full stack of capabilities, from bottom-level computing power to top-level application enablement.
[0063] The "intelligent computing center" mentioned in the present invention includes but is not limited to the "intelligent computing center".
[0064] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.
[0065] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0066] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide large-scale computing, storage and network services and other functions, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.
[0067] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.
[0068] The "computing power operation task" mentioned in the present invention refers to: a specific workload or job executed on computing power resources that requires a certain amount of computing power support, usually involving complex data processing, numerical calculations, model training or simulation scenarios.
[0069] The "Dingo Store" mentioned in the present invention refers to: an open source distributed key-value storage system that aims to provide high-performance, high-availability and scalable storage solutions, and is generally used in scenarios that require fast data access and high concurrency processing.
[0070] The "region" mentioned in the present invention refers to the smallest unit of Dingo Store scheduling and storage.
[0071] The present invention provides a method for preventing computing power operation task failures in an intelligent computing center. Figure 1 As shown, Figure 1 The flowchart of the method for preventing computing power operation task failures in the intelligent computing center of the present invention includes:
[0072] Step S1: Before multiple worker threads run the data recovery task in parallel, all currently uncreated regions required for running the data recovery task are combined to obtain a set of regions to be created;
[0073] Step S2: every preset waiting time, query whether the to-be-created region in the to-be-created region set has been created, and obtain the query result;
[0074] Step S3: If the query result indicates that there are regions to be created that have not been created, return to step S2 until the query result indicates that all regions to be created have been created;
[0075] Step S4: If the query result indicates that all regions to be created have been created, control multiple working threads to call the computing resources of the intelligent computing center to run data recovery tasks in parallel.
[0076] It should be noted that the data recovery task in the present invention can specifically be a full database recovery task in the Dingo Store system. Full database recovery tasks are used for empty databases and need to recover data as quickly as possible. Using multiple worker threads to run full database recovery tasks in parallel can significantly improve data recovery efficiency.
[0077] Specifically, the full database recovery task may include: restoring data of a specified region to the store / index / document node of Dingo Store.
[0078] The Dingo Store system includes dingodb_br nodes, coordinator nodes, store nodes, index nodes, and document nodes.
[0079] See also Figure 2 As shown, Figure 2The following is a diagram of the region creation process. After the coordinator node in Dingo Store receives the region creation instruction ("CreateRegionRequest") sent by the dingodb_br node, it returns the region creation completion result ("CreateRegionResponse") to the dingodb_br node. However, the region creation is not actually completed at this time. Afterwards, the coordinator node sends the real region creation instruction ("RealCreateRegionRequest") to the store node, index node, and document node respectively. After each node in the store node, index node, and document node completes the region creation, the coordinator node receives the real region creation completion result ("RealCreateRegionResponse") returned by the corresponding node. Until the coordinator node receives all the real region creation completion results ("RealCreateRegionResponse"), it means that all the required regions to be created have been created.
[0080] In the present invention, before multiple worker threads run the data recovery task in parallel, all currently uncreated regions required for running the data recovery task are combined to obtain a set of regions to be created, and in subsequent steps S2 to S4, the completion of creation of all to-be-created regions in the set of regions to be created is used as a condition for running the data recovery task in parallel. That is, the asynchronous thread that waits for the completion of creation of the to-be-created regions for each worker thread is changed to a synchronous thread that waits for the completion of creation of all to-be-created regions in the set of regions to be created before running the data recovery task in parallel, thereby avoiding the long time consumed by each worker thread waiting for the completion of creation of the to-be-created regions, thereby improving the efficiency of running the data recovery task.
[0081] In the present invention, before and during the execution of the present invention, see Figure 2 As shown, the thread responsible for creating the region continues to run. Step S2 of the present invention is to check the creation progress of the region to be created once every preset waiting time.
[0082] In the present invention, the preset waiting time can be specifically set by the user according to his actual needs, and the present invention does not limit this. In some optional embodiments, the preset waiting time can be any of the following times: 1 minute, 5 minutes, 1 second, 5 seconds.
[0083] In the present invention, step S3: if the query result indicates that there are regions to be created that have not been created, return to step S2 until the query result indicates that all regions to be created have been created; that is, as long as there is any region to be created that has not been created, continue waiting until all regions to be created are created.
[0084] In the present invention, step S4: if the query result indicates that all the regions to be created have been created, multiple working threads are controlled to call the computing power resources of the intelligent computing center to run the data recovery tasks in parallel, and the abundant computing power of the intelligent computing center is used to support multiple working threads to run the data recovery tasks in parallel, which can significantly improve the operating efficiency of the data recovery tasks.
[0085] In the present invention, through step S1: before multiple working threads run the data recovery task in parallel, all the currently uncreated regions required for running the data recovery task are combined to obtain a set of regions to be created; step S2: every preset waiting time, query whether the regions to be created in the set of regions to be created have been created, and obtain the query result; step S3: if the query result indicates that there are regions to be created that have not been created, return to step S2 until the query result indicates that all the regions to be created have been created; step S4: if the query result indicates that all the regions to be created have been created, control multiple working threads to call the computing power resources of the intelligent computing center to run the data recovery task in parallel. The present invention can avoid the system failure risk brought by running the data recovery task when there are regions to be created that have not been created, realize computing power running task failure defense, and ensure that multiple working threads run the data recovery task safely and efficiently with the sufficient computing power support of the intelligent computing center.
[0086] In some embodiments of the present invention, optionally, step S2 includes:
[0087] Step S21: before entering the step of querying whether the to-be-created region in the to-be-created region set has been created, verifying whether the number of currently completed queries exceeds a preset query number threshold, and obtaining a verification result;
[0088] Step S22: If the check result indicates that the number of completed queries is less than or equal to the query number threshold, proceed to the step of querying whether the to-be-created region in the to-be-created region set has been created;
[0089] Step S23: If the verification result indicates that the number of completed queries is greater than the query number threshold, the step of querying whether the to-be-created region in the to-be-created region set has been created is not entered, an error message is generated, and the error message is sent to the interactive terminal associated with the user.
[0090] In the present invention, the preset query number threshold can be specifically set by the user according to his or her actual needs, and the present invention does not impose any limitation on this.
[0091] In some embodiments of the present invention, optionally, a preset query number threshold represents the maximum waiting number. Correspondingly, step S22: if the check result indicates that the number of queries currently completed is less than or equal to the query number threshold, continue to enter the step of querying whether the region to be created in the region set to be created has been created, indicating that the maximum query number has not been reached, and step S2 can be continued. Correspondingly, step S23: if the check result indicates that the number of queries currently completed is greater than the query number threshold, do not enter the step of querying whether the region to be created in the region set to be created has been created, generate an error message, and send the error message to the interactive terminal associated with the user, indicating that there are still regions to be created that have not been created even though the maximum query number has been reached. Therefore, it is impossible to control multiple working threads to call the computing resources of the intelligent computing center to run the data recovery task in parallel. The present invention cannot jump to step S4, the task fails, and an error message is generated and sent to the interactive terminal associated with the user.
[0092] It should be noted that, in some optional embodiments, the error message may include information about regions that have not yet been created, so that the user can evaluate the running error and determine the next step.
[0093] In the present invention, through step S21: before entering the step of querying whether the region to be created in the set of regions to be created is completed, verify whether the number of currently completed queries exceeds a preset query number threshold to obtain a verification result; step S22: if the verification result indicates that the number of currently completed queries is less than or equal to the query number threshold, continue to enter the step of querying whether the region to be created in the set of regions to be created is completed; step S23: if the verification result indicates that the number of currently completed queries is greater than the query number threshold, do not enter the step of querying whether the region to be created in the set of regions to be created is completed, generate an error message, and send the error message to the interactive terminal associated with the user. The present invention sets an exit mechanism to avoid looping between step S3 and step S2 indefinitely, thereby avoiding multiple working threads waiting indefinitely to run data recovery tasks in parallel, which is beneficial to improving the efficiency of running data recovery tasks.
[0094] In some embodiments of the present invention, optionally, after step S1 and before step S2, the following steps are included:
[0095] Step S01: determining whether the number of regions to be created in the region set to be created is 0, and obtaining a determination result;
[0096] Step S02: If the judgment result indicates that the number of regions to be created is 0, the process terminates and proceeds to step S2, where multiple worker threads are controlled to call the computing resources of the intelligent computing center to run data recovery tasks in parallel;
[0097] Step S03: If the judgment result indicates that the number of regions to be created is not 0, proceed to step S2.
[0098] In the present invention, the judgment result of step S02 indicates that the number of regions to be created is 0, which means that all regions to be created have been created and no further query is required. The process ends and enters step S2, controlling multiple working threads to call the computing power resources of the intelligent computing center to run data recovery tasks in parallel.
[0099] In the present invention, the judgment result of step S03 indicates that the number of regions to be created is not 0, which means that there are regions to be created that have not been created yet, and further query is required, and the process goes to step S2.
[0100] In some embodiments of the present invention, optionally, multiple worker threads are created by the main thread;
[0101] During the parallel execution of the data recovery task, each worker thread executes at least one computing power subtask of the data recovery task;
[0102] After step S4, the method further includes:
[0103] Step S5: During the parallel execution of the data recovery task, the main thread is instructed to monitor the fault mark of each worker thread; if a fault mark is detected in any worker thread, the main thread is controlled to exit, and the worker thread is controlled to exit the execution of the computing power running subtask; if a fault mark is not detected in any worker thread, after waiting for the completion of all computing power running subtasks, the main thread is instructed to control multiple worker threads to exit the execution of the computing power running subtask.
[0104] See also Figure 3As shown, the main thread is responsible for starting and managing the execution of worker threads, including assigning tasks, scheduling the execution order of tasks, and monitoring the status of tasks. In multi-threaded parallel computing, the main thread also needs to coordinate data sharing and communication between different worker threads to ensure data consistency and correctness. The main thread is also responsible for capturing and processing faults or exceptions that occur in worker threads to ensure system stability. The worker thread executes the specific computing tasks assigned to it, including but not limited to data recovery and data backup. By having multiple worker threads execute tasks in parallel, computing efficiency can be significantly improved and task completion time can be shortened. During the execution of tasks, the worker thread can also capture and report faults or exceptions so that the main thread can perform corresponding processing. After completing the task, the worker thread usually returns the calculation results to the main thread or stores them in a shared data structure for subsequent processing or aggregation, thereby optimizing resource utilization and execution efficiency.
[0105] In the present invention, the fault mark helps the main thread to promptly discover abnormal or error conditions in the working thread, prevent the expansion of potential problems, ensure the stability of the system, and improve the reliability of the overall system. The main thread can decide whether to perform error handling or take other measures by checking the fault mark without having to deeply analyze the status of each thread, making the code clearer and easier to maintain, quickly responding to faults and providing feedback, and more effectively utilizing computing resources to avoid idle or wasted resources due to faults.
[0106] In some embodiments of the present invention, optionally, step S5 includes:
[0107] Step S51: Determine whether the monitoring time of the main thread monitoring the fault mark of each worker thread exceeds a preset monitoring time threshold;
[0108] Step S52: If no fault mark is detected on the worker thread after the monitoring time threshold is exceeded, and the computing power running subtask in the worker thread is not fully executed, the main thread is controlled to exit, and the worker thread is controlled to exit executing the computing power running subtask.
[0109] In the present invention, the monitoring duration threshold can be set by the user according to his / her actual needs, and the present invention does not impose any limitation on this.
[0110] In the present invention, by setting the listening time, it is possible to ensure that the main thread exits within a reasonable time, thereby releasing resources, avoiding the system from losing response due to long waiting times, and improving the overall availability of the system. After the timeout, the monitoring and alarm mechanism can also be triggered to promptly notify the operation and maintenance personnel to deal with potential problems, ensuring that the system can continue to process other tasks and avoiding the overall performance being affected by the delay of a certain thread.
[0111] The present invention provides a device for preventing computing power operation task failure in an intelligent computing center, see Figure 4 As shown, Figure 4 This is a principle block diagram of the computing power operation task failure prevention device of the intelligent computing center of the present invention. The computing power operation task failure prevention device 30 of the intelligent computing center includes:
[0112] A combining module 31 is configured to combine all currently uncreated regions required for running the data recovery task to obtain a set of regions to be created before multiple worker threads run the data recovery task in parallel;
[0113] A query module 32 is configured to query whether a region to be created in the set of regions to be created has been created every preset waiting time, and obtain a query result;
[0114] The first execution module 33 is configured to, if the query result indicates that any of the regions to be created have not been created, return the query module to cause the query module to query whether the regions to be created in the set of regions to be created have been created every preset waiting time, and obtain query results until the query result indicates that all of the regions to be created have been created;
[0115] The second execution module 34 is configured to control the multiple worker threads to call computing resources of the intelligent computing center to run the data recovery task in parallel if the query result indicates that all the regions to be created have been created.
[0116] In some embodiments of the present invention, optionally,
[0117] The query module 32 is further configured to verify whether the number of completed queries exceeds a preset query number threshold before entering the step of querying whether the to-be-created region in the to-be-created region set has been created, and obtain a verification result;
[0118] The query module 32 is further configured to, if the verification result indicates that the number of completed queries is less than or equal to the query number threshold, proceed to the step of querying whether the to-be-created region in the to-be-created region set has been created;
[0119] The query module 32 is further configured to, if the verification result indicates that the number of currently completed queries is greater than the query number threshold, not proceed to the step of querying whether the to-be-created region in the to-be-created region set has been created, generate an error message, and send the error message to the interaction terminal associated with the user.
[0120] In some embodiments of the present invention, optionally,
[0121] The query module 32 is further configured to determine whether the number of regions to be created in the set of regions to be created is 0, and obtain a determination result;
[0122] The query module 32 is further configured to, if the judgment result indicates that the number of regions to be created is 0, terminate the step of querying whether the regions to be created in the set of regions to be created have been created every preset waiting time, obtain the query result, and control the multiple worker threads to call the computing resources of the intelligent computing center to run the data recovery task in parallel;
[0123] The query module 32 is further configured to, if the judgment result indicates that the number of regions to be created is not 0, enter a step of querying whether the regions to be created in the set of regions to be created are created every preset waiting time to obtain a query result.
[0124] In some embodiments of the present invention, optionally, the multiple worker threads are created by a main thread;
[0125] During the parallel execution of the data recovery task, each of the worker threads executes at least one computing power execution subtask of the data recovery task;
[0126] The computing power operation task failure prevention device 30 includes:
[0127] The third execution module is used to instruct the main thread to monitor the fault mark of each of the working threads during the parallel execution of the data recovery task; if a fault mark is detected in any of the working threads, the main thread is controlled to exit, and the working threads are controlled to exit the execution of the computing power operation subtask; if a fault mark is not detected in any of the working threads, the main thread is controlled to control the multiple working threads to exit the execution of the computing power operation subtask after waiting for the completion of the execution of all the computing power operation subtasks.
[0128] In some embodiments of the present invention, optionally,
[0129] The third execution module is further configured to determine whether a monitoring time duration for which the main thread monitors the fault flag of each worker thread exceeds a preset monitoring time duration threshold;
[0130] The third execution module is also used to control the main thread to exit and control the worker thread to exit executing the computing power running subtask if no fault mark of the worker thread is detected after exceeding the monitoring time threshold, and the computing power running subtask in the worker thread is not fully executed.
[0131] The computing power operation task failure prevention device of the intelligent computing center provided by the present invention can achieve Figures 1 to 3 The various processes implemented by the method embodiment achieve the same technical effect and are not described here again to avoid repetition.
[0132] The present invention provides an electronic device 40, see Figure 5 As shown, Figure 5 This is a principle block diagram of the electronic device 40 of the present invention, including a processor 41, a memory 42, and a program or instruction stored in the memory 42 and executable on the processor 41. When the program or instruction is executed by the processor, the steps of any one of the methods for preventing computing power operation task failures in the intelligent computing center of the present invention are implemented.
[0133] The present invention provides a readable storage medium, which stores programs or instructions. When the programs or instructions are executed by a processor, the various processes of the embodiments of the computing power operation task failure prevention method of the intelligent computing center such as any of the above-mentioned items are implemented, and the same technical effects can be achieved. To avoid repetition, they will not be repeated here.
[0134] The readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. In some examples, the readable storage medium may be a non-transitory readable storage medium.
[0135] The present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the various processes of any of the above-mentioned embodiments of the method for preventing computing power operation task failures in an intelligent computing center, and can achieve the same technical effect. To avoid repetition, they will not be repeated here.
[0136] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0137] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0138] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A method for preventing computing power operation task failures in an intelligent computing center, characterized in that: include: Step S1: Before multiple worker threads run the data recovery task in parallel, all currently uncreated regions required for running the data recovery task are combined to obtain a set of regions to be created; Step S2: every preset waiting time, query whether the to-be-created region in the to-be-created region set has been created, and obtain the query result; Step S3: If the query result indicates that there are regions to be created that have not been created, return to step S2 until the query result indicates that all regions to be created have been created; Step S4: If the query result indicates that all the regions to be created have been created, control the multiple working threads to call the computing resources of the intelligent computing center to run the data recovery tasks in parallel.
2. The method for preventing computing power operation task failures in an intelligent computing center according to claim 1, characterized in that: The step S2 comprises: Step S21: before entering the step of querying whether the to-be-created region in the to-be-created region set has been created, verifying whether the number of currently completed queries exceeds a preset query number threshold, and obtaining a verification result; Step S22: If the verification result indicates that the number of completed queries is less than or equal to the query number threshold, proceed to the step of querying whether the to-be-created region in the to-be-created region set has been created; Step S23: If the verification result indicates that the number of completed queries is greater than the query number threshold, the step of querying whether the to-be-created region in the to-be-created region set has been created is not entered, an error message is generated, and the error message is sent to the interactive terminal associated with the user.
3. The method for preventing computing power operation task failures in an intelligent computing center according to claim 1, characterized in that: After step S1 and before step S2, the following steps are included: Step S01: determining whether the number of regions to be created in the region set to be created is 0, and obtaining a determination result; Step S02: If the judgment result indicates that the number of regions to be created is 0, the process terminates and proceeds to step S2, controlling the multiple worker threads to call the computing resources of the intelligent computing center to run the data recovery task in parallel; Step S03: If the judgment result indicates that the number of regions to be created is not 0, proceed to step S2.
4. The method for preventing computing power operation task failures in an intelligent computing center according to claim 1, characterized in that: The multiple worker threads are created by the main thread; During the parallel execution of the data recovery task, each of the worker threads executes at least one computing power execution subtask of the data recovery task; After step S4, the method further includes: Step S5: During the parallel execution of the data recovery task, the main thread is instructed to monitor the fault mark of each of the working threads; if a fault mark is detected in any of the working threads, the main thread is controlled to exit, and the working threads are controlled to exit the execution of the computing power operation subtask; if a fault mark is not detected in any of the working threads, the main thread is instructed to control the multiple working threads to exit the execution of the computing power operation subtask after all the computing power operation subtasks are completed.
5. The method for preventing computing power operation task failures in an intelligent computing center according to claim 4, characterized in that: The step S5 comprises: Step S51: determining whether the monitoring time of the main thread monitoring the fault mark of each worker thread exceeds a preset monitoring time threshold; Step S52: If no fault mark is detected on the working thread after the monitoring time threshold is exceeded, and the computing power running subtask in the working thread has not been fully executed, control the main thread to exit, and control the working thread to exit executing the computing power running subtask.
6. A device for preventing computing power operation task failures in an intelligent computing center, characterized in that: include: A combining module is used to combine and run all currently uncreated regions required for the data recovery task before multiple worker threads run the data recovery task in parallel to obtain a set of regions to be created; A query module, configured to query whether a region to be created in the set of regions to be created has been created every preset waiting time, and obtain a query result; a first execution module configured to, if the query result indicates that any of the regions to be created have not been created, return the query module to cause the query module to query whether the regions to be created in the set of regions to be created have been created every preset waiting time, and obtain query results until the query result indicates that all of the regions to be created have been created; The second execution module is used to control the multiple working threads to call the computing resources of the intelligent computing center to run the data recovery task in parallel if the query result indicates that all the regions to be created have been created.
7. The computing power operation task failure prevention device of the intelligent computing center according to claim 6 is characterized in that: The query module is further configured to verify whether the number of completed queries exceeds a preset query number threshold before entering the step of querying whether the to-be-created region in the to-be-created region set has been created, and obtain a verification result; The query module is further configured to, if the verification result indicates that the number of currently completed queries is less than or equal to the query number threshold, proceed to the step of querying whether the to-be-created region in the to-be-created region set has been created; The query module is further configured to, if the verification result indicates that the number of currently completed queries is greater than the query number threshold, not proceed to the step of querying whether the to-be-created region in the to-be-created region set has been created, generate an error message, and send the error message to an interactive terminal associated with the user.
8. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein when the program or instruction is executed by the processor, the steps of the method for preventing computing power operation task failures of an intelligent computing center as described in any one of claims 1 to 5 are implemented.
9. A readable storage medium, characterized in that: The readable storage medium stores programs or instructions, which, when executed by a processor, implement the steps of the method for preventing computing power operation task failures in an intelligent computing center as described in any one of claims 1 to 5.
10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the method for preventing computing power operation task failures of an intelligent computing center as described in any one of claims 1 to 5.