Multi-engine task retry method, device and equipment for disk array card and medium

By adopting a multi-engine task retry method based on microservice architecture in disk array cards, the problems of high system load, long retry cycle and high hardware cost in traditional methods are solved, and more efficient task retry and system load optimization are achieved.

CN119988100AActive Publication Date: 2025-05-13SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510198704.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-13
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

In complex business scenarios, the traditional disk array card multi-engine task retry method leads to increased system load and long task retry cycles, which easily causes host IO timeout and hardware usage costs to be too high.

Method used

The disk array card-multi-engine task retry method based on the microservice architecture is adopted. The first microservice component monitors the task execution of the hardware engine and obtains target events to determine task exception information. The second microservice component determines the retry task based on the exception information and writes it to the target queue. The third microservice component determines the task retry strategy based on the urgency, queue priority and usage, and realizes rapid retry of the abnormal task.

Benefits of technology

Reduces system load and task retry cycles, avoids retry of the entire IO data link, shortens development time and reduces hardware usage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988100A_ABST
    Figure CN119988100A_ABST
Patent Text Reader

Abstract

The invention discloses a disk array card multi-engine task retry method, device and equipment and a medium, relates to the technical field of computers, and is applied to a task retry system constructed based on micro-services, the system comprises a first micro-service component, a second micro-service component and a third micro-service component, the method comprises the steps that the task execution condition of hardware engines is monitored through a first micro-service component, when any hardware engine fails to execute a task, a target event defined by the hardware engine is obtained, and task exception information is obtained according to the target event; determining a retry task based on the task exception information through the second micro-service component, determining a corresponding target queue according to the task exception information, and writing the retry task into the target queue; and through the third micro-service component, determining a task retry strategy based on the task emergency degree, the queue priority and the queue use condition, and re-executing the retry task in the target queue. According to the invention, the system load and the task retry period are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a disk array card multi-engine task retry method, device, equipment and medium. Background Art

[0002] With the rapid development of information technologies such as artificial intelligence, the Internet of Things, and cloud computing, more and more data is running and stored on large servers. At the same time, the requirements for data storage speed, security, and reliability are also getting higher and higher. The continuous increase in storage demand has also promoted the rapid development and continuous improvement of RAID (Redundant Array of Independent Disks) technology.

[0003] The hardware RAID card receives a host IO ( , ), according to business needs, multiple engines of the RAID card "permutate and combine" to complete a certain IO task. The result of these engine "permutations and combinations" is usually called an "IO data chain". In complex business scenarios, it usually takes dozens of engines to work closely together to complete the disk placement task. Each engine processing link may have an exception, resulting in a certain IO sent by the host to the RAID card not being correctly placed in the disk under extreme pressure scenarios. Therefore, the IO sent by the host to the RAID card must be retried. The general traditional solution is to rebuild the "IO data chain" and retry the entire chain. The basic model is as follows: Figure 1 However, this will bring many negative effects. First, rebuilding the "IO data chain" itself is a very complicated process. It not only requires online analysis of IO anomalies, but also the reconstruction of the "IO data chain". After a series of operations, the system load is greatly increased. Second, the host IO timeout threshold is not infinite. Under complex business requirements, such a high-load "IO data chain" retry operation often involves dozens of engines, resulting in a long retry cycle and easy host IO timeout. Third, the traditional method is highly dependent on hardware customization and cannot flexibly adjust the strategy, resulting in excessively high hardware costs. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a disk array card multi-engine task retry method, device, equipment and medium, which can reduce the system load and task retry cycle, and its specific scheme is as follows:

[0005] In a first aspect, the present application discloses a disk array card multi-engine task retry method, which is applied to a task retry system built based on a microservice architecture, wherein the task retry system includes a first microservice component, a second microservice component, and a third microservice component, and the method includes:

[0006] The task execution status of each hardware engine in the disk array card is monitored by the first microservice component, and when any hardware engine fails to execute a task, a target event defined and fed back by the any hardware engine is obtained, so as to obtain corresponding task exception information according to the target event;

[0007] Determine a retry task based on the task exception information through the second microservice component, then determine a target queue corresponding to any one of the hardware engines from a plurality of preset queues according to the task exception information, and write the retry task into the target queue;

[0008] A task retry strategy is determined by the third microservice component based on task urgency, queue priority, and queue usage, and then the retry task in the target queue is re-executed according to the task retry strategy.

[0009] Optionally, the acquiring of a target event defined and fed back by any one of the hardware engines includes:

[0010] Obtain the target event defined and fed back by any of the hardware engines based on a target format; wherein the target format is obtained based on the hardware engine number, the reason for task execution failure, the task number, the task storage address, the task data size, the number of engine retries, and a target cyclic redundancy check value, and the target cyclic redundancy check value is a check value obtained by performing a cyclic redundancy check on the cumulative sum of the hardware engine number, the reason for task execution failure, the task number, the task storage address, the task data size, and the number of engine retries.

[0011] Optionally, before writing the retry task into the target queue, the method further includes:

[0012] The queue priority of each target queue corresponding to each hardware engine is determined according to the engine scheduling order, engine allocation weight and engine retry number of each hardware engine; wherein the engine scheduling order is the order in which each hardware engine executes corresponding tasks, and the engine retry number is the number of times each hardware engine has re-executed the corresponding task.

[0013] Optionally, determining the queue priority of each target queue corresponding to each hardware engine according to the engine scheduling order, engine allocation weight and engine retry count of each hardware engine includes:

[0014] Determine a first proportion of the engine scheduling order, a second proportion of the engine allocation weight, and a third proportion of the engine retry count for each of the hardware engines;

[0015] Determine a first calculation result based on the engine scheduling order and the first proportion, determine a second calculation result based on the engine allocation weight and the second proportion, and determine a third calculation result based on the engine retry count and the third proportion;

[0016] The queue priority of each of the target queues corresponding to each of the hardware engines is determined according to the first calculation result, the second calculation result, and the third calculation result.

[0017] Optionally, each of the target queues includes a first ring queue and a second ring queue, and the step of writing the retry task into the target queue includes:

[0018] Determine whether there is free space in the first ring queue in the target queue;

[0019] If there is free space in the first ring queue in the target queue, writing the retry task into the first ring queue;

[0020] If there is no free space in the first ring queue in the target queue, the retry task is written into the second ring queue.

[0021] Optionally, re-executing the retry task in the target queue according to the task retry strategy includes:

[0022] Determine a current target queue with the highest urgency from among the target queues;

[0023] Determine whether the retry task exists in the first ring queue in the current target queue;

[0024] If the retry task exists in the first ring queue in the current target queue, re-execute the retry task;

[0025] If the first ring queue in the current target queue does not have the retry task, determining a new current target queue from the remaining target queues according to the queue priority, and determining whether the first ring queue in the new current target queue has the retry task;

[0026] If the retry task exists in the first ring queue in the new current target queue, re-execute the retry task;

[0027] If the retry task does not exist in the first ring queue in the new current target queue, jump to the step of determining the current target queue with the highest urgency from among the target queues, so as to determine whether the retry task exists in the second ring queue in each target queue, and re-execute the retry task of the second ring queue.

[0028] Optionally, after re-executing the retry task in the target queue according to the task retry strategy, the method further includes:

[0029] If the re-execution process fails, determining whether the engine retry count of any of the hardware engines reaches a preset count threshold;

[0030] If the engine retry count of any of the hardware engines does not reach the preset count threshold, jump to the step of re-executing the retry task in the target queue;

[0031] If the engine retry times of any of the hardware engines reaches the preset times threshold, the any of the hardware engines is processed based on a preset exception handling mechanism.

[0032] In a second aspect, the present application discloses a disk array card multi-engine task retry device, which is applied to a task retry system built based on a microservice architecture, wherein the task retry system includes a first microservice component, a second microservice component and a third microservice component, including:

[0033] A monitoring module, used to monitor the task execution status of each hardware engine in the disk array card through the first microservice component, and when any hardware engine fails to execute a task, obtain a target event defined and fed back by any hardware engine, so as to obtain corresponding task exception information according to the target event;

[0034] A feedback module, configured to determine a retry task through the second microservice component and based on the task exception information, then determine a target queue corresponding to any one of the hardware engines from a plurality of preset queues according to the task exception information, and write the retry task into the target queue;

[0035] A retry module is used to determine a task retry strategy through the third microservice component and based on task urgency, queue priority, and queue usage, and then re-execute the retry task in the target queue according to the task retry strategy.

[0036] In a third aspect, the present application discloses an electronic device, comprising:

[0037] Memory, used to store computer programs;

[0038] The processor is used to execute the computer program to implement the aforementioned disk array card multi-engine task retry method disclosed above.

[0039] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned disk array card multi-engine task retry method is implemented.

[0040] It can be seen that the present application proposes a disk array card multi-engine task retry method, which is applied to a task retry system built based on a microservice architecture, wherein the task retry system includes a first microservice component, a second microservice component and a third microservice component, and the method includes: monitoring the task execution status of each hardware engine in the disk array card through the first microservice component, and when any hardware engine fails to execute a task, obtaining a target event defined and fed back by the any hardware engine, so as to obtain corresponding task exception information according to the target event; determining a retry task through the second microservice component and based on the task exception information, and then determining a target queue corresponding to any hardware engine from a plurality of preset queues according to the task exception information, and writing the retry task into the target queue; determining a task retry strategy through the third microservice component and based on task urgency, queue priority and queue usage, and then re-executing the retry task in the target queue according to the task retry strategy. From the above, it can be seen that the present application monitors the task execution status of each hardware engine through the first microservice component, and the hardware engine that fails to execute defines the target event, which is then fed back to the first microservice component so that the first microservice component can understand the corresponding task exception information, and determine the retry task and store it through the second microservice component, and then retry the retry task determined based on the task exception information through the third microservice component. In this way, retrying the entire target task is avoided. That is, compared with the traditional technology that retries the entire IO data chain, the present application retries the tasks in the hardware engine where the exception occurs based on the microservice architecture, which reduces the system load and the entire retry cycle, shortens the development time and thus reduces the hardware usage cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0042] Figure 1 A schematic diagram of a traditional task retry model;

[0043] Figure 2 A flow chart of a disk array card multi-engine task retry method disclosed in this application;

[0044] Figure 3 A schematic diagram of a hardware RAID card microservice layer disclosed in this application;

[0045] Figure 4 A schematic diagram of a multi-task abnormal monitoring and retry model for a hardware RAID card disclosed in this application;

[0046] Figure 5 A flowchart of a specific disk array card multi-engine task retry method disclosed in this application;

[0047] Figure 6 A schematic diagram of a DRQ storage rule disclosed in this application;

[0048] Figure 7 A schematic diagram of a task retry rule disclosed in this application;

[0049] Figure 8 This is a schematic diagram of the structure of a multi-engine task retry device for a disk array card disclosed in this application;

[0050] Fig. 9 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0051] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] In complex business scenarios, it usually takes dozens of engines to work closely together to complete the disk placement task. Each engine processing link may have anomalies, resulting in a certain IO sent by the host to the RAID card not being able to be correctly placed in the disk under extreme pressure scenarios. Therefore, the IO sent by the host to the RAID card must be retried. The general solution is to rebuild the "IO data chain" and retry the entire chain. However, this will bring many negative effects. First, rebuilding the "IO data chain" itself is a very complicated process. Not only does it require online analysis of IO anomalies, but it also requires the reconstruction of the "IO data chain", which greatly increases the system load; second, the entire retry cycle is long, which can easily cause host IO timeouts; third, traditional methods are highly dependent on hardware customization and cannot flexibly adjust strategies, resulting in excessively high hardware usage costs.

[0053] To this end, the embodiment of the present application proposes a disk array card multi-engine task retry solution, which can reduce system load and task retry cycle, and reduce hardware usage costs.

[0054] The present application embodiment discloses a disk array card multi-engine task retry method, which is applied to a task retry system built based on a microservice architecture. The task retry system includes a first microservice component, a second microservice component and a third microservice component. Figure 2 As shown, the method includes:

[0055] Step S11: monitor the task execution status of each hardware engine in the disk array card through the first microservice component, and when any hardware engine fails to execute a task, obtain the target event defined and fed back by the any hardware engine, so as to obtain corresponding task exception information according to the target event.

[0056] This application is based on the idea of ​​microservice architecture, adopts the design concept of microservice technology architecture of HWE (hardware event, events generated by hardware engine) exception monitoring, DRQ (Dual Ring Queue, two ring queues) dynamic storage, WRR (WRR With Urgent BasedOn DRQ, weighted polling scheduling priority algorithm based on DRQ storage mechanism) fast retry, and constructs three services of exception perception, conversion storage and retry decision. Furthermore, through the above three services, the abnormal tasks in the hardware engine with abnormal feedback are retried, thereby avoiding retrying the entire IO data chain.

[0057] Figure 3 This is a schematic diagram of a hardware RAID card microservice layer disclosed in this application. Figure 4 This is a schematic diagram of a hardware RAID card multi-task abnormal monitoring and senseless retry model disclosed in this application. This application divides the entire retry model into three levels from top to bottom, namely the host, the hardware RAID acceleration card function group and the hardware physical disk entity. Specifically, the host is used to provide IO data management services for managing the host side, for example: the user needs to write data to the disk attached to the hardware RAID card, or needs to read the specified data of the disk attached to the RAID card; the RAID acceleration card function group includes engines and modules such as event monitoring, dynamic storage, senseless retry, hardware engine group, and abnormal management engine, among which the hardware engine group includes various hardware engines, such as memory application engine, RAID acceleration engine, stripe mapping engine, IO processing engine, data encryption engine, abnormal management engine, etc.; the hardware physical disk entity is used to access the IO data sent by the host.

[0058] Figure 5The present invention discloses a specific disk array card multi-engine task retry method flow chart. First, the host sends IO data, and the task distributor sends a specific IO task to the hardware engine group based on the IO data, and monitors the task execution status of each hardware engine in the hardware engine group through the first microservice component. If the hardware engine task is successfully executed, the task distributor sends a specific IO task to the done entry (i.e. Figure 4 The write entry in the microservice component represents the feedback interface for completing the task) writes a completion signal. If any hardware engine fails to execute the task, the hardware engine that failed to execute fills in the target event in the target format, so as to obtain the corresponding task exception information and retry task according to the target event, and stores the corresponding retry task based on the target queue through the second microservice component, and then retries the corresponding retry task based on the task retry strategy through the third microservice component.

[0059] The hardware engine mainly includes the task acceptance unit, the result report unit, the task processing unit and the hardware channel resources, so that a single disk engine can independently complete a specific task. Specifically, the task acceptance unit is used to receive tasks to be executed; the result report unit is used to select whether to report the target event. If the task execution fails, the target event is filled in according to the target format and the completion status is reported; the task processing unit is used to process the specified task.

[0060] The target format is based on the hardware event ID (Identity document, ID card identification number), hardware engine number ( )、Task execution failure reason( )、Task Number( )、task storage address( )、task data size( ), engine retry times (Retrytimes) and a target cyclic redundancy check value (Parity bits), wherein the target cyclic redundancy check value is a check value obtained by performing a cyclic redundancy check on the cumulative sum of the hardware event ID, the hardware engine number, the reason for the task execution failure, the task number, the task storage address, the task data size and the engine retry times.

[0061] In this embodiment, a target event defined and fed back by any of the hardware engines is obtained so as to obtain corresponding task exception information according to the target event. It can be understood that the task exception information includes the hardware engine number, task storage address, task data size, engine retry count, etc.

[0062] Step S12: determining a retry task through the second microservice component and based on the task exception information, then determining a target queue corresponding to any one of the hardware engines from a plurality of preset queues according to the task exception information, and writing the retry task into the target queue.

[0063] In this embodiment, before writing the retry task into the target queue, it is necessary to determine the queue priority of each target queue corresponding to each hardware engine according to the engine scheduling order, engine allocation weight and engine retry number of each hardware engine; wherein, the engine scheduling order is the order in which each hardware engine executes the corresponding task, and the engine retry number is the number of times each hardware engine has re-executed the corresponding task. Specifically, determine the first proportion of the engine scheduling order, the second proportion of the engine allocation weight and the third proportion of the engine retry number of each hardware engine; determine the first operation result based on the engine scheduling order and the first proportion, determine the second operation result based on the engine allocation weight and the second proportion, and determine the third operation result based on the engine retry number and the third proportion; determine the queue priority of each target queue corresponding to each hardware engine through the first operation result, the second operation result and the third operation result.

[0064] Exemplary, defined The emergency queue has the highest priority and is used to handle the most urgent retry tasks. For example, the exception handling engine needs to handle the RAID card system-level exceptions, and the generated exception tasks must be responded to urgently. arrive The priority is lower in sequence. It should be noted that the higher the engine scheduling order, the higher its priority is expected to be. The higher the assigned weight, the higher its priority is expected to be. The more the engine has accumulated retries, the higher its priority is expected to be. In this embodiment, each proportion is configured according to actual business needs and is not a mandatory constraint. , that is, assuming arrive Stored in arrive , and the scheduling priority of the retry engine is consistent with the DRQ storage priority.

[0065] In this embodiment, after determining the priority of each queue, the retry task is determined through the second microservice component and based on the task exception information, and then the target queue corresponding to any hardware engine is determined from several preset queues according to the task exception information, and the retry task is written into the target queue.

[0066] Step S13: determining a task retry strategy through the third microservice component and based on task urgency, queue priority, and queue usage, and then re-executing the retry task in the target queue according to the task retry strategy.

[0067] In this embodiment, a DRQ Ping-Pong (which can use two task buffers at the same time to achieve the effect of continuous task reception and uninterrupted processing) storage mechanism is adopted. When the Q0 queue is used to receive or process retry IO, the Q1 queue is used to prepare the next set of data. During a round of retry in the Q0 queue, a new abnormal IO task is injected into Q1. In this way, even if multi-threaded scheduling is performed, there is no need to lock and wait. The DRQ storage rules are as follows: Figure 6 As shown, in this embodiment, Q0 is called the first ring queue and Q1 is called the second ring queue.

[0068] Based on the above-mentioned circular queue, when writing a retry task into the queue, this embodiment specifically determines whether there is free space in the first circular queue in the target queue; if there is free space in the first circular queue in the target queue, the retry task is written into the first circular queue; if there is no free space in the first circular queue in the target queue, the retry task is written into the second circular queue.

[0069] Based on the above-mentioned task writing method, in this embodiment, when reprocessing the retry task, the current target queue with the highest urgency is determined from each of the target queues; it is determined whether the retry task exists in the first ring queue in the current target queue; if the retry task exists in the first ring queue in the current target queue, the retry task is re-executed; if the retry task does not exist in the first ring queue in the current target queue, a new current target queue is determined from the remaining target queues according to the queue priority, and it is determined whether the retry task exists in the first ring queue in the new current target queue; if the retry task exists in the first ring queue in the new current target queue, the retry task is re-executed; if the retry task does not exist in the first ring queue in the new current target queue, the step of determining the current target queue with the highest urgency from each of the target queues is jumped to, so as to determine whether the retry task exists in the second ring queue in each of the target queues, and re-execute the retry task of the second ring queue.

[0070] See also Figure 7 As shown, when retrying a retry task, the urgency, priority, The three-dimensional perspective of the mechanism determines the task retry strategy. Specifically, it includes the following three steps: a) Scheduling the most urgent Each DRQ has two internal circular queues, Q0 and Q1. The scheduling starts from Q0. When the Q0 queue has a storage IO task, the retry task of its Q0 queue is completed. When the queue is empty, this round of scheduling retry is skipped. b) Scheduling the highest priority , start scheduling execution from Q0. When the Q0 queue has a storage IO task, complete the retry task of its Q0 queue. When the queue is empty, skip this round of scheduling and retry; c) After a round of Q0 scheduling is completed, execute another round of Q1 queue scheduling, and repeat this cycle to periodically retry abnormal IO tasks. In this embodiment, the queue with a smaller priority number has a higher urgency and will be retried first. In addition, this embodiment can also add a fourth dimension on the basis of the above three-dimensional angles, that is, the determination time of the retry task. The time obtained by different retry tasks and the above three dimensions jointly determine the retry order of each retry task. In this way, the accuracy of the task retry strategy is improved.

[0071] Furthermore, after re-executing the retry task in the target queue according to the task retry strategy, it also includes: if the re-execution fails, determining whether the engine retry times of any of the hardware engines have reached a preset times threshold; if the engine retry times of any of the hardware engines have not reached the preset times threshold, jumping to the step of re-executing the retry task in the target queue; if the engine retry times of any of the hardware engines have reached the preset times threshold, processing any of the hardware engines based on a preset exception handling mechanism. That is, this embodiment sets a retry times threshold, and if the retry times are less than the retry times threshold and the retry fails, it can be retried again, and if the retry times are greater than the retry times threshold and the retry fails, no retry is performed, but processing is performed based on a preset exception handling mechanism.

[0072] In summary, this application proposes the idea of ​​adopting a microservice architecture, using an event monitoring mechanism to perceive the execution status of hardware engine tasks, converting hardware events into engine executable tasks, and completing dynamic storage according to DRQ definition rules, and then achieving fast retry through the WRR mechanism, thereby achieving real-time interaction with the hardware engine, fast closed loop, and efficient retry, improving hardware stability, optimizing hardware system load, reducing hardware usage costs, shortening function development cycle, and accelerating product launch.

[0073] It can be seen that the present application proposes a disk array card multi-engine task retry method, which is applied to a task retry system built based on a microservice architecture, wherein the task retry system includes a first microservice component, a second microservice component and a third microservice component, and the method includes: monitoring the task execution status of each hardware engine in the disk array card through the first microservice component, and when any hardware engine fails to execute a task, obtaining a target event defined and fed back by the any hardware engine, so as to obtain corresponding task exception information according to the target event; determining a retry task through the second microservice component and based on the task exception information, and then determining a target queue corresponding to any hardware engine from a plurality of preset queues according to the task exception information, and writing the retry task into the target queue; determining a task retry strategy through the third microservice component and based on task urgency, queue priority and queue usage, and then re-executing the retry task in the target queue according to the task retry strategy. From the above, it can be seen that the present application monitors the task execution status of each hardware engine through the first microservice component, and the hardware engine that fails to execute defines the target event, which is then fed back to the first microservice component so that the first microservice component can understand the corresponding task exception information, and determine the retry task and store it through the second microservice component, and then retry the retry task determined based on the task exception information through the third microservice component. In this way, retrying the entire target task is avoided. That is, compared with the traditional technology that retries the entire IO data chain, the present application retries the tasks in the hardware engine where the exception occurs based on the microservice architecture, which reduces the system load and the entire retry cycle, shortens the development time and thus reduces the hardware usage cost.

[0074] Correspondingly, the embodiment of the present application also discloses a disk array card multi-engine task retry device, which is applied to a task retry system built based on a microservice architecture. The task retry system includes a first microservice component, a second microservice component and a third microservice component. Figure 8 As shown, the device comprises:

[0075] The monitoring module 11 is used to monitor the task execution status of each hardware engine in the disk array card through the first microservice component, and when any hardware engine fails to execute a task, obtain a target event defined and fed back by any hardware engine, so as to obtain corresponding task exception information according to the target event;

[0076] A feedback module 12 is used to determine a retry task based on the task exception information through the second microservice component, and then determine a target queue corresponding to any one of the hardware engines from a plurality of preset queues according to the task exception information, and write the retry task into the target queue;

[0077] The retry module 13 is used to determine the task retry strategy through the third microservice component and based on the task urgency, queue priority and queue usage, and then re-execute the retry task in the target queue according to the task retry strategy.

[0078] Among them, for more specific working processes of the above-mentioned modules, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.

[0079] It can be seen that the present application proposes a disk array card multi-engine task retry method, which is applied to a task retry system built based on a microservice architecture, wherein the task retry system includes a first microservice component, a second microservice component and a third microservice component, and the method includes: monitoring the task execution status of each hardware engine in the disk array card through the first microservice component, and when any hardware engine fails to execute a task, obtaining a target event defined and fed back by the any hardware engine, so as to obtain corresponding task exception information according to the target event; determining a retry task through the second microservice component and based on the task exception information, and then determining a target queue corresponding to any hardware engine from a plurality of preset queues according to the task exception information, and writing the retry task into the target queue; determining a task retry strategy through the third microservice component and based on task urgency, queue priority and queue usage, and then re-executing the retry task in the target queue according to the task retry strategy. From the above, it can be seen that the present application monitors the task execution status of each hardware engine through the first microservice component, and the hardware engine that fails to execute defines the target event, which is then fed back to the first microservice component so that the first microservice component can understand the corresponding task exception information, and determine the retry task and store it through the second microservice component, and then retry the retry task determined based on the task exception information through the third microservice component. In this way, retrying the entire target task is avoided. That is, compared with the traditional technology that retries the entire IO data chain, the present application retries the tasks in the hardware engine where the exception occurs based on the microservice architecture, which reduces the system load and the entire retry cycle, shortens the development time and thus reduces the hardware usage cost.

[0080] Furthermore, an embodiment of the present application also provides an electronic device. Fig. 9 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.

[0081] Fig. 9The present invention provides a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the disk array card multi-engine task retry method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0082] In this embodiment, the power supply 26 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 25 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 24 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0083] In addition, the memory 22, as a carrier for storing resources, may be a read-only memory, a random access memory, a disk or an optical disk, etc., and the resources stored thereon may include a computer program 221, and the storage method may be a temporary storage or a permanent storage. Among them, the computer program 221 includes not only a computer program that can be used to complete the disk array card multi-engine task retry method executed by the electronic device 20 disclosed in any of the aforementioned embodiments, but also a computer program that can be used to complete other specific tasks.

[0084] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned disk array card multi-engine task retry method disclosed above is implemented.

[0085] For the specific steps of the method, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be described in detail here.

[0086] Furthermore, an embodiment of the present application also discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the disk array card multi-engine task retry method disclosed above.

[0087] The various embodiments in this application are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.

[0088] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0089] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0090] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0091] The above is a detailed introduction to a disk array card multi-engine task retry method, device, equipment, and storage medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A disk array card multi-engine task retry method, characterized in that: Applied to a task retry system built based on a microservice architecture, the task retry system includes a first microservice component, a second microservice component and a third microservice component, and the method includes: The task execution status of each hardware engine in the disk array card is monitored by the first microservice component, and when any hardware engine fails to execute a task, a target event defined and fed back by the any hardware engine is obtained, so as to obtain corresponding task exception information according to the target event; Determine a retry task based on the task exception information through the second microservice component, then determine a target queue corresponding to any one of the hardware engines from a plurality of preset queues according to the task exception information, and write the retry task into the target queue; A task retry strategy is determined by the third microservice component based on task urgency, queue priority, and queue usage, and then the retry task in the target queue is re-executed according to the task retry strategy.

2. The disk array card multi-engine task retry method according to claim 1, characterized in that: The obtaining of the target event defined and fed back by any one of the hardware engines includes: Obtain the target event defined and fed back by any of the hardware engines based on a target format; wherein the target format is obtained based on the hardware engine number, the reason for task execution failure, the task number, the task storage address, the task data size, the number of engine retries, and a target cyclic redundancy check value, and the target cyclic redundancy check value is a check value obtained by performing a cyclic redundancy check on the cumulative sum of the hardware engine number, the reason for task execution failure, the task number, the task storage address, the task data size, and the number of engine retries.

3. The disk array card multi-engine task retry method according to claim 2, characterized in that: Before writing the retry task into the target queue, the method further includes: The queue priority of each target queue corresponding to each hardware engine is determined according to the engine scheduling order, engine allocation weight and engine retry number of each hardware engine; wherein the engine scheduling order is the order in which each hardware engine executes corresponding tasks, and the engine retry number is the number of times each hardware engine has re-executed the corresponding task.

4. The disk array card multi-engine task retry method according to claim 3, characterized in that: The determining the queue priority of each target queue corresponding to each hardware engine according to the engine scheduling order, engine allocation weight and engine retry count of each hardware engine includes: Determine a first proportion of the engine scheduling order, a second proportion of the engine allocation weight, and a third proportion of the engine retry count for each of the hardware engines; Determine a first calculation result based on the engine scheduling order and the first proportion, determine a second calculation result based on the engine allocation weight and the second proportion, and determine a third calculation result based on the engine retry count and the third proportion; The queue priority of each of the target queues corresponding to each of the hardware engines is determined according to the first calculation result, the second calculation result, and the third calculation result.

5. The disk array card multi-engine task retry method according to any one of claims 1 to 4, characterized in that: Each of the target queues includes a first ring queue and a second ring queue, and writing the retry task into the target queue includes: Determine whether there is free space in the first ring queue in the target queue; If there is free space in the first ring queue in the target queue, writing the retry task into the first ring queue; If there is no free space in the first ring queue in the target queue, the retry task is written into the second ring queue.

6. The disk array card multi-engine task retry method according to claim 5, characterized in that: The re-executing the retry task in the target queue according to the task retry strategy includes: Determine a current target queue with the highest urgency from among the target queues; Determine whether the retry task exists in the first ring queue in the current target queue; If the retry task exists in the first ring queue in the current target queue, re-execute the retry task; If the first ring queue in the current target queue does not have the retry task, determining a new current target queue from the remaining target queues according to the queue priority, and determining whether the first ring queue in the new current target queue has the retry task; If the retry task exists in the first ring queue in the new current target queue, re-execute the retry task; If the retry task does not exist in the first ring queue in the new current target queue, jump to the step of determining the current target queue with the highest urgency from among the target queues, so as to determine whether the retry task exists in the second ring queue in each target queue, and re-execute the retry task of the second ring queue.

7. The disk array card multi-engine task retry method according to claim 6, characterized in that: After re-executing the retry task in the target queue according to the task retry strategy, the method further includes: If the re-execution process fails, determining whether the engine retry count of any of the hardware engines reaches a preset count threshold; If the engine retry count of any of the hardware engines does not reach the preset count threshold, jump to the step of re-executing the retry task in the target queue; If the engine retry times of any of the hardware engines reaches the preset times threshold, the any of the hardware engines is processed based on a preset exception handling mechanism.

8. A disk array card multi-engine task retry device, characterized in that: The task retry system is applied to a task retry system built based on a microservice architecture, the task retry system includes a first microservice component, a second microservice component and a third microservice component, including: A monitoring module, used to monitor the task execution status of each hardware engine in the disk array card through the first microservice component, and when any hardware engine fails to execute a task, obtain a target event defined and fed back by any hardware engine, so as to obtain corresponding task exception information according to the target event; A feedback module, configured to determine a retry task through the second microservice component and based on the task exception information, then determine a target queue corresponding to any one of the hardware engines from a plurality of preset queues according to the task exception information, and write the retry task into the target queue; A retry module is used to determine a task retry strategy through the third microservice component and based on task urgency, queue priority, and queue usage, and then re-execute the retry task in the target queue according to the task retry strategy.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is used to execute the computer program to implement the disk array card multi-engine task retry method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein, when the computer program is executed by a processor, the disk array card multi-engine task retry method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Soft and hard engine switching method for NVMe SSD disk management, terminal and storage medium

    CN115658385A

  • Server-free workflow function exception retry method and device

    CN115686933A

  • Processing an operation with a plurality of processing steps

    US10362097B1