Processor Synchronization Equalization Method, System, Device, Equipment, Medium and Product

By using the credit value management of the submission queue and completion queue in a multi-core CPU environment, synchronous balance between processors is achieved, the IO command loss caused by task asymmetry is solved, and the service quality of IO commands is improved.

CN119690682BActive Publication Date: 2025-07-25INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510196112.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-25
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

In a multi-core CPU environment, the asymmetry of tasks assigned on each core leads to the loss of step and imbalance of IO commands, affecting the quality of service (QoS) of IO commands.

Method used

By distributing the commands in the submission queue to the processor, the initial credit value is determined based on the completion queue and the number of processor cores, and synchronous equalization operations are performed when the update credit value reaches the preset value, adjust the credit value and processor computing power, and transfer the command processing between the processors to achieve balance.

Benefits of technology

The command processing balance between different processors is achieved, the quality of the processor is ensured, the jitter of IO commands is reduced, and the stability and consistency of the server is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119690682B_ABST
    Figure CN119690682B_ABST
Patent Text Reader

Abstract

The present disclosure provides a processor synchronization and balancing method, system, device, equipment, medium and product. Among them, the method includes: distributing a first command in a first submission queue to a first processor; determining an initial credit value of the first processor based on a first completion queue corresponding to the first submission queue and the number of cores of the first processor; decreasing and updating the initial credit value based on the number of first commands processed by the first processor to obtain an updated credit value; when the updated credit value reaches a first preset value, performing a synchronization and balancing operation on the updated credit value of the first processor. The method of the present disclosure can adaptively adjust the credit value during processing and the ceding of processor computing power by performing a synchronization and balancing operation on the updated credit value of the first processor when the updated credit value reaches the first preset value, so as to balance the command processing among different processors and ensure the service quality of the processors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computers, and in particular, to a processor synchronization and balancing method, system, device, equipment, medium and product. Background Art

[0002] With the rapid development of information technology, the central processing unit (CPU) has long entered the multi-core era. Each core can process the same or different data processing tasks and coordinate with each other through an interconnected information network to achieve fast processing of data.

[0003] However, due to the incomplete symmetry of the tasks allocated and running on each core, in the actual operation process, there will be an out-of-step and imbalance of input / output (IO) commands on different CPUs of the solid-state drive, which in turn causes a large jitter in the IO commands on the server side, affecting the quality of service (QoS) of the IO commands. Summary of the Invention

[0004] The present disclosure provides a processor synchronization and balancing method, system, device, equipment, medium and product, which can adaptively adjust the credit value in the processing and the ceding of the processor computing power, so as to balance the command processing between different processors and ensure the quality of service of the processors.

[0005] The first aspect embodiment of the present disclosure proposes a processor synchronization and balancing method, including: distributing a first command in a first submission queue to a first processor; determining an initial credit value of the first processor based on a first completion queue corresponding to the first submission queue and the number of cores of the first processor; decreasing and updating the initial credit value based on the number of first commands completed by the first processor to obtain an updated credit value; when the updated credit value reaches a first preset value, performing a synchronization and balancing operation on the updated credit value of the first processor.

[0006] In some embodiments of the present disclosure, distributing the first command in the first submission queue to the first processor includes: based on a polling algorithm, evenly distributing the first command to the first processor in sequence.

[0007] In some embodiments of the present disclosure, determining the initial credit value of the first processor based on the first completion queue corresponding to the first submission queue and the number of cores of the first processor includes: determining the initial credit value based on the ratio between the first queue depth of the first completion queue and the number of cores of the first processor.

[0008] In some embodiments of the present disclosure, when the update credit value reaches a first preset value, synchronously balancing the update credit value of the first processor includes: determining the remaining credit value of the second completion queue based on the head pointer value of the second completion queue, the tail pointer value of the second completion queue, and the second queue depth of the second completion queue, where the second completion queue is the queue in the first completion queue whose update credit value reaches the first preset value; determining the total number of pending processes of the first processor based on the expected transmission count of the first processor and the completed transmission count of the first processor, where the expected transmission count is the number of commands that the first processor expects to send to the second completion queue, and the completed transmission count is the number of commands that the first processor has already sent to the second completion queue; determining whether there is available credit value in the second completion queue based on the remaining credit value and the total number of pending processes; and performing a synchronous balancing operation on the update credit value of the first processor based on whether there is available credit value in the second completion queue.

[0009] In some embodiments of the present disclosure, determining the remaining credit value of the second completion queue based on the head pointer value of the second completion queue, the tail pointer value of the second completion queue, and the second queue depth of the second completion queue includes: when the head pointer value is greater than the tail pointer value, determining the remaining credit value by using the second queue depth; when the head pointer value is less than the tail pointer value, determining the remaining credit value based on the head pointer value and the tail pointer value; and when the head pointer value is equal to the tail pointer value, determining the remaining credit value based on the head pointer value, the tail pointer value, and the second queue depth.

[0010] In some embodiments of the present disclosure, determining the total number of pending processes of the first processor based on the expected transmission count of the first processor and the completed transmission count of the first processor includes: determining the total number of pending processes based on the difference between the expected transmission count of the first processor and the completed transmission count of the first processor.

[0011] In some embodiments of the present disclosure, determining whether there is available credit value in the second completion queue based on the remaining credit value and the total number of pending processes includes: when the remaining credit value is greater than the total number of pending processes, determining that there is available credit value in the second completion queue; and when the remaining credit value is less than or equal to the total number of pending processes, determining that there is no available credit value in the second completion queue.

[0012] In some embodiments of the present disclosure, performing a synchronous balancing operation on the update credit value of the first processor based on whether there is available credit value in the second completion queue includes: when there is available credit value, determining the redistributed credit value based on the remaining credit value, the total number of pending processes, and the number of cores of the first processor, and distributing the redistributed credit value to the first processor; when there is no available credit value, the second processor stops sending commands to the second completion queue, and the second processor is the first processor whose update credit value reaches the first preset value.

[0013] In some embodiments of the present disclosure, determining the reallocation credit value based on the remaining credit value, the total number of tasks to be processed, and the number of cores of the first processor includes: determining the difference between the remaining credit value and the total number of tasks to be processed; and determining the ratio between the difference and the number of cores of the first processor as the reallocation credit value.

[0014] An embodiment of the second aspect of the present disclosure provides a processor synchronization and balancing system, including: a first device, a first processor, and a solid-state drive. The solid-state drive at least includes: a submission command arbitration module and a management module. The first processor at least includes: a command sending synchronization and balancing module. Among them, the first device is configured to send a first submission queue to the submission command arbitration module. The submission command arbitration module is configured to distribute the first command in the first submission queue to the first processor based on a polling algorithm. The management module is configured to determine the initial credit value of the first processor based on the first completion queue corresponding to the first submission queue and the number of cores of the first processor; and is configured to decrementally update the initial credit value based on the number of first commands processed by the first processor to obtain an updated credit value. The command sending synchronization and balancing module is configured to perform a synchronization and balancing operation on the updated credit value of the first processor when the updated credit value reaches a first preset value.

[0015] In some embodiments of the present disclosure, the first device is further configured to store the first command in a submission queue with an idle slot to generate the first submission queue.

[0016] An embodiment of the third aspect of the present disclosure provides a processor synchronization and balancing device, including: a distribution unit configured to distribute the first command in the first submission queue to the first processor; a determination unit configured to determine the initial credit value of the first processor based on the first completion queue corresponding to the first submission queue and the number of cores of the first processor; an update unit configured to decrementally update the initial credit value based on the number of first commands processed by the first processor to obtain an updated credit value; and an equalization unit configured to perform a synchronization and balancing operation on the updated credit value of the first processor when the updated credit value reaches a first preset value.

[0017] An embodiment of the fourth aspect of the present disclosure provides an electronic device, including: a processor and a memory for storing a computer program that can run on the processor. Among them, when the processor is used to run the computer program, it executes the method described in the first aspect embodiment of the present disclosure.

[0018] An embodiment of the fifth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the method described in the first aspect embodiment of the present disclosure.

[0019] A sixth aspect embodiment of the present disclosure provides a computer program product, including a computer program which, when executed by a processor, implements the method described in the first aspect embodiment of the present disclosure.

[0020] In summary, according to a processor synchronization and balancing method provided by the present disclosure, it includes: distributing a first command in a first submission queue to a first processor; determining an initial credit value of the first processor based on a first completion queue corresponding to the first submission queue and the number of cores of the first processor; decreasing and updating the initial credit value based on the number of first commands processed by the first processor to obtain an updated credit value; when the updated credit value reaches a first preset value, performing a synchronization and balancing operation on the updated credit value of the first processor. The method of the present disclosure can adaptively adjust the credit value during processing and the ceding of processor computing power by performing a synchronization and balancing operation on the updated credit value of the first processor when the updated credit value reaches the first preset value, so as to balance the command processing between different processors and ensure the service quality of the processor.

[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0023] Figure 1 is a flowchart of a processor synchronization and balancing method provided by an embodiment of the present disclosure;

[0024] Figure 2 is a flowchart of another processor synchronization and balancing method provided by an embodiment of the present disclosure;

[0025] Figure 3 is a framework diagram of a processor synchronization and balancing system provided by an embodiment of the present disclosure;

[0026] Figure 4 is a topology diagram of a processor synchronization and balancing system proposed by the present disclosure;

[0027] Figure 5 is a flowchart example of a processor synchronization and balancing method provided by an embodiment of the present disclosure;

[0028] Figure 6 is a flowchart example of a processor synchronization and balancing method provided by an embodiment of the present disclosure;

[0029] Figure 7A flowchart example of a processor synchronization and balancing method provided by an embodiment of the present disclosure;

[0030] Figure 8 A schematic structural diagram of a processor synchronization and balancing device provided by an embodiment of the present disclosure;

[0031] Figure 9 A schematic diagram of the hardware composition structure of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0032] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0033] As the transmission speed of the Peripheral Component Interconnect express (PCIe) is getting faster and faster, in order to improve the concurrent processing ability on the host server side, the CPU has long entered the multi-core era. When applying the Non-Volatile Memory Express (NVMe) technology, a multi-queue method for high-concurrent read and write access to NVMe SSDs is adopted, that is, the number of created multi-queues is determined by the comprehensive minimum value of the number of cores of the server CPU and the number of queues supported by the NVMe solid-state drive (SSD). In this way, a multi-core CPU architecture is also adopted on the NVMe SSD side to increase the concurrent processing ability, and the task allocation on these CPUs directly affects the processing efficiency of NVMe commands in the queues corresponding to different CPUs on the host side.

[0034] In the related art, different instances are processed by different CPU cores. For example, an IO command main task, IO_Dispatcher_Task, is created on all cores to receive IO commands and perform event scheduling and processing of various subtasks. An FTL Manager_Task is created on some cores to complete tasks such as mapping management from logical LBA addresses to physical PBA addresses and management of the cache Buffer. Some cores are responsible for garbage collection (GC) of flash blocks and wear leveling. Some cores run a front-end management task, Front_End_Task, for executing management commands and accessing host devices. Some cores run a log management task, MemLog_Task, to complete the saving and output of system logs. Some cores need to run an out-of-band management task, NVMe_MI_Task, to be responsible for the interaction management with the BMC of the server. Each core performs its own duties and coordinates with each other through an interconnected information network. This multi-core architecture better ensures the concurrent ability of IO data processing and the normal management and operation of devices.

[0035] However, due to the incomplete symmetry of the tasks allocated and run on each core, during actual operation, there will be an out-of-step and imbalance of IO commands on different CPUs of the SSD, which in turn causes a large jitter in the IO commands on the server side, affecting the QoS of the IO commands (QoS refers to the ability to complete all requests with stable and consistent performance within a specified time. A relatively common SSD QoS quantization index gives the maximum latency with a confidence level of 99% or 99.99%). Consistency problems occur, further affecting the high service quality of the disk. To ensure QoS, a certain scheduling mechanism is required for the processing of IO commands by multi-core CPUs on the SSD. The currently solved engineering applications mainly perform priority sorting and command distribution on different types of requests on the back-end Flash channel to reduce latency jitter. But what also affects QoS consistency is

[0036] Regarding the consistency of the CPU on the host side with respect to the IO commands from the disk, if the command completion messages from the NVMe disk are more consistent, the provided service stability is better. Therefore, the consistency of the IO command completion messages distributed on different cores of the NVMe SSD determines the consistency of the upper-layer host with respect to the IO commands. However, the different tasks and loads on different CPUs of the NVMe SSD will result in different speeds of return of IO commands on different CPUs and poor consistency.

[0037] To solve the technical problems existing in the related art, an embodiment of the present disclosure provides a processor synchronization and balancing method.

[0038] The embodiments of the present disclosure will be described in detail below.

[0039] As shown Figure 1 in the following, an embodiment of the present disclosure provides a processor synchronization and balancing method, including the following steps:

[0040] Step 101: Distribute the first command in the first submission queue to the first processor.

[0041] In some embodiments, the first submission queue is a submission queue sent by a first device to the device executing this method. The first device may be, for example, a host, a server, or other devices, and the present disclosure does not limit this.

[0042] In some embodiments, the first command is any command included in the first submission queue.

[0043] Exemplarily, the host places the first command into the corresponding free entry of the submission queue through multiple submission queues SQ0, SQ1, SQ2... SQN to generate the first submission queue, where the first submission queue is a circular queue maintained by a head pointer Head and a tail pointer Tail.

[0044] In some embodiments, the first command can be distributed to the first processor by an average distribution method; or the first command can be distributed to the first processor according to the capabilities of the first processor. The present disclosure does not limit this.

[0045] Step 102: Determine the initial credit value of the first processor based on the first completion queue corresponding to the first submission queue and the number of cores of the first processor.

[0046] In some embodiments, the first completion queue is used to store the instructions produced after the first processor finishes processing the first command.

[0047] In some embodiments, the initial credit value of the first processor can be determined according to the first queue depth of the first completion queue and the number of cores of the first processor; for example, the ratio between the first queue depth of the first completion queue and the number of cores of the first processor is determined as the initial credit value.

[0048] It should be understood that each initial credit value corresponds to a first completion queue. In other words, a first submission queue may contain one or more initial credit values.

[0049] For example, according to the depth of the first completion queue #1 and the number of cores of the first processor, it is determined that the initial credit value assigned to the first processor #1 by the first completion queue #1 is 6; according to the depth of the first completion queue #2 and the number of cores of the first processor, it is determined that the initial credit value assigned to the first processor #1 by the first completion queue #2 is 10.

[0050] Step 103: Based on the number of first commands processed by the first processor, decrementally update the initial credit value to obtain an updated credit value.

[0051] In some embodiments, according to the number of first commands processed by the first processor, the initial credit value is decrementally updated. For example, when the first processor finishes processing each first command, the initial credit value corresponding to the first command is decremented by one (e.g., subtract 1 from the initial credit value corresponding to the first command).

[0052] Step 104: When the updated credit value reaches a first preset value, perform a synchronization and balancing operation on the updated credit value of the first processor.

[0053] In some embodiments, when the updated credit value reaches the first preset value, it indicates that there is a situation where the first processor cannot continue to send commands to the first completion queue. At this time, to ensure synchronization and balance among the first processors, a synchronization and balancing operation is performed on the updated credit value of the first processor.

[0054] In some embodiments, the synchronization and balancing operation may include: when there is available credit value, based on the remaining credit value, the total number of commands to be processed, and the number of cores of the first processor, determine the redistributed credit value to distribute the redistributed credit value to the first processor; when there is no available credit value, the second processor stops sending commands to the second completion queue, and the second processor is the first processor whose updated credit value reaches the first preset value.

[0055] In summary, the processor synchronization and balancing method proposed according to the present disclosure includes: distributing the first commands in the first submission queue to the first processor; determining the initial credit value of the first processor based on the first completion queue corresponding to the first submission queue and the number of cores of the first processor; decrementally updating the initial credit value based on the number of first commands processed by the first processor to obtain an updated credit value; when the updated credit value reaches the first preset value, perform a synchronization and balancing operation on the updated credit value of the first processor. The method of the present disclosure can adaptively adjust the credit value during processing and the ceding of processor computing power by performing a synchronization and balancing operation on the updated credit value of the first processor when the updated credit value reaches the first preset value, thereby achieving balance in command processing among different processors and ensuring the service quality of the processors.

[0056] Figure 2 Further shows a flowchart of a processor synchronization and balancing method proposed by the present disclosure. Based on Figure 1 the embodiments shown, further explain Figure 2 It may include the following steps.

[0057] Step 201: Based on the polling algorithm, evenly distribute the first commands to the first processor in sequence.

[0058] In some implementations, the polling algorithm is, for example, the Round-Robin algorithm, but is not limited thereto, and may also be other algorithms that can achieve even distribution.

[0059] Specifically, different first commands in the same first submission queue are evenly distributed to different first processors in turn and cyclically.

[0060] Specifically, different first commands in different first submission queues are evenly distributed to different first processors in turn and cyclically.

[0061] Step 202: Determine an initial credit value based on the ratio between the first queue depth of the first completion queue and the number of cores of the first processor.

[0062] In some embodiments, when initializing the creation of the first completion queue, the initial credit value assigned by the first completion queue to each first processor can be determined by dividing the first queue depth qSize of each first completion queue by the number of cores (the number) of the first processor.

[0063] Optionally, the initial credit value can also be recorded in a preset variable in the memory, for example, recorded in the variable remainingCplQEntries[MAX_CORES][MAX_CPL_QUEUES], where MAX_CORES is the number of cores of the first processor and [MAX_CPL_QUEUES] is the total number of the first completion queues.

[0064] Step 203: Decrease and update the initial credit value based on the number of first commands processed by the first processor to obtain an updated credit value;

[0065] In some embodiments, whenever the first processor finishes processing a first command (that is, the first processor generates an IO command message according to the first command and sends the IO command message to a first completion queue), the initial credit value corresponding to the first completion queue in the first processor is decreased and updated.

[0066] Exemplarily, for example, whenever the first processor #1 sends an IO command message to the first completion queue #1, the initial credit value corresponding to the first completion queue #1 in the first processor is decreased by 1, and the decreased credit value is determined as the updated credit value.

[0067] In other words, when a certain first processor sends the first command to the corresponding first completion queue after internal IO command processing, the number of successfully sent completion commands for the first processor is incremented. Among them, the increment variable can be set to sendCplqMsgCnt[MAX_CORES][MAX_CPL_QUEUES]++.

[0068] Step 204: When the updated credit value reaches the first preset value, determine the remaining credit value of the second completion queue based on the head pointer value of the second completion queue, the tail pointer value of the second completion queue, and the second queue depth of the second completion queue.

[0069] In some embodiments, the first preset value can be 0, but is not limited thereto, and can also be other preset values.

[0070] In some embodiments, the second completion queue is the queue in the first completion queue whose updated credit value reaches the first preset value.

[0071] In some embodiments, when the head pointer value is greater than the tail pointer value, the second queue depth is used to determine the remaining credit value.

[0072] For example, when Head = Tail, the total number of available credits (i.e., the remaining credit value) is the completion queue depth qSize, that is, cq_j_remainCredits = qSize, where cq_j_remainCredits is the remaining credit value.

[0073] In some embodiments, when the head pointer value is less than the tail pointer value, the remaining credit value is determined based on the head pointer value and the tail pointer value.

[0074] For example, when Head>Tail, the total number of available credits cq_j_remainCredits is cq_head – cq_tail – 1, that is, cq_j_remainCredits = cq_head – cq_tail – 1, where cq_head is the value of the head pointer Head (i.e., the head pointer value), and cq_tail is the value of the tail pointer Tail (i.e., the tail pointer value).

[0075] In some embodiments, when the head pointer value is equal to the tail pointer value, the remaining credit value is determined based on the head pointer value, the tail pointer value, and the second queue depth.

[0076] Exemplarily, when Head < Tail, the total number of available credits cq_j_remainCredits is qSize - (cq_tail - cq_head) – 1, that is, cq_j_remainCredits = qSize - (cq_tail – cq_head) - 1.

[0077] Step 205: Determine the total number of commands to be processed by the first processor based on the expected number of commands to be sent by the first processor and the number of commands that have been sent by the first processor.

[0078] In some embodiments, the total number of commands to be processed can be determined based on the difference between the expected number of commands to be sent and the number of commands that have been sent.

[0079] In some embodiments, the expected number of commands to be sent is the number of commands that the first processor expects to send to the second completion queue, and the number of commands that have been sent is the number of commands that the first processor has already sent to the second completion queue.

[0080] In some embodiments, the total number of commands to be processed can be determined by the following formula:

[0081]

[0082] Where cpu_m represents the first processor identified as m, cq_j represents the cq completion queue identified as j, otherCpusOutstandingIOSum represents the total number of commands to be processed, gSqeRecv[cpu_m][cq_j] represents the expected number of commands to be sent, and sendCplgMsgCnt[cpu_m][cq_j] represents the number of commands that have been sent.

[0083] It should be understood that when calculating the sum of the differences for all the first processors above, the difference of the i-th first processor (i.e., the first processor whose credit value update reaches the first preset value) is not included, that is, m != i.

[0084] It should be understood that under normal circumstances, gSqeRecv[cpu_m][cq_j] ≥ sendCplgMsgCnt[cpu_m][cq_j], that is, the above difference cannot be negative.

[0085] Step 206: Determine whether there is available credit in the second completion queue based on the remaining credit value and the total number of commands to be processed.

[0086] In some embodiments, it can be determined whether there is available credit in the second completion queue by comparing the remaining credit value and the total number of commands to be processed.

[0087] Specifically, when the remaining credit value is greater than the total number of items to be processed, it is determined that there is available credit in the second completion queue; when the remaining credit value is less than or equal to the total number of items to be processed, it is determined that there is no available credit in the second completion queue.

[0088] In some embodiments, if the remaining credit value is greater than the total number of items to be processed, that is, cq_j_remainCredits > otherCpusOutstandingIOSum, this indicates that there is available credit in this second completion queue. That is, the i-th first processor (i.e., the first processor whose updated credit value reaches the first preset value) can continue to send the IO command completion messages of the second completion queue.

[0089] In some embodiments, if the remaining credit value is less than or equal to the total number of items to be processed, that is, cq_j_remainCredits < otherCpusOutstandingIOSum, this indicates that there is no available credit in this second completion queue. That is, the second processor processes the IO commands of the second completion queue faster than the other first processors process the IO commands of the second completion queue.

[0090] Step 207, based on whether there is available credit in the second completion queue, perform a synchronization and balancing operation on the updated credit value of the first processor.

[0091] In some embodiments, when there is available credit, based on the remaining credit value, the total number of items to be processed, and the number of cores of the first processor, determine the redistributed credit value to distribute the redistributed credit value to the first processor.

[0092] Specifically, it is possible to determine the difference between the remaining credit value and the total number of items to be processed; and then determine the ratio between the difference and the number of cores of the first processor as the redistributed credit value.

[0093] Exemplarily, the update method of the redistributed credit value remainingCplQEntries[cpu_i][cplq_j] can be as follows:

[0094] remainingCplQEntries[cpu_i][cplq_j]=(cq_j_remainCredits - otherCpusOutstandingIOSum) / MAX_CORES,

[0095] where cq_j_remainCredits represents the number of available credits, otherCpusOutstandingIOSum represents the total number of items to be processed, and MAX_CORES represents the number of cores of the first processor.

[0096] In some embodiments, when there is no available credit value, the second processor stops sending commands to the second completion queue, and the second processor is the first processor whose updated credit value reaches the first preset value.

[0097] In other words, when there is no available credit value in the second completion queue, it means that the second processor processes the IO commands of the second completion queue faster than other first processors. To ensure synchronization between the first processor and the second processor and balance the IO load, it is necessary to stop the second processor from sending IO completion commands to the second completion queue, cede the computing power of the second processor, and perform corresponding processing of IO commands and other tasks on other queues.

[0098] In summary, the processor synchronization and balancing method proposed by the present disclosure includes: based on the polling algorithm, evenly distributing the first commands to the first processors in sequence; determining the initial credit value based on the ratio between the first queue depth of the first completion queue and the number of cores of the first processors; decreasingly updating the initial credit value based on the number of first commands completed by the first processors to obtain the updated credit value; when the updated credit value reaches the first preset value, determining the remaining credit value of the second completion queue based on the head pointer value, tail pointer value, and second queue depth of the second completion queue; determining whether there is available credit value in the second completion queue based on the remaining credit value and the total number of pending processes; and performing a synchronization and balancing operation on the updated credit value of the first processors based on whether there is available credit value in the second completion queue. In the present disclosure, by determining the remaining credit value and the total number of pending processes, it is thus determined whether there is available credit value in the second completion queue, and based on whether there is available credit value, it is determined whether to perform credit value redistribution processing on the first processors or perform computing power ceding processing on the first processors, so as to balance the command processing between different processors and ensure the service quality of the processors.

[0099] Figure 3 Further, a system block diagram of a processor synchronization and balancing system proposed by the present disclosure is shown, which can be applicable to Figure 1 or Figure 2 the method shown.

[0100] In some embodiments, the processor synchronization and balancing system 300 may include a first device 301, a first processor 302, and a solid-state drive 303. The solid-state drive 303 at least includes: a submission command arbitration module 3031 and a management module 3032. The first processor 302 at least includes: a command sending synchronization and balancing module 3021.

[0101] In some embodiments, the first processor 302 and the solid-state drive 303 are physically connected (i.e., directly connected through an interface, a data cable, etc.).

[0102] In some embodiments, there is a communication connection between the first device 301 and the first processor 302, and a communication connection between the first processor 302 and the solid-state drive 303.

[0103] In some embodiments, the first device 301 is configured to send a first submission queue to the submission command arbitration module.

[0104] In some embodiments, the submission command arbitration module 3031 is configured to distribute the first command in the first submission queue to the first processor based on a polling algorithm.

[0105] In some embodiments, the management module 3032 is configured to determine an initial credit value of the first processor based on the first completion queue corresponding to the first submission queue and the number of cores of the first processor; and is configured to decrement and update the initial credit value based on the number of first commands processed by the first processor to obtain an updated credit value.

[0106] In some embodiments, the command sending synchronization and balancing module 3021 is configured to perform a synchronization and balancing operation on the updated credit value of the first processor when the updated credit value reaches a first preset value.

[0107] Optionally, in some embodiments, the first device 301 is further configured to store the first command in a submission queue with an idle slot to generate the first submission queue.

[0108] The following is an exemplary description of the present disclosure:

[0109] Figure 4 The topology schematic diagram of a processor synchronization and balancing system proposed by the present disclosure includes the following:

[0110] Step 1: The host places commands into the corresponding free entries of the submission queues through multiple submission queues SQ0, SQ1, SQ2... SQN. The submission queues are circular queues maintained by a head pointer Head and a tail pointer Tail; inside the SSD, the internal completed IO command messages are sent to the completion queues through the corresponding completion queues CQ0, CQ1, CQ2... CQN. Each completion queue is also maintained by a head pointer Head and a tail pointer Tail.

[0111] Step 2. The ASIC main control chip inside the SSD uses the Round-Robin method to obtain each submission queue command and distribute the commands to each CPU internally (that is, the IO commands in the same submission queue SQ will be evenly scheduled to each CPU by the submission command arbitration module in a circular manner; the IO commands in different submission queues SQs will be circularly scheduled to each CPU by the submission command arbitration module), and perform an increment count of the number of IO submission commands: gSqeRecv[MAX_CORES][MAX_SUB_QUEUES].

[0112] Among them, MAX_CORES is the total number of CPU cores of CPU0 / CPU1… / CPUN, and MAX_SUB_QUEUES is the total number of SQ submission queues of SQ0 / SQ1… / SQN;

[0113] Step 3. When initializing each queue inside the SSD, the initial value of the credit on each CPU of the flow control mechanism will be completed according to the queue depth of each CQ completion queue. The initial value is selected as the queue depth qSize divided by the number of CPUs, that is, the method of evenly dividing the Credit value; in this way, when the host creates multiple pairs of queues, each CPU will have different initial Credit values for sending according to different queue depths and record them in the memory variable remainingCplQEntries[MAX_CORES][MAX_CPL_QUEUES].

[0114] Step 4. After a certain CPU processes the IO command obtained in Step 2 and performs internal IO command processing, it is sent to the corresponding CQ completion queue, and an increment count of the number of successfully sent completion commands on the current CPU is performed, sendCplqMsgCnt[MAX_CORES][MAX_CPL_QUEUES]++, and at the same time, a decrement operation of the sending Credit value is performed;

[0115] Step 5. When the sending Credit value of the completion command of the jth completion queue cq_j used by the ith CPU decreases to 0, that is, remainingCplQEntries[cpu_i][cplq_j] is 0 and there is no available credit value, an update operation of the remaining credit value of the completion queue is required. During the update process, the flow control synchronization and IO balancing operations between this CPU_i and other CPUs are completed. The update method is as follows:

[0116] (1) The ith CPU reads the total value of the remaining credit value of the jth completion queue, cq_j_remainCredits, which is obtained from the sum of the head pointer Head value, tail pointer Tail value, and the depth qSize of the corresponding jth completion queue. The determination method is as follows:

[0117] When Head = Tail, the total number of available credits Credit is the completion queue depth qSize, that is, cq_j_remainCredits = qSize;

[0118] When Head > Tail, the total number of available credits cq_j_remainCredits is cq_head – cq_tail – 1, that is, cq_j_remainCredits = cq_head–cq_tail–1;

[0119] When Head < Tail, the total number of available credits cq_j_remainCredits is qSize - (cq_tail - cq_head) – 1, that is, cq_j_remainCredits = qSize-(cq_tail–cq_head)-1;

[0120] (2)The i-th CPU calculates the total number of other CPUs' currently executing IO commands otherCpusOutstandingIOSum, and the calculation method is as follows:

[0121]

[0122] Among them, cpu_m represents the cpu marked as m, cq_j represents the cq completion queue marked as j. When calculating the difference count sum for all cpus, the i-th cpu is not included, that is, m != i; gSqeRecv[MAX_CORES][MAX_CQ_QUEUES] represents the initial number of IO commands sent to the CQ completion queue corresponding to each CPU, and sendCplgMsgCnt[MAX_CORES][MAX_CQ_QUEUES] represents the format of the completion commands that have been sent to the CQ completion queue. Normally, gSqeRecv[cpu_m][cq_j] ≥ sendCplgMsgCnt[cpu_m][cq_j], so the above difference will not be negative.

[0123] (3)The i-th CPU completes the comparison and judgment of cq_j_remainCredits and otherCpusOutstandingIOSum:

[0124] If cq_j_remainCredits > otherCpusOutstandingIOSum, it means there are available remaining transmission credits in this completion queue. The i-th CPU can continue to send the IO command completion messages of the j-th completion queue. Then, the remaining available credits remainingCplQEntries[cpu_i][cplq_j] are updated as follows:

[0125] remainingCplQEntries[cpu_i][cplq_j]=(cq_j_remainCredits - otherCpusOutstandingIOSum) / MAX_CORES,

[0126] If cq_j_remainCredits ≤ otherCpusOutstandingIOSum, it means there are no available remaining transmission credits in this completion queue. Correspondingly, the processing of the IO commands on queue j of this CPU_i is faster than that of the commands on queue j of other CPUs. To ensure synchronization between different CPUs and balance the IO load, it is necessary to stop sending the IO completion commands on queue j of CPU_i, yield the CPU computing power, and perform corresponding processing of IO commands and other tasks on other queues.

[0127] Figure 5 It is a flowchart of a processor synchronization and balancing method proposed by the present disclosure, which is applied to Figure 4 the system shown in the figure, and is executed by the above management module 3032 to implement flow control variable initialization, including the following steps:

[0128] 1. Initialize the IO SQ to the IO command distribution mode as the Round-Robin cyclic scheduling method.

[0129] 2. Initialize the mapping relationship between the internal SQ and CQ according to the association relationship between the host-created IO SQ and CQ.

[0130] 3.1. Initialize the initial values of the flow control credits for multiple queues on multiple CPUs: remainingCplQEntries[MAX_CORES][MAX_CPL_QUEUES].

[0131] 3.2. Initialize the parameters of the IO synchronization and balancing module on multiple CPUs: gSqeRecv[MAX_CORES][MAX_CQ_QUEUES] and sendCplqMsgCnt[MAX_CORES][MAX_CQ_QUEUES].

[0132] Figure 6The flowchart of a processor synchronization and balancing method proposed by the present disclosure, which is applied to Figure 4 the system shown in the figure, and is executed by the first processor 302 as above to achieve the synchronization and balancing of the first processor 302, including the following steps:

[0133] Step 1: CPU_i processes the IO command to be sent to CQ_j.

[0134] Step 2: CPU_i sends the processed IO command to its own sending synchronization and balancing module for execution.

[0135] Step 3: Determine whether CPU_i has exited the sending to the cq_j queue. If so, go to Step 4; if not, go to Step 1.

[0136] Step 4: Complete the synchronization / balancing between CPUs by calculating the remaining credit value of the cq_j queue and processing other tasks of the local CPU.

[0137] Step 5: CPU_j processes the IO command to be sent to CQ_j.

[0138] Step 6: CPU_j sends the processed IO command to its own command sending synchronization and balancing module for execution.

[0139] Step 3: Determine whether CPU_j has exited the sending to the cq_j queue. If so, go to Step 4; if not, go to Step 1.

[0140] Figure 7 The flowchart of a processor synchronization and balancing method proposed by the present disclosure, which is applied to Figure 4 the system shown in the figure, and is executed by the sending synchronization and balancing module 3031 as above to achieve the synchronization and balancing of the first processor 302, including the following steps:

[0141] Step 1: Perform an increment count according to the IO command identifier distributed by the Round-Robin submission command arbitration module, that is, gSqeRecv[cpu_i][cq_j]++.

[0142] Step 2: After the IO command is processed, it is sent to the corresponding cq_j completion queue.

[0143] Step 3: Determine whether there is an available Credit credit value in the current cq_j completion queue, that is, determine whether remainingCplQEntries[cpu_i][cplq_j]>0. If so, go to Step 4; if not, go to Step 5.

[0144] Step 4, increment the count of successfully sent completion commands on the current CPU: sendCplqMsgCnt[MAX_CORES][MAX_CPL_QUEUES]++, and decrement the available credit value: remainingCplQEntries[cpu_i][cplq_j]--.

[0145] Step 5, update the available credit value credit on the cq_j queue for cpu_i:

[0146] 1) Calculate the sum of outstanding I / Os for the cq_j queue on other cpus except cpu_i (m = 0,... m!= i,…, m = MAX_CORES): otherCpusOutstandingIOSum = ∑((gSqeRecv[cpu_m][cq_j] - sendCplqMsgCnt[cpu_m][cq_j])), where cpu_m represents the cpu identified as m, and cq_j represents the completion queue cq identified as j;

[0147] 2) Read the current total remaining available credit value of the cq_j queue: cq_j_remainCredits = {head, tail, qSize}, where head and tail are obtained by reading the CQ hardware queue registers.

[0148] Step 6, compare the total remaining available credit value of the cq_j queue.

[0149] If cq_j_remainCredits > otherCpusOutstandingIOSum is satisfied, go to Step 7; otherwise, go to Step 8.

[0150] Step 7, calculate the available credit value credit that can be used by cpu_i on cq_j.

[0151] The available credit value credit is: remainingCplQEntries[cpu_i][cplq_j] = (cq_j_remainCredits - otherCpusOutstandingIOSum) / MAX_CORES

[0152] Step 8, exit the sending, wait for the scheduling of the next IO completion command to be sent, and complete the synchronization of the cq_j commands between CPU_i and other CPUs.

[0153] To implement the processor synchronization and balancing method provided by the embodiments of the present disclosure, the embodiments of the present disclosure also provide a processor synchronization and balancing device, asFigure 8 As shown, the processor synchronization and balancing device 800 includes:

[0154] A distribution unit 801 for distributing the first command in the first submission queue to the first processor;

[0155] A determination unit 802 for determining the initial credit value of the first processor based on the first completion queue corresponding to the first submission queue and the number of cores of the first processor;

[0156] An update unit 803 for decrementally updating the initial credit value based on the number of first commands processed by the first processor to obtain an updated credit value;

[0157] A balancing unit 804 for performing a synchronization and balancing operation on the updated credit value of the first processor when the updated credit value reaches a first preset value.

[0158] In some embodiments, the distribution unit 801 is further configured to: based on a polling algorithm, evenly distribute the first commands to the first processor in sequence.

[0159] In some embodiments, the determination unit 802 is further configured to: determine the initial credit value based on the ratio between the first queue depth of the first completion queue and the number of cores of the first processor.

[0160] In some embodiments, the balancing unit 804 is further configured to: determine the remaining credit value of the second completion queue based on the head pointer value, the tail pointer value, and the second queue depth of the second completion queue, where the second completion queue is the queue in the first completion queue whose updated credit value reaches the first preset value; determine the total number of commands to be processed by the first processor based on the expected transmission number of the first processor and the number of completed transmissions of the first processor, where the expected transmission number is the number of commands that the first processor expects to send to the second completion queue, and the number of completed transmissions is the number of commands that the first processor has sent to the second completion queue; determine whether there is available credit value in the second completion queue based on the remaining credit value and the total number of commands to be processed; perform a synchronization and balancing operation on the updated credit value of the first processor based on whether there is available credit value in the second completion queue.

[0161] In some embodiments, the balancing unit 804 is further configured to: when the head pointer value is greater than the tail pointer value, determine the remaining credit value as the second queue depth; when the head pointer value is less than the tail pointer value, determine the remaining credit value based on the head pointer value and the tail pointer value; when the head pointer value is equal to the tail pointer value, determine the remaining credit value based on the head pointer value, the tail pointer value, and the second queue depth.

[0162] In some embodiments, the balancing unit 804 is further configured to: determine the total number of tasks to be processed based on the difference between the expected transmission count and the completed transmission count of the first processor.

[0163] In some embodiments, the balancing unit 804 is further configured to: when the remaining credit value is greater than the total number of tasks to be processed, determine that there is available credit value in the second completion queue; when the remaining credit value is less than or equal to the total number of tasks to be processed, determine that there is no available credit value in the second completion queue.

[0164] In some embodiments, the balancing unit 804 is further configured to: when there is available credit value, determine the redistributed credit value based on the remaining credit value, the total number of tasks to be processed, and the number of cores of the first processor, and distribute the redistributed credit value to the first processor; when there is no available credit value, the second processor stops sending commands to the second completion queue, and the second processor is the first processor whose updated credit value reaches the first preset value.

[0165] In some embodiments, the balancing unit 804 is further configured to: determine the difference between the remaining credit value and the total number of tasks to be processed; determine the ratio between the difference and the number of cores of the first processor as the redistributed credit value.

[0166] In summary, the processor synchronization and balancing device proposed according to the present disclosure includes: a distribution unit configured to distribute the first command in the first submission queue to the first processor; a determination unit configured to determine the initial credit value of the first processor based on the first completion queue corresponding to the first submission queue and the number of cores of the first processor; an update unit configured to perform a decremental update on the initial credit value based on the number of first commands processed by the first processor to obtain an updated credit value; and a balancing unit configured to perform a synchronization and balancing operation on the updated credit value of the first processor when the updated credit value reaches the first preset value. The device of the present disclosure can adaptively adjust the credit value during processing and the ceding of processor computing power by performing a synchronization and balancing operation on the updated credit value of the first processor when the updated credit value reaches the first preset value, so as to balance the command processing between different processors and ensure the service quality of the processors.

[0167] It should be noted that: when the processor synchronization and balancing device provided in the above embodiments performs processor synchronization and balancing, only the above-mentioned division of each program module is used for illustration. In actual applications, the above-mentioned processing can be allocated to different program modules according to needs, that is, the internal structure of the processor synchronization and balancing device is divided into different program modules to complete all or part of the above-described processing. In addition, the processor synchronization and balancing device provided in the above embodiments and the method embodiments of the processor synchronization and balancing method provided in the embodiments of the present disclosure belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.

[0168] Figure 9 The following is a schematic diagram of the hardware composition structure of the electronic device provided by the embodiments of the present disclosure. As Figure 9 shown, the electronic device 900 includes at least one processor 902; and a memory 901 communicatively connected to the at least one processor 902; wherein, the memory 901 stores instructions executable by the at least one processor 902, and the instructions are executed by the at least one processor 902 to implement the steps of the processor synchronization and equalization method described in the embodiments of the present disclosure.

[0169] Optionally, the electronic device may specifically be the processor synchronization and equalization device of the embodiments of the present application, and the electronic device can implement the corresponding processes implemented by the processor synchronization and equalization device in the various methods of the embodiments of the present application. For the sake of brevity, details are not described herein again.

[0170] It can be understood that the electronic device further includes a communication interface 903. Each component in the electronic device is coupled together through a bus system 904. It can be understood that the bus system 904 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 904 further includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 9 all kinds of buses are labeled as the bus system 904.

[0171] It can be understood that the memory 901 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), direct rambus random access memory (DRRAM).The memory 901 described in the embodiments of the present invention is intended to include but not limited to these and any other suitable types of memories.

[0172] The method disclosed in the above embodiments of the present disclosure can be applied to or implemented by the processor 902. The processor 902 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 902 or the instructions in the form of software. The above-mentioned processor 902 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 902 can implement or execute each method, step, and logic block diagram disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the memory 901. The processor 902 reads the information in the memory 901 and combines its hardware to complete the steps of the foregoing method.

[0173] In an exemplary embodiment, the electronic device may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components for performing the foregoing method.

[0174] The embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to implement the steps of the processor synchronization and equalization method described in the embodiments of the present disclosure when executed.

[0175] The embodiments of the present disclosure also provide a computer program product, including a computer program, and the computer program implements the steps of the processor synchronization and equalization method described in the embodiments of the present disclosure when executed by the processor.

[0176] Optionally, the computer-readable storage medium can be applied to the processor synchronization and equalization device in the embodiments of the present application, and the computer instructions cause the computer to execute the corresponding processes implemented by the processor synchronization and equalization device in each method of the embodiments of the present application. For the sake of brevity, it will not be elaborated here.

[0177] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the displayed or discussed components can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0178] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0179] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0180] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes: various media such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks that can store program codes.

[0181] Alternatively, if the above-mentioned integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present invention. The foregoing storage medium includes: various media such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks that can store program codes.

[0182] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the said claims.

Claims

1. A processor synchronization and equalization method, characterized in that, The method includes: Based on a polling algorithm, sequentially and evenly distributing the first commands in the first submission queue to the first processors; Determining an initial credit value of the first processor based on a ratio between a first queue depth of a first completion queue corresponding to the first submission queue and the number of cores of the first processor; Based on the number of first commands processed by the first processor, performing a decrement update on the initial credit value to obtain an updated credit value; When the updated credit value reaches a first preset value, performing a synchronization and balancing operation on the updated credit value of the first processor; Wherein, when the updated credit value reaches the first preset value, performing the synchronization and balancing operation on the updated credit value of the first processor includes: Based on a head pointer value of a second completion queue, a tail pointer value of the second completion queue, and a second queue depth of the second completion queue, determining a remaining credit value of the second completion queue, where the second completion queue is the queue in the first completion queue for which the updated credit value reaches the first preset value; Based on an expected transmission number of the first processor and a completed transmission number of the first processor, determining a total number of commands to be processed by the first processor, where the expected transmission number is the number of commands that the first processor expects to transmit to the second completion queue, and the completed transmission number is the number of commands that the first processor has already transmitted to the second completion queue; Based on the remaining credit value and the total number of commands to be processed, determining whether there is available credit value in the second completion queue to obtain a judgment result; Based on the judgment result, performing a synchronization and balancing operation on the updated credit value of the first processor.

2. The processor synchronization and equalization method according to claim 1, wherein The determining the remaining credit value of the second completion queue based on the head pointer value of the second completion queue, the tail pointer value of the second completion queue, and the second queue depth of the second completion queue includes: When the head pointer value is greater than the tail pointer value, using the second queue depth to determine the remaining credit value; When the head pointer value is less than the tail pointer value, determining the remaining credit value based on the head pointer value and the tail pointer value; When the head pointer value is equal to the tail pointer value, determining the remaining credit value based on the head pointer value, the tail pointer value, and the second queue depth.

3. The processor synchronization and equalization method according to claim 1, wherein The determining the total number of commands to be processed by the first processor based on the expected transmission number of the first processor and the completed transmission number of the first processor includes: Based on a difference between the expected transmission number of the first processor and the completed transmission number of the first processor, determining the total number of commands to be processed.

4. The processor synchronization equalization method according to claim 1, wherein The determining whether there is available credit value in the second completion queue based on the remaining credit value and the total number of commands to be processed includes: When the remaining credit value is greater than the total number of commands to be processed, determining that there is available credit value in the second completion queue; When the remaining credit value is less than or equal to the total number of commands to be processed, determining that there is no available credit value in the second completion queue.

5. The processor synchronization equalization method according to claim 1, characterized in that The performing a synchronization and balancing operation on the updated credit value of the first processor based on whether there is available credit value in the second completion queue includes: When there is the available credit value, determine a reallocated credit value based on the remaining credit value, the total number of items to be processed, and the number of cores of the first processor, so as to distribute the reallocated credit value to the first processor; When there is no available credit value, the second processor stops sending commands to the second completion queue, where the second processor is the first processor whose updated credit value reaches a first preset value.

6. The processor synchronization and equalization method according to claim 5, wherein The determining the reallocated credit value based on the remaining credit value, the total number of items to be processed, and the number of cores of the first processor includes: Determine the difference between the remaining credit value and the total number of items to be processed; Determine the ratio between the difference and the number of cores of the first processor as the reallocated credit value.

7. A processor synchronization and equalization system, characterized in that, The system includes: a first device, a first processor, and a solid-state drive, where the solid-state drive at least includes: a submission command arbitration module and a management module, and the first processor at least includes: a command sending synchronization and balancing module; Among them, the first device is used to send a first submission queue to the submission command arbitration module; The submission command arbitration module is used to evenly distribute the first commands in the first submission queue to the first processor in sequence based on a polling algorithm; The management module is used to determine the initial credit value of the first processor based on the ratio between the first queue depth of the first completion queue corresponding to the first submission queue and the number of cores of the first processor; It is used to decrement and update the initial credit value based on the number of first commands processed by the first processor to obtain an updated credit value; The command sending synchronization and balancing module is used to perform a synchronization and balancing operation on the updated credit value of the first processor when the updated credit value reaches a first preset value; Among them, when the updated credit value reaches the first preset value, performing the synchronization and balancing operation on the updated credit value of the first processor includes: Determine the remaining credit value of the second completion queue based on the head pointer value of the second completion queue, the tail pointer value of the second completion queue, and the second queue depth of the second completion queue, where the second completion queue is the queue in the first completion queue whose updated credit value reaches the first preset value; Determine the total number of items to be processed by the first processor based on the expected number of commands to be sent by the first processor and the number of commands already sent by the first processor to the second completion queue, where the expected number of commands to be sent is the number of commands that the first processor expects to send to the second completion queue, and the number of commands already sent is the number of commands that the first processor has sent to the second completion queue; Determine whether there is available credit value in the second completion queue based on the remaining credit value and the total number of items to be processed, and obtain a judgment result; Based on the judgment result, perform a synchronization and balancing operation on the updated credit value of the first processor.

8. The processor synchronization and equalization system according to claim 7, wherein The first device is further used to store the first command in a submission queue with an idle slot to generate the first submission queue.

9. A processor synchronization and equalization device, characterized in that, Includes: A distribution unit, configured to evenly distribute the first commands in the first submission queue to the first processor in sequence based on a polling algorithm; A determination unit, configured to determine an initial credit value of the first processor based on a ratio between a first queue depth of a first completion queue corresponding to the first submission queue and the number of cores of the first processor; An update unit, configured to perform a decrement update on the initial credit value based on the number of first commands processed by the first processor, so as to obtain an updated credit value; An equalization unit, configured to perform a synchronization and equalization operation on the updated credit value of the first processor when the updated credit value reaches a first preset value; Wherein, when the updated credit value reaches the first preset value, performing a synchronization and equalization operation on the updated credit value of the first processor includes: Determining a remaining credit value of the second completion queue based on a head pointer value of the second completion queue, a tail pointer value of the second completion queue, and a second queue depth of the second completion queue, where the second completion queue is a queue in the first completion queue whose updated credit value reaches the first preset value; Determining a total number of pending processes of the first processor based on an expected transmission number of the first processor and a completed transmission number of the first processor, where the expected transmission number is the number of commands that the first processor expects to send to the second completion queue, and the completed transmission number is the number of commands that the first processor has sent to the second completion queue; Determining whether there is available credit value in the second completion queue based on the remaining credit value and the total number of pending processes, to obtain a judgment result; Performing a synchronization and equalization operation on the updated credit value of the first processor based on the judgment result.

10. An electronic device, characterized in that, Comprising: A processor and a memory for storing a computer program that can run on the processor, Wherein, when the processor is used to run the computer program, it executes the processor synchronization and equalization method according to any one of claims 1-6.

11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the processor synchronization and equalization method according to any one of claims 1-6.

12. A computer program product, characterized in that, Comprising a computer program, where the computer program, when executed by a processor, implements the processor synchronization and equalization method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Message scheduling method and device

    CN106856459A

  • NVMe SSD read-write method and system suitable for multiple cores

    CN118131996A