Heterogeneous computing scheduling method, device and equipment based on proxy thread and storage medium

By introducing proxy threads and lock-free communication queues in the Ceph distributed storage system, the problems of thread blocking and queue number limitation in the single-threaded synchronous IO mode are solved, high concurrency performance and wide device compatibility are achieved, and system performance and applicability are optimized.

CN120849038APending Publication Date: 2025-10-28JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510866738.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

When using SPDK, the existing Ceph distributed storage system suffers from thread blocking and queue size limitations due to the single-threaded synchronous I/O mode, which restricts the high concurrency performance of NVMe devices and the compatibility with low-end devices.

Method used

A heterogeneous computing scheduling method based on agent threads is adopted. By creating worker thread groups and agent threads, and using lock-free communication queues for connection, the worker thread groups handle upper-layer business tasks, and the agent threads manage the IO operations of the SPDK user-space NVMe driver, thereby achieving asynchronous IO and reducing the number of IO queues.

Benefits of technology

It improves the high-concurrency performance of NVMe devices, enhances compatibility with low-end devices, reduces the overhead of thread switching and synchronization waiting, and improves system efficiency and applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849038A_ABST
    Figure CN120849038A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous computing scheduling method and device based on a proxy thread, computer equipment and a storage medium, the scheduling method comprises the steps that a working thread group and the proxy thread are created, the working thread group comprises at least one working thread, and a lockless communication queue is established to connect the working thread group and the proxy thread; in response to the received to-be-processed task of the execution unit, the working thread submits an operation request of the to-be-processed task to the proxy thread through the lock-free communication queue; in response to successful submission of the operation request, performing local data calculation by the working thread, and entering dormancy if the local data calculation is completed; in response to the operation request received by the proxy thread, the proxy thread submits the operation request to an execution unit and polls whether the to-be-processed task is processed or not; in response to completion of processing of the to-be-processed task, the proxy thread wakes up a corresponding working thread through a callback function and returns a processing result; by means of the method, the high concurrency performance and the compatibility of heterogeneous equipment can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data storage management technology, and in particular to a heterogeneous computing scheduling method, apparatus, device and storage medium based on agent threads. Background Technology

[0002] The Ceph distributed storage system manages data storage through a decentralized RADOS cluster, supporting block / file / object services. It utilizes the CRUSH algorithm for automatic data balancing and fault recovery. Core components include a Monitor cluster, OSD nodes, and backend storage devices. OSD threads access block devices in kernel mode via system calls, triggering user-mode-kernel context switching and interrupt handling, leading to latency jitter and scheduling overhead. Multi-threaded concurrency also causes queue lock contention. The SPDK user-mode driver directly manipulates hardware registers and polls device status, avoiding kernel switching overhead but employing a single-threaded synchronous processing mechanism: the Reactor thread simultaneously handles business logic processing, I / O submission, and polling, resulting in thread blocking and a queue depth limited to 1.

[0003] In traditional Ceph distributed storage systems, the Object Storage Daemon (OSD) initiates system calls through the block device interface provided by the operating system to write data to the storage device. This process requires switching from user mode to kernel mode, resulting in context switching and system interrupts, increasing scheduling overhead and causing performance bottlenecks. The SPDK user-mode NVMe driver improves performance by directly manipulating device registers and polling I / O to complete the status, avoiding kernel mode switching and interrupt overhead.

[0004] However, existing Ceph support for SPDK is primarily based on SPDK's Reactor-Core model, where a single thread handles both upper-layer business logic and lower-layer I / O simultaneously, achieving asynchronous I / O through loop polling. This is incompatible with Ceph's traditional multi-threaded asynchronous I / O model. Currently, Ceph OSDs rely on thread pools to schedule multiple threads to execute tasks asynchronously. When using SPDK, a single thread synchronously initiates I / O and polls for completion. During this waiting period, the thread cannot handle other tasks, leading to frequent thread switching. Furthermore, the I / O queue depth is only 1, limiting the high-concurrency performance of NVMe devices. In addition, each thread needs an independent I / O queue to avoid concurrent access conflicts, resulting in an excessive number of queues and poor compatibility with low-end NVMe devices that support a limited number of queues. This invention introduces a proxy thread to uniformly manage SPDK user-space driven I / O operations, achieving asynchronous I / O, reducing the number of I / O queues, thereby improving concurrency performance, reducing thread switching overhead, and enhancing compatibility with low-end NVMe devices. This significantly optimizes the performance and applicability of Ceph distributed storage systems. Summary of the Invention

[0005] Therefore, it is necessary to provide a heterogeneous computing scheduling method, apparatus, device, and storage medium based on proxy threads that can eliminate thread blocking and queue number limitations, significantly improve high-concurrency performance and compatibility with heterogeneous devices, in order to address the above-mentioned technical problems.

[0006] Firstly, a heterogeneous computing scheduling method based on proxy threads is provided, including:

[0007] Create a worker thread group and a proxy thread. The worker thread group contains at least one worker thread. The worker thread is used to process local data calculations and control the execution unit to process tasks.

[0008] Establish a lock-free communication queue to connect the worker thread group and the agent thread;

[0009] In response to receiving a task to be processed by the execution unit, the worker thread submits the operation request of the task to be processed to the agent thread through a lock-free communication queue;

[0010] In response to a successful operation request submission, the worker thread performs local data calculations. Once the local data calculations are complete, it enters a sleep state.

[0011] When the agent thread receives an operation request, it submits the operation request to the execution unit and polls to see if the pending task has been completed.

[0012] In response to the completion of the pending task, the agent thread wakes up the corresponding worker thread through the callback function and returns the processing result.

[0013] Secondly, a heterogeneous computing scheduling device is provided, applied to the heterogeneous computing scheduling method based on agent threads described in the first aspect, including:

[0014] The thread creation module is used to create worker thread groups and agent threads. A worker thread group contains at least one worker thread, which is used to process local data calculations and control the execution unit to process tasks.

[0015] A communication queue creation module is used to establish a lock-free communication queue connection between the worker thread group and the proxy thread;

[0016] The task receiving module is used to respond to the receiving of a task to be processed by the execution unit. In this case, the worker thread submits the operation request of the task to be processed to the agent thread through a lock-free communication queue.

[0017] The request submission module is used to respond to the successful submission of an operation request. In this case, the worker thread performs local data calculations, and if the local data calculations are completed, it enters a sleep state.

[0018] The execution module is used to respond to the agent thread receiving an operation request. The agent thread submits the operation request to the execution unit and polls to see if the pending task has been completed.

[0019] The result return module is used to respond to the completion of the pending task. In this case, the agent thread wakes up the corresponding worker thread through the callback function and returns the processing result.

[0020] Thirdly, a computer device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the heterogeneous computing scheduling method based on proxy threads described in the first aspect.

[0021] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the heterogeneous computing scheduling method based on proxy threads described in the first aspect is implemented.

[0022] By implementing the aforementioned heterogeneous computing scheduling method, apparatus, device, and storage medium based on proxy threads, this method creates a worker thread group containing at least one worker thread to handle upper-layer business tasks, while simultaneously creating a single proxy thread to manage the IO operations of the SPDK user-space NVMe driver. The worker thread group and the proxy thread are connected via a lock-free communication queue. Worker threads submit asynchronous IO operation requests to the proxy thread through this queue, ensuring high efficiency and concurrency safety. Upon receiving a request, the proxy thread submits it to the NVMe device and polls for completion status, avoiding the overhead of thread blocking or frequent switching in the traditional Ceph OSD synchronous IO mode. After the operation is completed, the proxy thread wakes up the corresponding worker thread through a callback function and returns the result, achieving rapid task response. This mechanism overcomes the limitation of the SPDK single-threaded IO queue depth of 1, improves the high-concurrency performance of NVMe devices, and reduces the number of queues by uniformly managing the IO queue through a single proxy thread, enhancing compatibility with low-end NVMe devices with limited queue support. Lock-free communication queues and callback functions further reduce the overhead of thread switching and synchronization waiting, improve system efficiency, and provide Ceph distributed storage systems with a block device implementation solution that offers higher performance, lower latency, and wider adaptability. Attached Figure Description

[0023] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1A flowchart illustrating a heterogeneous computing scheduling method based on proxy threads provided in this application embodiment;

[0025] Figure 2 A structural block diagram of a heterogeneous computing scheduling device based on proxy threads provided in this application embodiment;

[0026] Figure 3 A timing diagram of a heterogeneous computing scheduling method based on proxy threads provided in an embodiment of this application;

[0027] Figure 4 This is a diagram showing the internal structure of a computer device in an embodiment of this application. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0029] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0030] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0031] In one embodiment, such as Figure 1 As shown, a heterogeneous computing scheduling method based on proxy threads is provided, including:

[0032] S100: Create a worker thread group and a proxy thread, wherein the worker thread group contains at least one worker thread, and the worker thread is used to process local data calculations and control the execution unit to process tasks.

[0033] S200: Establish a lock-free communication queue connection between the worker thread group and the agent thread;

[0034] S300: In response to receiving a task to be processed by the execution unit, the worker thread submits the operation request of the task to be processed to the proxy thread through the lock-free communication queue;

[0035] S400: In response to the successful submission of the operation request, the worker thread performs the local data calculation, and enters sleep mode if the local data calculation is completed;

[0036] S500: In response to the proxy thread receiving the operation request, the proxy thread submits the operation request to the execution unit and polls whether the task to be processed has been completed.

[0037] S600: In response to the completion of the pending task, the agent thread wakes up the corresponding worker thread through the callback function and returns the processing result.

[0038] In this context, a worker thread group refers to a collection of threads consisting of at least one worker thread, responsible for executing upper-layer business logic tasks of the Object Storage Daemon (OSD) in the Ceph distributed storage system, such as handling client requests or data management operations; a worker thread refers to a single thread within the worker thread group; a proxy thread refers to a single thread specifically responsible for managing the IO operations of the SPDK user-space NVMe driver, receiving operation requests from worker threads, submitting them to the NVMe device, polling the completion status, and returning results through callback functions; a lock-free communication queue is an efficient and concurrently safe communication mechanism used to connect the worker thread group and the proxy thread, allowing worker threads to submit operation requests asynchronously without the need for lock mechanisms to reduce synchronization overhead; an operation request refers to the IO task submitted by the worker thread to the proxy thread, typically involving read and write operations on the NVMe device; an execution unit refers to the NVMe device hardware responsible for actually executing the IO operations submitted by the proxy thread; polling refers to the proxy thread repeatedly checking the status of the NVMe device to determine whether the operation request is complete, avoiding interruption or blocking; a callback function is a function called by the proxy thread after the operation request is completed; and the processing result refers to the data or status returned by the NVMe device after the operation request is executed.

[0039] Specifically, a worker thread group is created, containing at least one worker thread, to handle upper-layer business tasks of the Ceph OSD. Simultaneously, a single agent thread manages the IO operations of the SPDK user-space NVMe driver. The worker thread group and the agent thread establish a connection through a lock-free communication queue. The worker threads submit asynchronous IO operation requests to the agent thread through this queue, ensuring high efficiency and concurrency safety. Upon receiving a request, the agent thread submits it to the execution unit of the NVMe device and monitors the operation completion status via polling, avoiding the overhead of thread blocking or frequent context switching in the traditional Ceph OSD synchronous IO mode. After the operation is completed, the agent thread wakes up the corresponding worker thread through a callback function and returns the processing result, achieving rapid task response. This mechanism overcomes the limitation of the SPDK single-threaded IO queue depth of 1, significantly improving the high-concurrency performance of NVMe devices. Furthermore, by uniformly managing the IO queue through a single agent thread, the number of queues is reduced, enhancing compatibility with low-end NVMe devices with limited queue support. The combination of lock-free communication queues and callback functions further reduces the overhead of thread switching and synchronization waiting, improves the overall system efficiency, and provides Ceph distributed storage system with a block device implementation solution that offers higher performance, lower latency, and wider adaptability.

[0040] In one embodiment, establishing a lock-free communication queue connection between the worker thread group and the agent thread includes:

[0041] By calling a thread-safe data structure function, shared memory is set up, and a lock-free communication queue is constructed in the shared memory for transmitting the operation request between the worker thread group and the proxy thread;

[0042] Initialize the producer head pointer and consumer tail pointer of the lock-free communication queue;

[0043] The operation permissions of the consumer tail pointer are bound to the agent thread, and the operation permissions of the producer head pointer are assigned to the worker thread group.

[0044] Thread-safe data structure functions refer to a set of functions used to safely operate on shared data in a multi-threaded environment, ensuring that data races or inconsistencies are avoided when constructing a lock-free communication queue. Shared memory refers to a memory area shared by worker threads and agent threads, used to store the data structures of the lock-free communication queue, enabling efficient cross-thread communication. The lock-free communication queue is a concurrently safe queue built on shared memory. Worker threads and agent threads pass operation requests through this queue without using lock mechanisms to reduce synchronization overhead. The producer head pointer points to the head of the lock-free communication queue, used by worker threads to insert new operation requests and identify a writable position in the queue. The consumer tail pointer points to the tail of the lock-free communication queue, used by agent threads to read operation requests and identify a readable position in the queue. Operation permissions refer to access control over the producer head pointer or consumer tail pointer, restricting specific threads' operations on the pointers to ensure the queue's concurrency safety. The operation request data template is a predefined structure containing fields for operation requests, used to standardize the IO tasks submitted by worker threads to agent threads. The thread identifier is a unique identifier used to identify the worker thread submitting the request within the operation request. This allows the agent thread to wake up the corresponding thread via a callback function after the operation is completed. The operation type refers to the category of I / O operation specified in the operation request, such as read or write, which guides the agent thread to execute specific NVMe device operations. The request data refers to the specific data of the I / O task included in the operation request, such as the data to be written or the target address to be read. The callback function pointer is a pointer to a callback function in the operation request. The agent thread calls this function after the operation is completed to wake up the worker thread and return the processing result.

[0045] Specifically, a lock-free communication queue is constructed in shared memory by calling thread-safe data structure functions, enabling efficient transmission of operation requests between worker thread groups and agent threads. The producer head pointer and consumer tail pointer of the lock-free communication queue are initialized. Operation permissions for the producer head pointer are assigned to the worker thread group for inserting operation requests; operation permissions for the consumer tail pointer are bound to the agent thread for reading requests. Based on the structure of the lock-free communication queue, an operation request data template is designed, pre-setting fields including thread identifier, operation type, request data, and callback function pointer to ensure the standardization and completeness of requests. Worker threads write operation requests to the queue via the producer head pointer, and agent threads read requests and submit them to the NVMe device via the consumer tail pointer. After polling the operation completion status, the corresponding worker thread is woken up via a callback function and the result is returned. This mechanism achieves asynchronous I / O operations, breaking the limitation of the SPDK single-threaded I / O queue depth of 1, and significantly improving the high-concurrency performance of NVMe devices. The lock-free communication queue, through shared memory and pointer permission allocation, eliminates lock contention and synchronization overhead, reducing thread switching costs. A single agent thread manages the IO queues uniformly, reducing the number of queues and enhancing compatibility with low-end NVMe devices that support a limited number of queues. The standardized design of operation request data templates improves request processing efficiency and system stability, providing Ceph distributed storage systems with a higher-performance, lower-latency, and more adaptable block device implementation solution.

[0046] In one embodiment, in response to a worker thread submitting an operation request to a proxy thread via a lock-free communication queue, the following steps are included:

[0047] The worker thread sets the request message structure based on the currently pending operation requests;

[0048] When an available enqueue slot is available in the lock-free communication queue, the worker thread writes the operation request into the specified slot of the lock-free communication queue through the lock-free mechanism and updates the producer head pointer.

[0049] The standardized request message structure refers to a predefined data structure containing fields for the operation request (such as thread identifier, operation type, request data, and callback function pointer), used to standardize the I / O tasks submitted by worker threads. Available enqueue slots refer to storage locations in the lock-free communication queue that have not yet been occupied by operation requests; worker threads can write new requests to these locations. The lock-free mechanism refers to the safe updating of the queue state (such as the producer head pointer) in a multi-threaded environment through techniques such as atomic operations or memory barriers, avoiding the use of locks to reduce overhead. The producer head pointer is a pointer to the head of the lock-free communication queue, manipulated by worker threads, used to identify the position where an operation request can be written. Updating the producer head pointer means that after a worker thread writes an operation request, it adjusts the producer head pointer to point to the next available slot, ensuring correct queue operation.

[0050] Specifically, worker threads construct standardized request message structures based on the currently pending I / O tasks. These structures include fields such as thread identifier, operation type, request data, and callback function pointers, ensuring the standardization and completeness of the requests. When an available enqueue slot exists in the lock-free communication queue, the worker thread writes the operation request to the designated slot in the queue using a lock-free mechanism and atomically updates the producer head pointer, completing the request submission. The proxy thread reads the operation requests from the queue using the consumer tail pointer, submits them to the NVMe device execution unit, and polls for completion status. After the operation is completed, it wakes up the corresponding worker thread through a callback function and returns the result. This mechanism achieves efficient asynchronous I / O operations, significantly improving the high-concurrency performance of NVMe devices. The standardized request message structure standardizes the operation request format, improving the stability and efficiency of request processing. The lock-free communication queue eliminates lock contention through a lock-free mechanism, reducing the overhead of thread switching and synchronization waiting, and ensuring efficient communication between worker threads and proxy threads.

[0051] In one embodiment, in response to the agent thread receiving an operation request, the agent thread submits the operation request to the execution unit and polls whether the pending tasks have been completed, including:

[0052] The proxy thread obtains operation requests through the consumer interface of the lock-free communication queue and updates the consumer tail pointer;

[0053] The proxy thread submits the operation request to the execution unit and records the operation request status mapping;

[0054] In response to the arrival of a preset time interval, the agent thread queries the execution unit to see if the pending task has been completed.

[0055] Updating the consumer tail pointer refers to the agent thread atomically adjusting the consumer tail pointer to point to the next readable slot after reading an operation request, ensuring the correctness of the queue operation. The request status mapping is a record structure maintained by the agent thread to track the submission status (e.g., pending, completed) of each operation request and its associated information (e.g., callback function pointer). The preset time interval is a fixed period for the agent thread to poll the execution unit status, used to balance polling frequency and resource consumption. The operation request status refers to the execution status of the IO task returned by the execution unit, such as whether the operation is completed or the result data.

[0056] Specifically, the proxy thread reads operation requests submitted by worker threads through the consumer interface of the lock-free communication queue and atomically updates the consumer's tail pointer to ensure the queue's concurrency safety. After reading, the proxy thread submits the operation request to the execution unit of the NVMe device, and simultaneously records the request's identifier and status in the request status map for easy tracking later. At preset time intervals, the proxy thread queries the execution unit's operation request status to check if the I / O task is complete, and wakes up the corresponding worker thread and returns the result via a callback function upon completion. This step achieves efficient request retrieval through lock-free operations using the consumer interface and tail pointer, reducing the proxy thread's communication overhead. The introduction of the request status map allows the proxy thread to accurately track the progress of multiple concurrent requests, avoiding resource waste from unordered queries.

[0057] In one embodiment, such as Figure 3 As shown, in response to the completion of the pending task, the agent thread wakes up the corresponding worker thread through a callback function and returns the processing result, including:

[0058] Based on the preset mapping relationship between operation handles and operation requests, the request message structure corresponding to the operation request is retrieved. The request message structure contains: a callback function pointer and worker thread context information, wherein the callback function pointer points to a preset callback function.

[0059] The processing result and the callback function pointer are encapsulated into a callback parameter structure, and the callback parameter structure is bound to the request message structure;

[0060] The callback function is invoked, which wakes up the corresponding worker thread based on the synchronization object in the worker thread context information.

[0061] In response to a worker thread being woken up, the corresponding processing result is obtained from the request message structure;

[0062] In response to the completion of the callback function, the proxy thread releases the resources corresponding to the operation request.

[0063] The mapping relationship refers to the record table maintained by the agent thread, which stores the correspondence between operation handles and the original request message structure, used to track the status and associated information of the request. The original request message structure is a standardized data structure for the operation request submitted by the worker thread, containing fields such as thread identifier, operation type, request data, and callback function pointer. The callback parameter structure is a data structure encapsulating the result of the operation request processing, containing the return data or status of the IO operation, which is passed to the worker thread by the callback function. Resource usage refers to the memory or record entries occupied by the operation request in the mapping relationship table, which must be released after the callback is completed to avoid resource leaks.

[0064] Specifically, based on the preset operation handles and mapping relationships, the corresponding original request message structure is retrieved, and its thread identifier and callback function pointer are obtained. The proxy thread encapsulates the processing result (such as read data or operation status) into a callback parameter structure and calls the callback function in the original request message structure, passing the callback parameter structure as a parameter. The callback function wakes up the corresponding worker thread, which extracts the processing result based on the callback parameter structure and continues to execute subsequent business logic. After the callback is completed, the proxy thread releases the resource occupation of the operation request in the mapping relationship table, ensuring efficient resource management. Accurate retrieval of operation handles and mapping relationships enables rapid request matching and reduces the overhead of callback processing. The standardized encapsulation of the callback parameter structure ensures the complete transmission of processing results, improving the reliability and consistency of data processing. The callback function calling mechanism allows worker threads to quickly resume execution without additional synchronization waiting, significantly improving task response speed. The resource release mechanism effectively prevents memory leaks and optimizes the long-term operational stability of the system.

[0065] In one embodiment, the method further includes:

[0066] In response to the worker thread submitting the operation request, the worker thread performs local data calculations to obtain the processed value;

[0067] Once local data computation is complete, the worker thread enters a sleep state.

[0068] In response to the worker thread being awakened, the worker thread will perform secondary processing on the processed value and the processed result.

[0069] The processing value refers to the result data generated by the worker thread after submitting an operation request and performing data processing operations, such as the preprocessing, format conversion, or calculation results of business data.

[0070] The blocked state refers to the state in which the worker thread pauses execution and releases CPU resources after completing data processing, waiting for the agent thread to wake it up through a callback function, in order to avoid busy waiting and resource waste.

[0071] Secondary processing refers to the operation of integrating or further processing the previously generated processing values ​​with the processing results (return data or status of IO operations) returned by the proxy thread after the worker thread is awakened, such as data merging, verification, or business logic calculation.

[0072] Specifically, processing values ​​are generated based on the content of the operation request, such as preprocessing or format conversion of business data. After data processing is complete, the worker thread enters a blocked state, releasing resources and avoiding meaningless busy waiting, thereby improving system resource utilization. In response to the agent thread waking up the worker thread through a callback function, the worker thread resumes execution, performing secondary processing on the previously generated processing value and the processing result returned by the agent thread, such as merging data or executing verification logic to complete the business task. This step, by parallelizing data processing and I / O operations, allows the worker thread to continue executing data processing tasks after submitting an operation request, making full use of the idle time waiting for I / O completion and significantly improving thread task processing efficiency. The design of the worker thread entering a blocked state avoids resource waste and optimizes resource scheduling in a multi-threaded environment, especially reducing thread contention in high-concurrency scenarios. The secondary processing mechanism, by integrating the processing value and I / O result, enhances the flexibility and integrity of business logic and ensures the seamless connection between data processing and I / O operations.

[0073] In one embodiment, it also includes:

[0074] The proxy thread submits the operation request to the execution unit and records the submission timestamp of the operation request.

[0075] Write the submission timestamp into the data structure corresponding to the operation request;

[0076] Set corresponding timeout thresholds for operation requests submitted to the execution unit;

[0077] If the result of the query from the proxy thread to the execution unit is that the operation is not completed, then the time difference between the current system time and the submission timestamp in the data structure is calculated.

[0078] If the time difference is greater than the timeout threshold and the operation request has not been completed, the operation request will be judged as a timeout request.

[0079] For the operation request corresponding to the timeout request, construct a processing result carrying error code and timeout description information, and encapsulate it into a callback parameter structure;

[0080] Invoke the callback function corresponding to the operation request, wake up the worker thread bound to the operation request through the thread synchronization mechanism and pass the timeout processing result;

[0081] In response to the worker thread recognizing a timeout state, the waiting process is terminated and a new operation request is received;

[0082] In response to the completion of the callback function, the proxy thread releases the resources corresponding to the operation request.

[0083] The submission timestamp refers to the system time recorded when the proxy thread submits the operation request to the execution unit, used to track the submission time of the request. The timeout threshold is the maximum waiting time set for the operation request; failure to complete within this time is considered a timeout. The time difference is the difference between the current system time and the operation request submission timestamp, used to determine if the operation has timed out. A timeout request is an operation request whose time difference exceeds the timeout threshold and is still not completed; it is judged as an abnormal state. The error code is an identifier included in the timeout request processing result, used to indicate the specific type of timeout error. The timeout description information is the text or data included in the timeout request processing result, describing the timeout reason or status. The thread synchronization mechanism is a mechanism used to coordinate the communication between the proxy thread and the worker thread, ensuring that the callback function safely wakes up the worker thread.

[0084] Specifically, when the proxy thread submits an operation request to the execution unit, it records the submission timestamp and writes it into the operation request's data structure, while simultaneously setting a timeout threshold for the request. When the execution unit query result shows that the operation is incomplete, the proxy thread calculates the time difference between the current system time and the submission timestamp. If the time difference exceeds the timeout threshold and the operation is still incomplete, it is determined to be a timeout request, and a processing result containing error codes and timeout description information is constructed and encapsulated in a callback parameter structure. A callback function is called through a thread synchronization mechanism to wake up the bound worker thread and deliver the timeout result. After recognizing the timeout status, the worker thread terminates the waiting process and receives new operation requests. After the callback is completed, the proxy thread releases the resources of the operation request. By using the submission timestamp and timeout threshold to provide precise tracking and proactive detection of request duration, infinite waiting caused by device latency is avoided. Time difference calculation and timeout determination quickly identify abnormal requests, and combined with clear error information delivery, enhance the transparency of fault handling. The thread synchronization mechanism ensures that worker threads respond quickly to timeout statuses, maintain workflow continuity, and optimize memory usage and long-term operational stability.

[0085] In one embodiment, the proxy thread invokes a callback function and wakes up the corresponding worker thread, including:

[0086] Before the proxy thread calls the callback function, the processing result of the operation request is written to the shared result buffer bound to the worker thread, and the result flag is marked as completed. The shared result buffer is a pre-allocated lock-free circular queue structure, and the worker thread corresponds to an independent buffer segment.

[0087] The proxy thread writes the worker thread identifier to its ready thread queue through an atomic operation and triggers the thread scheduler's wake-up logic.

[0088] The awakened worker thread checks the status of the flag bit in its result buffer;

[0089] In response to the flag indicating completion, the corresponding processing result structure is read and the buffer segment is cleared;

[0090] In response to an abnormal flag bit, a fault recovery process is triggered, which includes: logging, rolling back the status and resubmitting the request;

[0091] In response to a successful read result, the worker thread selects whether to proceed with subsequent processing logic based on the result status.

[0092] Specifically, after processing is complete or an exception occurs, the proxy thread writes the processing result structure of the operation request into the buffer segment of the target worker thread and marks the result flag as "completed" as a signal that the data writing is complete. The proxy thread writes the thread identifier of the target worker thread into the ready thread queue maintained by the scheduler through an atomic operation and explicitly triggers the thread scheduler to wake up the corresponding thread. After being woken up, the worker thread checks the status of the flag in its corresponding buffer segment. If the flag is "completed", it indicates that the proxy thread has successfully written the data. The worker thread then continues to read the processing result structure in that segment and actively clears the buffer segment after successful reading to release resources and avoid dirty data pollution. If the flag is not set correctly or an error code or other abnormal state occurs, the worker thread marks the request as abnormal and actively executes the fault recovery process. This process includes recording the fault log, rolling back part of the processing state, and resubmitting the request to the proxy thread or task queue to achieve automatic compensation or fault-tolerant retry of the request. After successfully reading the result structure, the worker thread will determine whether to proceed to the next step of the business process based on the result content, including logic such as business feedback, context cleanup, or resource reclamation. An efficient asynchronous communication mechanism is implemented between the proxy thread and the worker thread, significantly reducing the performance overhead of thread synchronization while maintaining data consistency. Furthermore, the robustness and availability of the system are enhanced through a coordinated mechanism of flag control and fault recovery procedures. The combination of lock-free structures and atomic operations effectively avoids thread blocking and lock waiting, making it suitable for high-concurrency distributed scenarios.

[0093] In one embodiment, such as Figure 2 As shown, a heterogeneous computing scheduling device based on proxy threads is provided, including: a thread creation module 610, a communication queue creation module 620, a task receiving module 630, a request submission module 640, an execution module 650, and a result return module 660, used for:

[0094] The thread creation module 610 is used to create worker thread groups and agent threads. The worker thread group contains at least one worker thread, which is used to process local data calculations and control the execution unit to process tasks.

[0095] The communication queue creation module 620 is used to establish a lock-free communication queue connection between the worker thread group and the agent thread.

[0096] The task receiving module 630 is used to respond to receiving a task to be processed by the execution unit, in which the worker thread submits the operation request of the task to be processed to the agent thread through a lock-free communication queue.

[0097] The request submission module 640 is used to respond to the successful submission of the operation request. If the work thread performs local data calculation, it will enter sleep mode if the local data calculation is completed.

[0098] The execution module 650 is used to respond to the agent thread receiving an operation request, in which case the agent thread submits the operation request to the execution unit and polls whether the pending task has been completed.

[0099] The result return module 660 is used to respond to the completion of the pending task. In this case, the agent thread wakes up the corresponding worker thread through the callback function and returns the processing result.

[0100] In one embodiment, the thread creation module 610 is used for:

[0101] By calling thread-safe data structure functions, shared memory is set up, and a lock-free communication queue is constructed in the shared memory to pass operation requests between the worker thread group and the agent thread.

[0102] Initialize the producer head pointer and consumer tail pointer of the lock-free communication queue;

[0103] Bind the consumer's tail pointer operation permissions to the agent thread, and assign the producer's head pointer operation permissions to the worker thread group.

[0104] In one embodiment, the request submission module 640 is used for:

[0105] The worker thread sets the request message structure based on the currently pending operation requests;

[0106] When an available enqueue slot is available in the lock-free communication queue, the worker thread writes the operation request into the specified slot of the lock-free communication queue through the lock-free mechanism and updates the producer head pointer.

[0107] In one embodiment, result returning module 660 is used for:

[0108] Based on the preset mapping relationship between operation handles and operation requests, the request message structure corresponding to the operation request is retrieved. The request message structure contains: a callback function pointer and worker thread context information, wherein the callback function pointer points to a preset callback function.

[0109] The processing result and the callback function pointer are encapsulated into a callback parameter structure, and the callback parameter structure is bound to the request message structure;

[0110] The callback function is invoked, which wakes up the corresponding worker thread based on the synchronization object in the worker thread context information.

[0111] In response to a worker thread being woken up, the corresponding processing result is obtained from the request message structure;

[0112] In response to the completion of the callback function, the proxy thread releases the resources corresponding to the operation request.

[0113] It should be understood that, although Figure 2 The steps in the device block diagram are shown sequentially as indicated by the arrows; however, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order requirement for the execution of these steps, and they can be executed in other orders. Furthermore, Figure 2 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0114] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the heterogeneous computing scheduling method based on proxy threads when running.

[0115] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0116] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in the embodiments of the heterogeneous computing scheduling method based on proxy threads described above.

[0117] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both, such as Figure 4 As shown, to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the above description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0118] The heterogeneous computing scheduling method based on proxy threads provided in this application has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A heterogeneous computing scheduling method based on proxy threads, characterized in that, include: Create a worker thread group and a proxy thread, wherein the worker thread group contains at least one worker thread, and the worker thread is used to process local data calculations and control the execution unit to process tasks. Establish a lock-free communication queue to connect the worker thread group and the proxy thread; In response to receiving a task to be processed by the execution unit, the worker thread submits the operation request of the task to be processed to the proxy thread through the lock-free communication queue; In response to the successful submission of the operation request, the worker thread performs local data calculation, and enters sleep mode if the local data calculation is completed. In response to the proxy thread receiving the operation request, the proxy thread submits the operation request to the execution unit and polls to see if the task to be processed has been completed. In response to the completion of the pending task, the proxy thread wakes up the corresponding worker thread through a callback function and returns the processing result.

2. The heterogeneous computing scheduling method based on proxy threads according to claim 1, characterized in that, The step of establishing a lock-free communication queue connection between the worker thread group and the proxy thread includes: By calling a thread-safe data structure function, shared memory is set up, and a lock-free communication queue is constructed in the shared memory for transmitting the operation request between the worker thread group and the proxy thread; Initialize the producer head pointer and consumer tail pointer of the lock-free communication queue; The operation permissions of the consumer tail pointer are bound to the agent thread, and the operation permissions of the producer head pointer are assigned to the worker thread group.

3. The heterogeneous computing scheduling method based on proxy threads according to claim 2, characterized in that, The response to the at least one worker thread submitting an operation request to the proxy thread through the lock-free communication queue includes: The worker thread sets the request message structure according to the operation request to be processed; In response to the availability of an enqueue slot in the lock-free communication queue, the worker thread writes the operation request into the specified slot of the lock-free communication queue using a lock-free mechanism and updates the producer head pointer.

4. The heterogeneous computing scheduling method based on proxy threads according to claim 2, characterized in that, In response to the agent thread receiving an operation request, the agent thread submits the operation request to the execution unit and polls whether the pending tasks have been completed, including: The proxy thread obtains the operation request through the consumer interface of the lock-free communication queue and updates the consumer tail pointer; The proxy thread submits the operation request to the execution unit and records the operation request status mapping; in response to the arrival of a preset time interval, the proxy thread queries whether the execution unit has completed the task to be processed.

5. A heterogeneous computing scheduling method based on proxy threads according to claim 3, characterized in that, In response to the completion of the pending task, the proxy thread wakes up the corresponding worker thread through a callback function and returns the processing result, including: Based on the preset mapping relationship between the operation handle and the operation request, the request message structure corresponding to the operation request is retrieved. The request message structure includes: a callback function pointer and worker thread context information, wherein the callback function pointer points to a preset callback function. The processing result and the callback function pointer are encapsulated into a callback parameter structure, and the callback parameter structure is bound to the request message structure; The callback function is invoked, and the callback function wakes up the corresponding worker thread based on the synchronization object in the worker thread context information; In response to the worker thread being awakened, the corresponding processing result is obtained from the request message structure; in response to the completion of the callback function, the proxy thread releases the resources corresponding to the operation request.

6. The heterogeneous computing scheduling method based on proxy threads according to claim 1, characterized in that, The method further includes: In response to the worker thread submitting the operation request, the worker thread performs the local data calculation to obtain the processed value; Upon completion of the local data calculation, the worker thread enters a sleep state. In response to the worker thread being awakened, the worker thread performs secondary processing on the processed value and the processed result.

7. The heterogeneous computing scheduling method based on proxy threads according to claim 5, characterized in that, The method further includes: The proxy thread submits the operation request to the execution unit and records the submission timestamp of the operation request. Write the submission timestamp into the data structure corresponding to the operation request; Set a corresponding timeout threshold for the operation request submitted to the execution unit; If the result of the query from the proxy thread to the execution unit is that the operation is not completed, then the time difference between the current system time and the submission timestamp in the data structure is calculated. If the time difference is greater than the timeout threshold and the operation request is not completed, the operation request is determined to be a timeout request. For the operation request corresponding to the timeout request, construct a processing result carrying error code and timeout description information, and encapsulate it into a callback parameter structure; The callback function corresponding to the operation request is invoked to wake up the worker thread bound to the operation request through the thread synchronization mechanism and pass the timeout processing result. In response to the worker thread recognizing a timeout state, the waiting process is terminated and a new operation request is received; In response to the completion of the callback function, the proxy thread releases the resources corresponding to the operation request.

8. A heterogeneous computing scheduling device based on proxy threads, characterized in that, The device includes: The thread creation module is used to create worker thread groups and agent threads. A worker thread group contains at least one worker thread, which is used to process local data calculations and control the execution unit to process tasks. A communication queue creation module is used to establish a lock-free communication queue connection between the worker thread group and the proxy thread; The task receiving module is used to respond to the receiving of a task to be processed by the execution unit. In this case, the worker thread submits the operation request of the task to be processed to the agent thread through a lock-free communication queue. The request submission module is used to respond to the successful submission of an operation request. In this case, the worker thread performs local data calculations, and if the local data calculations are completed, it enters a sleep state. The execution module is used to respond to the agent thread receiving an operation request. The agent thread submits the operation request to the execution unit and polls to see if the pending task has been completed. The result return module is used to respond to the completion of the pending task. In this case, the agent thread wakes up the corresponding worker thread through the callback function and returns the processing result.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Asynchronous IO calling method and device for service codes, storage medium and computer equipment

    CN122132127A

  • A processing method and system for asynchronous read IO of a Ceph OSD service end

    CN122470128A