Task processing apparatus, task processing method, electronic device, and storage medium
By integrating computing scheduling granules and functional granules through a multi-die solution, unified task scheduling is achieved, which solves the bottlenecks of traditional SoCs in terms of scalability and energy efficiency. It supports a variety of heterogeneous computing tasks, improves computing efficiency and flexibility, and is suitable for AI acceleration and high-performance computing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI BIREN TECH CO LTD
- Filing Date
- 2026-01-27
- Publication Date
- 2026-04-28
AI Technical Summary
When faced with the slowdown of Moore's Law and the demand for high-performance computing, traditional monolithic SoCs suffer from soaring costs, high design complexity, and limited functional scalability, making it difficult to efficiently support heterogeneous and diverse computing workloads.
Employing a multi-die approach, this solution integrates computation scheduling kernels and functional kernels into a single package using advanced packaging technology. This enables unified task scheduling, integrates computation scheduling kernels, functional kernels, memory, and interconnect structures, supports quantum computing management kernels and ray tracing core kernels, and provides efficient allocation of computing resources and collaborative computing.
It significantly improves adaptability in hybrid task scenarios on the edge, supports a variety of heterogeneous workloads, including scientific computing, AI training and inference, quantum computing, image rendering, etc., and solves the bottlenecks of traditional SoCs in scalability, energy efficiency and task scheduling fragmentation. It is suitable for AI acceleration, high-performance computing and data center scenarios.
Smart Images

Figure CN121579417B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to the field of artificial intelligence, specifically to a task processing apparatus, a task processing method, an electronic device, and a storage medium. Background Technology
[0002] A System-on-Chip (SoC) is an integrated circuit architecture that integrates multiple functional modules, such as processors, memory, peripheral interfaces, dedicated acceleration units, and interconnect structures, onto a single silicon chip. As Moore's Law gradually slows down, and applications such as artificial intelligence and high-performance computing place higher demands on computing power, bandwidth, and energy efficiency, traditional monolithic SoCs face challenges such as soaring costs, high design complexity, and limited functional scalability. Against this backdrop, how to construct a system architecture that can efficiently support heterogeneous and diverse computing workloads has become one of the important technical issues. Summary of the Invention
[0003] At least one embodiment of this disclosure provides a task processing apparatus including a computation scheduling core and a functional core integrated in a single package. The computation scheduling core includes a hardware subsystem and a computation core. The hardware subsystem is configured to receive and parse task information to obtain at least one task, and to assign the at least one task to at least one of the functional core and the computation core according to the task type. The computation core is configured to perform computational operations according to the tasks assigned by the hardware subsystem. The functional core includes at least one of a quantum computing management core and a ray tracing core core. The quantum computing management core is configured to control a quantum computing device to perform quantum computing operations according to the tasks assigned by the hardware subsystem, wherein the quantum computing device is connected to the quantum computing management core. The ray tracing core core is configured to perform ray tracing computational operations according to the tasks assigned by the hardware subsystem.
[0004] In a task processing apparatus provided in at least one embodiment of this disclosure, the computing scheduling kernel further includes a bus interface, through which the computing scheduling kernel receives task information from external devices.
[0005] In a task processing apparatus provided in at least one embodiment of this disclosure, the computing scheduling core further includes an inter-core interconnection interface, and the computing scheduling core is connected to the functional core through the inter-core interconnection interface.
[0006] In a task processing apparatus provided in at least one embodiment of this disclosure, a memory is also integrated within the single package, and the computing scheduling chip further includes a memory control interface, through which the computing scheduling chip is connected to the memory.
[0007] In a task processing apparatus provided in at least one embodiment of this disclosure, the hardware subsystem includes a command processor, a memory management unit, and a data transfer unit. The command processor is configured to receive and parse task information to obtain at least one task, and to allocate the at least one task to at least one of the functional core and the computing core according to the task type. The memory management unit is configured to manage access to the memory. The data transfer unit is configured to perform data transfer operations between the computing scheduling core and the memory.
[0008] In the task processing apparatus provided in at least one embodiment of this disclosure, the functional modules in the computing scheduling chip are interconnected through an on-chip interconnect structure.
[0009] In the task processing apparatus provided in at least one embodiment of this disclosure, the computing core includes at least one of a tensor computing core and a vector computing core.
[0010] In a task processing apparatus provided in at least one embodiment of the present disclosure, the functional core further includes an input / output core configured to enable interconnection between a plurality of the task processing apparatuses.
[0011] This disclosure provides at least one embodiment of a task processing method applied to a task processing apparatus provided in at least one embodiment of this disclosure, comprising: the hardware subsystem receiving and parsing task information to obtain at least one task, and allocating the at least one task to at least one of the functional core and the computing core according to the task type; in response to receiving a first type of task allocated by the hardware subsystem, the computing core performing a computation operation according to the first type of task; if the functional core includes a quantum computing management core, in response to receiving a second type of task allocated by the hardware subsystem, the quantum computing management core controlling a quantum computing device to perform a quantum computing operation according to the second type of task; if the functional core includes a ray tracing core core, in response to receiving a third type of task allocated by the hardware subsystem, the ray tracing core core performing a ray tracing computation operation according to the third type of task.
[0012] In a task processing method provided in at least one embodiment of this disclosure, the at least one task includes at least one of a first type of task, a second type of task, and a third type of task. The hardware subsystem allocates the at least one task to at least one of the computing cores of the functional core and the computing scheduling core according to the task type, including: the hardware subsystem allocating the first type of task to the computing core; the hardware subsystem allocating the second type of task to the quantum computing management core; or the hardware subsystem allocating the third type of task to the ray tracing core core.
[0013] The task processing method provided in at least one embodiment of this disclosure further includes: the computing core accessing the memory or input / output kernel according to the first type of task; the quantum computing management kernel accessing the memory or input / output kernel according to the second type of task; or the ray tracing core kernel accessing the memory or input / output kernel according to the third type of task.
[0014] In a task processing method provided in at least one embodiment of this disclosure, after the quantum computing management chip controls the quantum computing device to perform a quantum computing operation according to the second type of task, the task processing method further includes: in response to the completion of the quantum computing operation, the quantum computing management chip sends a second end signal to the hardware subsystem.
[0015] The task processing method provided in at least one embodiment of this disclosure further includes: in response to receiving the second termination signal, the hardware subsystem assigns the first type of task to the computing core, so that the computing core performs a computing operation based on the result of the quantum computing operation.
[0016] In at least one embodiment of the task processing method provided in this disclosure, after the computing core performs a computing operation according to the first type of task, the task processing method further includes: in response to the completion of the computing operation, the computing core sends a first end signal to the hardware subsystem.
[0017] The task processing method provided in at least one embodiment of this disclosure further includes: in response to receiving the first end signal, the hardware subsystem assigns the third type of task to the ray tracing core kernel, so that the ray tracing core kernel performs ray tracing calculation operations based on the result of the calculation operation.
[0018] At least one embodiment of this disclosure provides an electronic device, including a task processing device provided in at least one embodiment of this disclosure.
[0019] At least one embodiment of this disclosure provides an electronic device, including: at least one processor; at least one memory including one or more computer program modules; wherein the one or more computer program modules are stored in the at least one memory and configured to be executed by the at least one processor, and the one or more computer program modules are used to implement the task processing method provided by at least one embodiment of this disclosure.
[0020] At least one embodiment of this disclosure provides a non-transitory computer-readable storage medium storing computer-readable instructions thereon, wherein the computer-readable instructions, when executed by at least one processor, perform the task processing method provided in at least one embodiment of this disclosure.
[0021] The task processing apparatus, task processing method, electronic device, and storage medium provided in at least one embodiment of this disclosure are based on a multi-die approach. Through advanced packaging technologies (such as 2.5D / 3D packaging or silicon interposers), multiple functionally distinct chips are integrated into the same package (e.g., on the same substrate or interposer). These multiple chips collaborate in computation, presenting themselves as a single, logically unified, high-performance chip. This device significantly improves its adaptability in edge-side mixed-task scenarios through a unified task scheduling and allocation mechanism. This allows a single computing chip (or a single computing card) to efficiently support various heterogeneous workloads, covering scientific computing, AI training and inference, quantum computing, image rendering, and other computing tasks. It fully leverages the synergistic advantages of heterogeneous chips in terms of computing power, energy efficiency, and flexibility. Therefore, this task processing device effectively solves the bottlenecks of traditional monolithic SoCs in terms of scalability, energy efficiency, and task scheduling fragmentation, making it suitable for applications with high requirements for computing density and scheduling efficiency, such as AI acceleration, high-performance computing, and data centers. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.
[0023] Figure 1 This is a schematic block diagram of a task processing apparatus provided for at least one embodiment of the present disclosure.
[0024] Figure 2 This is a schematic block diagram of another task processing apparatus provided for at least one embodiment of the present disclosure.
[0025] Figure 3 A flowchart illustrating a task processing method provided in at least one embodiment of this disclosure.
[0026] Figure 4 A flowchart illustrating another task processing method provided in at least one embodiment of this disclosure.
[0027] Figure 5A A flowchart of yet another task processing method provided in at least one embodiment of this disclosure.
[0028] Figure 5B A flowchart illustrating yet another task processing method provided in at least one embodiment of this disclosure.
[0029] Figure 6 This is a schematic block diagram of an electronic device provided for at least one embodiment of the present disclosure.
[0030] Figure 7This is a schematic block diagram of another electronic device provided for at least one embodiment of the present disclosure.
[0031] Figure 8 This is a schematic block diagram of another electronic device provided for at least one embodiment of the present disclosure.
[0032] Figure 9 This is a schematic block diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0034] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as “comprising” or “including” mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as “connected” or “linked” are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as “upper,” “lower,” “left,” and “right” are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.
[0035] The present disclosure will now be described through several specific embodiments. To keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and known components may be omitted. When any component of an embodiment of the present disclosure appears in more than one drawing, the component is represented by the same or similar reference numerals in each drawing.
[0036] Chiplet technology is a design approach that breaks down the various functional modules (such as computing, storage, or input / output) integrated in a traditional SoC into multiple independently manufactured, reusable small dies. These dies are given specific functional roles in the chipplet architecture and are called "chips".
[0037] Some chips typically employ monolithic integration or architectures based on simple chip assembly. In such designs, various computing units (e.g., Central Processing Unit (CPU), Graphics Processing Unit (GPU), Neural Network Processing Unit (NPU)) are usually equipped with their own independent task schedulers, lacking a unified global scheduling mechanism. This results in low overall system resource utilization and uneven distribution of computational load. Furthermore, some chip architecture optimization schemes primarily focus on physical layer interconnect standards, lacking deep collaborative design at the system level, such as computational function integration and cross-chip task collaborative scheduling. In scenarios requiring multiple computing tasks, such as scientific computing, artificial intelligence (AI) training and inference, quantum computing, and image rendering, the demands on computing units vary significantly, making it difficult to efficiently complete all tasks on the edge using a single, fixed computing unit.
[0038] At least one embodiment of this disclosure provides a task processing apparatus, including a computation scheduling core and a functional core integrated in a single package. The computation scheduling core includes a hardware subsystem and a computation core. The hardware subsystem is configured to receive and parse task information to obtain at least one task, and to assign the at least one task to at least one of the functional core and the computation core according to the task type. The computation core is configured to perform computational operations according to the tasks assigned by the hardware subsystem. The functional core includes at least one of a quantum computing management core and a ray tracing core core. The quantum computing management core is configured to control a quantum computing device to perform quantum computing operations according to the tasks assigned by the hardware subsystem, wherein the quantum computing device is connected to the quantum computing management core. The ray tracing core core is configured to perform ray tracing computational operations according to the tasks assigned by the hardware subsystem.
[0039] The task processing device provided in at least one embodiment of this disclosure is based on a multi-die approach. Through advanced packaging technologies (such as 2.5D / 3D packaging or silicon interposers), multiple functionally distinct chips are integrated into a single package (e.g., on the same substrate or interposer). These chips collaborate in computation, presenting themselves as a single, logically unified, high-performance chip. This device significantly improves its adaptability to mixed-task scenarios on the edge through a unified task scheduling and allocation mechanism. This allows a single computing chip (or a single computing card) to efficiently support various heterogeneous workloads, covering scientific computing, AI training and inference, quantum computing, image rendering, and other computing tasks. It fully leverages the synergistic advantages of heterogeneous chips in terms of computing power, energy efficiency, and flexibility. Therefore, this task processing device effectively solves the bottlenecks of traditional monolithic SoCs in terms of scalability, energy efficiency, and task scheduling fragmentation, making it suitable for applications with high requirements for computing density and scheduling efficiency, such as AI acceleration, high-performance computing, and data centers.
[0040] It should be noted that the "heterogeneous core" mentioned above refers to two or more cores in a multi-core integrated system that differ in dimensions such as functional type. For example, in terms of functional type, a single core can be configured to implement one or more of the following computing units, including but not limited to: a central processing unit (CPU), a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), an accelerated processing unit (APU), a neural network processing unit (NPU), a quantum processing unit (QPU), a ray tracing core (RTcore), an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA), etc.
[0041] Figure 1 This is a schematic block diagram of a task processing apparatus provided for at least one embodiment of the present disclosure.
[0042] For example, such as Figure 1 As shown, the task processing apparatus 100 provided in at least one embodiment of the present disclosure includes a computing scheduling core 110 and a functional core 120 integrated in a single package.
[0043] For example, the computation scheduling core 110 includes a hardware subsystem 111 and a computation core 112. The hardware subsystem 111 is configured to receive and parse task information to obtain at least one task, and to assign at least one task to at least one of the functional core 120 and the computation core 112 according to the task type. The computation core 112 is configured to perform computation operations according to the task assigned by the hardware subsystem 111.
[0044] For example, a hardware subsystem is a collection of physical hardware within a task processing unit used to perform functions such as information reception, task parsing and scheduling, and data transmission. The hardware subsystem is configured to perform scheduling tasks outside of the core computing, thereby offloading non-computationally intensive system functions such as task management and communication control from the computing core.
[0045] For example, the computation core may include at least one of a tensor core (Tcore) and a vector core (Vcore).
[0046] For example, the vector computing core can execute Single Instruction Multiple Threads (SIMT) instructions to perform corresponding vector operations. SIMT instructions are common graphics processing unit (GPU) programming instructions widely used in massively parallel computing tasks such as graphics rendering and AI computation. Vector operations include, but are not limited to, floating-point operations, fixed-point operations, and logical operations. The vector computing core's pipelined architecture supports massively parallel computing tasks. The pipelined architecture includes multiple processes, such as instruction fetching, instruction scheduling, decoding, operand fetching, operation execution, and result writing back.
[0047] For example, the Tensor Computation Core can perform tensor-related computations, such as mixed-precision matrix multiply-accumulate (MMA) operations. Through massively parallel MMA operations, the Tensor Computation Core significantly improves the execution efficiency of tensor-intensive workloads.
[0048] For example, functional core 120 includes at least one of a quantum computing management core and a ray tracing core core. The quantum computing management core is configured to control a quantum computing device to perform quantum computing operations according to tasks assigned by the hardware subsystem, wherein the quantum computing device is connected to the quantum computing management core; the ray tracing core core is configured to perform ray tracing computing operations according to tasks assigned by the hardware subsystem.
[0049] In at least one embodiment of this disclosure, a quantum computing device refers to a physical hardware device for performing computational tasks based on the principles of quantum mechanics. One example includes the quantum processing unit (QPU) described below.
[0050] A quantum processing unit (QPU) is a hardware processor dedicated to performing quantum computing tasks and serves as the computing engine of a quantum computer. Its working principle is based on quantum mechanics, relying on fundamental laws of quantum mechanics (such as quantum superposition, entanglement, and tunneling) to achieve high-speed mathematical and logical operations. Unlike classical computers that use binary bits, QPUs use qubits as the basic unit of information and process information by manipulating the quantum states of qubits (such as creating superposition states and generating entangled states). Specifically, QPUs can process information by manipulating the quantum states of qubits (such as superposition and entanglement properties). As the physical device that carries out quantum algorithms and performs qubit manipulation and measurement, the QPU is essentially a dedicated chip or device for performing quantum computing, capable of solving complex problems that classical computers struggle to handle efficiently, such as large integer factorization, quantum many-body system simulation, and combinatorial optimization problems.
[0051] Quantum computing devices can be managed collaboratively by a controller, which can perform at least one of the following functions: scheduling and controlling instruction resources involved in quantum computing tasks, including compiling quantum gate sequences, mapping logic to physical qubits, and coordinating the timing of quantum operations; and performing data preparation operations, including initializing qubit states, generating quantum-encoded representations of classical input data, and configuring measurement bases, among other preprocessing steps. The controller can be configured to manage a single quantum computing device or a computing cluster comprising multiple quantum computing devices, thereby supporting distributed or parallel quantum computing architectures.
[0052] In at least one embodiment of this disclosure, the quantum computing management chip is a dedicated control module that integrates all or part of the functions of the aforementioned controller into a single chip. This quantum computing management chip can communicate with quantum computing devices through high-bandwidth, low-latency interconnect interfaces (e.g., including but not limited to silicon photonic interconnects, advanced packaged embedded interconnects, etc.), thereby realizing a modular, scalable, and easily integrated quantum control system architecture.
[0053] The Ray Tracing Core (RT core) is a hardware acceleration unit used to efficiently perform critical computational tasks in ray tracing algorithms, such as ray-tracing-related intensive geometric calculations. For example, the RT core can accelerate traversal of the Bounding Volume Hierarchy (BVH) to quickly filter scene regions where rays may intersect; it can accelerate ray-triangle intersection testing, determining whether a ray intersects a geometric surface after entering the target bounding volume; and it can assist in managing ray generation and the scheduling of multi-level bounce paths, including the emission and traversal control of reflected rays, refracted rays, and shadow rays. These operations can be processed in parallel by dedicated circuitry built into the RT Core, significantly reducing the time-consuming geometric computation overhead in ray tracing. This hardware acceleration mechanism is particularly suitable for applications with stringent real-time rendering performance requirements, including but not limited to: implementing dynamic global illumination, realistic reflections, and regional shadow effects in video games; high-fidelity lighting simulation in building information modeling and visualization software; real-time physically-based optical simulation in industrial digital twin systems; and cinematic pre-visualization and virtual production systems built on real-time rendering engines. Through the hardware acceleration mechanism provided by RT Core, a real-time frame rate that meets the needs of human-computer interaction can be achieved while maintaining a high degree of visual realism.
[0054] In the task processing device provided in at least one embodiment of the present disclosure, a high-performance chip architecture for multi-chip integration is implemented, which integrates multi-chip collaborative computing capabilities and a unified task scheduling mechanism, significantly improving its adaptability in edge-side mixed task scenarios, enabling a single computing chip (or a single computing card) to efficiently support a variety of heterogeneous workloads, covering a variety of computing tasks such as scientific computing, AI training and inference, quantum computing, and image rendering.
[0055] In the task processing apparatus provided in at least one embodiment of this disclosure, the computing scheduling core may further include a bus interface, through which the computing scheduling core receives task information from external devices.
[0056] In some examples, the bus interface may adopt the Peripheral Component Interconnect Express (PCIe) interface standard. The PCIe interface features high bandwidth, low latency, and good compatibility, enabling seamless integration with mainstream host systems. In other examples, other high-speed interconnect protocols such as Compute Express Link (CXL) and Advanced eXtensible Interface (AXI) may be used as the physical or logical implementation of the bus interface; this disclosure does not impose limitations on these implementations.
[0057] For example, an external device refers to a device located outside the task processing unit, distinct from the individual cores integrated within it. External devices can include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), a smart network interface card (SmartNIC), an FPGA accelerator card, or other coprocessors. For instance, when the external device is a CPU, the host operating system or user-mode driver can encapsulate the computational task to be processed and distribute it to the task processing unit via the PCIe bus, and then send it to the computation scheduling core through the bus interface; when the external device is a smart network interface card, network data packets can be directly parsed and transformed into computational tasks, which are then scheduled and executed in real-time by the computation scheduling core.
[0058] By integrating a bus interface into the computing scheduling core, the task processing device provided in at least one embodiment of the present disclosure not only achieves flexible interconnection with a variety of external devices, but also provides a hardware foundation for task distribution, resource coordination and performance optimization in heterogeneous computing environments.
[0059] In the task processing apparatus provided in at least one embodiment of this disclosure, the computing scheduling core further includes an inter-core interconnection interface, through which the computing scheduling core is connected to the functional core.
[0060] For example, a computing scheduling kernel may include one or more kernel interconnect interfaces for establishing high-speed, low-latency kernel communication connections with one or more functional kernels.
[0061] In some examples, the inter-chip interconnect interface can adopt the Universal Chiplet Interconnect Express (UCIe) standard. This interface can include a UCIe-compliant controller and a UCIe PHY, supporting a complete stack of functionality across the protocol layer, adapter layer, and physical layer. The UCIe interface features high bandwidth density, scalable topology, and native support for various upper-layer protocols (such as PCIe and CXL), effectively meeting the interconnection needs of heterogeneous chip integration scenarios.
[0062] In other examples, the inter-core interconnect interface may also employ other core interconnect technologies, such as bundle of wires (BoW), advanced interface bus (AIB), or custom high-speed serial / parallel interfaces, which can enable reliable and efficient data exchange between computing scheduling cores and functional cores. This disclosure does not limit this.
[0063] By integrating the inter-core interconnection interface in the computing scheduling core, the task processing device provided in at least one embodiment of the present disclosure not only decouples the computing scheduling logic from the dedicated functional hardware, but also significantly improves the system integration, maintainability and scalability, providing a flexible and efficient scheduling foundation for heterogeneous computing platforms based on core architecture.
[0064] In at least one embodiment of the task processing apparatus provided in this disclosure, a memory is also integrated within a single package. That is, the task processing apparatus includes a computing scheduling core, a functional core, and a memory integrated within a single package.
[0065] Correspondingly, the computation scheduling kernel may also include a memory controller interface (MCI), through which the kernel connects to the memory. The kernel can directly access the memory via the MCI with high throughput and low latency to read input data and write intermediate computation results. This memory can serve as the task processing device's local main memory or cache pool, significantly reducing data access latency and improving overall throughput performance.
[0066] In some examples, the memory is high-bandwidth memory (HBM), such as stacked DRAM memory like HBM2, HBM2E, or HBM3. Correspondingly, the memory control interface is an HBM control interface, including an HBM-compliant controller and a physical layer (HBM PHY). The HBM controller handles high-level functions such as memory access requests, address mapping, command scheduling, refresh management, and error correction; the HBM PHY provides high-speed electrical connectivity to the HBM stack, supporting multi-channel parallel data transmission, clock synchronization, and signal integrity optimization.
[0067] In other examples, the memory may also be other types of high-density, high-bandwidth storage media, such as phase-change memory (PCM), resistive random access memory (RRAM), etc., which are not limited in this disclosure.
[0068] By integrating the memory control interface within the computation scheduling core, the task processing apparatus provided in at least one embodiment of this disclosure can avoid frequent access to remote host memory, reduce system bus congestion, and support rapid switching of task contexts. Especially when processing large-scale parallel tasks (such as AI computation), it can significantly improve task scheduling efficiency.
[0069] In at least one embodiment of the task processing apparatus provided in this disclosure, the hardware subsystem may include a command processor, a memory management unit, and a data transfer unit.
[0070] For example, the command processor is configured to receive and parse task information to obtain at least one task, and assign the at least one task to at least one of the functional cores and computing cores according to the task type; the memory management unit is configured to manage access to memory; and the data transport unit is configured to perform data transport operations between the computing scheduling core and memory. The memory management unit and the data transport unit work together to manage data interaction between the computing scheduling core and memory.
[0071] For example, the data transfer unit is configured to perform efficient data transfer operations between the computation scheduling kernel and memory, including but not limited to reading data required for a task from memory and writing task execution results to memory; this disclosure does not limit this. In some examples, the data transfer unit can be a Direct Memory Access (DMA) controller, which can autonomously complete data transfer without processor intervention, significantly reducing scheduling overhead and freeing up computing resources.
[0072] For example, the memory management unit can be configured to perform address translation, access control, memory isolation, and virtualization support for memory access. In some examples, the memory management unit can be the System Memory Management Unit (SMMU). All DMA request addresses undergo address translation and access control checks by the SMMU, thereby ensuring the security and consistency of memory access while achieving high-performance data transfer.
[0073] By integrating a memory management unit and a data transport unit into the hardware subsystem, the task processing apparatus provided in at least one embodiment of this disclosure not only enables secure and efficient access to on-chip memory, but also provides a reliable data path for heterogeneous task processing, achieving low-latency, high-throughput task scheduling and execution.
[0074] In the task processing apparatus provided in at least one embodiment of this disclosure, the functional modules in the computing scheduling core are interconnected through an on-chip interconnect structure. For example, the computing scheduling core integrates an on-chip interconnect structure to connect the various functional modules within the core (such as the computing core, hardware subsystem, bus interface, memory control interface, and inter-core interconnect interface mentioned above) to achieve efficient and low-latency data communication and control signal transmission.
[0075] For example, an on-chip interconnect architecture refers to an interconnect architecture integrated within a single chip to enable data communication and signal transmission between multiple functional modules within that chip. In some examples, the on-chip interconnect architecture is a Network-on-Chip (NoC). NoC provides high-bandwidth, scalable point-to-point or broadcast communication capabilities, suitable for high-concurrency task scheduling scenarios. NoC may include two layers: High-Bandwidth Fabric (HBF) and Low-Bandwidth Fabric (LBF). In other examples, the on-chip interconnect architecture can be a crossbar, a ring bus, etc., and this disclosure does not limit this aspect.
[0076] By integrating an on-chip interconnect structure, the computing scheduling chip provided in at least one embodiment of this disclosure achieves tight coupling and collaboration of internal functional modules, effectively reducing communication overhead between modules, thereby significantly improving task processing efficiency and response capability.
[0077] In the task processing apparatus provided in at least one embodiment of this disclosure, in addition to the quantum computing management core and the ray tracing core core, one example of a functional core may also be an input / output (I / O) core, which is configured to enable interconnection between multiple task processing devices.
[0078] For example, input / output chips can provide high-speed, scalable external interconnect interfaces for connecting this task processing device to other similar devices (e.g., other task processing devices deployed on the same server board, accelerator card cluster, or multi-chip module). Through this interconnect mechanism, multiple task processing devices can collaboratively process large-scale distributed tasks, improving system scalability.
[0079] It should be noted that functional cores are not limited to the aforementioned quantum computing management cores, ray tracing cores, and input / output cores; they can also be any other core used to perform specialized or general-purpose computing tasks. In other words, specifically, functional cores can integrate different types of processing logic according to application scenario requirements to achieve diverse computing functions. For example, examples of functional cores include, but are not limited to, signal processing cores, general-purpose computing cores, and neural network computing cores. All types of functional cores can connect to the computing scheduling core through inter-core interface to complete their respective computing tasks under the same task scheduling mechanism.
[0080] Figure 2 This is a schematic block diagram of another task processing apparatus provided for at least one embodiment of the present disclosure.
[0081] For example, such as Figure 2 As shown, at least one embodiment of the present disclosure provides a task processing apparatus including a computing scheduling core, a memory, a quantum computing management core, a ray tracing core, an input / output core, and other cores integrated in a single package.
[0082] For example, such as Figure 2 As shown, the computation scheduling core includes a hardware subsystem, computation cores (tensor computation core and vector computation core), interconnect interfaces between cores, memory control interfaces, and bus interfaces. These modules are interconnected via an on-chip interconnect structure. Figure 2(Shaded in the middle) are interconnected. The hardware subsystem includes a command processor, a memory management unit, and a data transfer unit. Communication with external devices is achieved through a bus interface, communication with quantum computing devices is achieved through quantum computing management chips, and communication with other task processing devices is achieved through input / output chips.
[0083] In the task processing device provided in at least one embodiment of the present disclosure, a high-performance chip architecture for multi-chip integration is implemented, which integrates multi-chip collaborative computing capabilities and a unified task scheduling mechanism, significantly improving its adaptability in edge-side mixed task scenarios, enabling a single computing chip (or a single computing card) to efficiently support a variety of heterogeneous workloads, covering a variety of computing tasks such as scientific computing, AI training and inference, quantum computing, and image rendering.
[0084] This disclosure provides at least one embodiment of a task processing method, applied to the task processing apparatus provided in the at least one embodiment described above.
[0085] Figure 3 A flowchart illustrating a task processing method provided in at least one embodiment of this disclosure.
[0086] like Figure 3 As shown, the task processing method provided in at least one embodiment of this disclosure may include steps S101 to S104.
[0087] Step S101: The hardware subsystem receives and parses the task information to obtain at least one task, and assigns at least one task to at least one of the functional cores and computing cores according to the task type.
[0088] For example, the hardware subsystem can receive task information from external devices (such as the CPU), one example of which is an executable program. For instance, the external device can compile and merge code from different programming languages (such as those for GPUs / QPUs / RT Cores) into a single executable program based on a unified heterogeneous programming model. This allows it to simultaneously carry task information for computing cores, quantum computing management chips, and ray tracing core chips. The external device can then send the executable program to a bus port, where it is subsequently transmitted to the hardware subsystem via on-chip interconnects.
[0089] For example, in step S101, the hardware subsystem parses the task information into tasks for different execution units (e.g., computing cores, quantum computing management chips, ray tracing core chips), and sends the parsed tasks to the corresponding execution units according to the task type. For example, the parsed tasks may include at least one of a first type of task, a second type of task, and a third type of task.
[0090] For example, the task processing method provided in at least one embodiment of this disclosure can be applied to fields such as speech processing, image processing, text processing, video processing, virtual reality, and scientific computing.
[0091] For example, in the field of speech processing, the first type of task can be feature extraction, speech enhancement, speech recognition, etc. For example, in the field of image processing, the first type of task can be feature extraction, image segmentation, object detection, style transfer, etc. For example, in the field of text processing, the first type of task can be text classification, sentiment analysis, text generation, question answering systems, etc. For example, in the field of video processing, the first type of task can be optical flow estimation, object tracking, video enhancement, scene understanding, action recognition, etc. For example, in the field of virtual reality, the first type of task can be gesture tracking, eye tracking, facial expression synchronization, 3D reconstruction, spatial localization, virtual-real fusion rendering, etc. For example, in the field of scientific computing, the first type of task can be accelerated molecular dynamics simulations, climate model inference, solving physical information neural networks, etc. Of course, this disclosure is not limited to these; the task processing methods provided by at least one embodiment of this disclosure can also be applied to other application scenarios or fields, which will not be elaborated here.
[0092] In some examples, the first type of task is, for example, an AI computing task (such as an AI inference task, an AI decoding task, tensor operations, vector operations, etc., which are not limited in this disclosure embodiment), the second type of task is, for example, a quantum computing task, and the third type of task is, for example, a ray tracing computing task. Of course, the parsed task can have many more types, which are not limited in this disclosure embodiment.
[0093] For example, the hardware subsystem can assign the first type of task to the computing core, the second type of task to the quantum computing management core, and the third type of task to the ray tracing core core.
[0094] For example, the hardware subsystem can maintain a task queue to cache tasks to be processed, and retrieve tasks from the task queue in sequence according to a preset scheduling strategy (such as first-in-first-out, priority sorting or dependency constraints), and send the tasks to the corresponding functional cores or computing cores, thereby ensuring the orderliness, determinism and resource coordination of task execution.
[0095] Step S102: In response to receiving a first type of task assigned by the hardware subsystem, the computing core performs a computing operation according to the first type of task.
[0096] For example, in step S102, the first type of task can be further subdivided into a first type of subtask and a second type of subtask, which are respectively assigned to the tensor computation core and the vector computation core for processing. The first type of subtask is, for example, a matrix multiplication and addition task, and the second type of subtask is, for example, a vector computation task. After completing the computation operation, the computation core can send a termination signal to the hardware subsystem, notifying the hardware subsystem that the task has been completed.
[0097] Step S103: In the case where the functional core includes a quantum computing management core, in response to receiving a second type of task assigned by the hardware subsystem, the quantum computing management core controls the quantum computing device to perform quantum computing operations according to the second type of task.
[0098] For example, in step S103, the quantum computing management chip can send a control signal to the quantum computing device according to the received quantum computing task, thereby driving the quantum computing device to execute the quantum computing operation corresponding to the task. The control signal may include instructions for initializing the state of qubits, applying a specific quantum gate sequence, triggering readout operations, etc., thereby realizing coordinated control of the entire quantum computing process. After completing the quantum computing operation, the quantum computing management chip can send a termination signal to the hardware subsystem to notify the hardware subsystem that the task has been completed.
[0099] Step S104: In the case where the functional core includes a ray tracing core, in response to receiving a third type of task assigned by the hardware subsystem, the ray tracing core performs ray tracing calculation operations according to the third type of task.
[0100] For example, in step S104, the ray tracing core can perform the corresponding ray tracing calculation operation based on the received ray tracing calculation task. After completing the ray tracing calculation operation, the ray tracing core can send a completion signal to the hardware subsystem to notify the hardware subsystem that the task has been completed.
[0101] In the task processing method provided in at least one embodiment of this disclosure, if there is a data read / write requirement in the first type of task, the computing core can access the memory (e.g., HBM) to read and write data from the memory according to the first type of task; the computing core can also access the input / output kernel to receive data from other task processing devices or write data to other task processing devices through the input / output kernel; similarly, if there is a data read / write requirement in the second type of task, the quantum computing management kernel can access the memory or the input / output kernel according to the second type of task; if there is a data read / write requirement in the third type of task, the ray tracing core kernel can access the memory or the input / output kernel according to the third type of task.
[0102] Figure 4A flowchart illustrating another task processing method provided in at least one embodiment of this disclosure.
[0103] like Figure 4 As shown, the task processing method provided in at least one embodiment of this disclosure may include steps S201 to S203. Figure 4 An example of a task processing device performing AI computing tasks.
[0104] Step S201: The hardware subsystem receives and parses the task information to obtain at least one task, and assigns the first type of task to the computing core.
[0105] Step S202: In response to receiving a first type of task assigned by the hardware subsystem, the computing core performs a computing operation according to the first type of task.
[0106] Step S203: In response to the completion of the computation operation, the computing core sends a first termination signal to the hardware subsystem.
[0107] In step S201, the hardware subsystem parses and obtains the first type of task (AI computing task) and sends the task to the computing core. In step S202, the computing core executes the corresponding computing operation according to the first type of task. If the task requires data read / write operations, the computing core can access the memory or input / output chips to perform data read / write operations. In step S203, after completing the computing operation, the computing core can send a first end signal to the hardware subsystem to notify the hardware subsystem that the task has been completed.
[0108] Figure 5A A flowchart of yet another task processing method provided in at least one embodiment of this disclosure.
[0109] like Figure 5A As shown, the task processing method provided in at least one embodiment of this disclosure may include steps S301 to S304. Figure 5A An example of a task processing device performing quantum computing tasks + AI computing tasks.
[0110] Step S301: The hardware subsystem receives and parses the task information to obtain at least one task, and assigns the second type of task to the quantum computing management chip.
[0111] Step S302: In response to receiving a second type of task assigned by the hardware subsystem, the quantum computing management chip controls the quantum computing device to perform quantum computing operations according to the second type of task.
[0112] Step S303: In response to the completion of the quantum computing operation, the quantum computing management chip sends a second termination signal to the hardware subsystem.
[0113] Step S304: In response to receiving the second end signal, the hardware subsystem assigns the first type of task to the computing core so that the computing core performs computing operations based on the results of the quantum computing operations.
[0114] In step S301, the hardware subsystem parses and obtains a first type of task (AI computing task) and a second type of task (quantum computing task). Based on task analysis, it determines that the quantum computing task needs to be executed first, followed by AI decoding of the quantum computing result. Therefore, the hardware subsystem sends the second type of task to the quantum computing management chip and temporarily stores the first type of task in the task queue. In step S302, the quantum computing management chip controls the quantum computing device to execute the corresponding quantum computing operation based on the second type of task. If the task requires data read / write operations, the quantum computing management chip can access the memory or input / output chip to perform these operations. In step S303, after the quantum computing device completes the quantum computing operation, the quantum computing management chip can send a second end signal to the hardware subsystem, notifying it that the task has been completed. In step S304, upon receiving the second end signal, the hardware subsystem retrieves the first type of task from the task queue and sends it to the computing core, enabling the computing core to execute AI computing operations (e.g., AI decoding) based on the result of the quantum computing operation. If the task requires data read / write operations, the computing core can access memory or input / output chips to perform these operations. After completing the AI computation, the computing core can send a first termination signal to the hardware subsystem, notifying it that the task has been completed.
[0115] Figure 5B A flowchart illustrating yet another task processing method provided in at least one embodiment of this disclosure.
[0116] like Figure 5B As shown, the task processing method provided in at least one embodiment of this disclosure may include steps S401 to S406.
[0117] Step S401: The hardware subsystem receives and parses the task information to obtain at least one task, and assigns the second type of task to the quantum computing management chip.
[0118] Step S402: In response to receiving a second type of task assigned by the hardware subsystem, the quantum computing management chip controls the quantum computing device to perform quantum computing operations according to the second type of task.
[0119] Step S403: In response to the completion of the quantum computing operation, the quantum computing management chip sends a second termination signal to the hardware subsystem.
[0120] Step S404: In response to receiving the second end signal, the hardware subsystem assigns the first type of task to the computing core so that the computing core performs computing operations based on the results of the quantum computing operations.
[0121] Step S405: In response to the completion of the computation operation, the computing core sends a first termination signal to the hardware subsystem.
[0122] Step S406: In response to receiving the first end signal, the hardware subsystem assigns the third type of task to the ray tracing core kernel so that the ray tracing core kernel performs ray tracing computation operations based on the results of the computation operations.
[0123] In step S401, the hardware subsystem parses and obtains three types of tasks: a first type of task (AI computing task), a second type of task (quantum computing task), and a third type of task (ray tracing computing task). Based on the task analysis, it is determined that the quantum computing task must be executed first, followed by AI decoding of the quantum computing result, and then ray tracing computation is performed on the AI decoding result (or possibly the result obtained from running an AI model, all calculated by the computing core). Therefore, the hardware subsystem sends the second type of task to the quantum computing management chip and temporarily stores the first and third types of tasks in the task queue. In step S402, the quantum computing management chip controls the quantum computing device to execute the corresponding quantum computing operation according to the second type of task. If the task requires data read / write operations, the quantum computing management chip can access the memory or input / output chip to perform data read / write operations. In step S403, after the quantum computing device completes the quantum computing operation, the quantum computing management chip can send a second end signal to the hardware subsystem, notifying the hardware subsystem that the task has been completed. In step S404, after receiving the second end signal, the hardware subsystem retrieves a first-type task from the task queue and sends it to the computing core, enabling the computing core to perform AI computation operations (e.g., AI decoding operations) based on the results of the quantum computing operations. If the task requires data read / write operations, the computing core can access memory or input / output particles to perform these operations. In step S405, after completing the computation operation, the computing core can send a first end signal to the hardware subsystem, notifying it that the task has been completed. In step S406, after receiving the first end signal, the hardware subsystem retrieves a third-type task from the task queue and sends it to the ray tracing core particle, enabling the ray tracing core particle to perform ray tracing computation operations based on the results of the AI computing operations. If the task requires data read / write operations, the ray tracing core particle can access memory or input / output particles to perform these operations. After completing the ray tracing computation operation, the ray tracing core particle can send a third end signal to the hardware subsystem, notifying it that the task has been completed.
[0124] For example, before step S406, the hardware subsystem can further allocate the AI inference task to the computing core, so that the computing core can perform the AI inference operation; after the computing core has completed the above-mentioned AI decoding operation and AI inference operation, it sends a first end signal to the hardware subsystem. Correspondingly, in step S406, the ray tracing core can perform ray tracing calculation operation based on both the result of the AI decoding operation and the result of the AI inference operation.
[0125] Figure 5B This example illustrates a task processing device that performs quantum computing tasks, AI computing tasks, and ray tracing computing tasks. For instance, building or running a complex simulation of a "future world model" requires the synergy of AI computing, quantum computing, and ray tracing computing. For example, a computing core (equipped with a Large Language Model (LLM) or Vision-Language Model (VLM)) provides foundational intelligent reasoning capabilities for semantic understanding, environment generation, decision planning, or user interaction; quantum computing devices complete sub-tasks requiring quantum-enhanced modeling (e.g., quantized state-space search, quantum probabilistic reasoning, or quantum physics engine simulation); and ray tracing core particles enable real-time 3D scene rendering and visualization, thus presenting a dynamic view of the world model using ray tracing.
[0126] In the task processing method provided in at least one embodiment of the present disclosure, the three types of computing tasks are coordinated by a unified task scheduling module (hardware subsystem) and distributed to the corresponding functional chips or computing cores, so that the task processing device under a single package can support the next generation of hybrid workloads that integrate AI, graphics and quantum computing from end to end.
[0127] It should also be noted that the execution order of the steps of the task processing method is not limited in the various embodiments of this disclosure. Although the execution process of each step has been described in a specific order above, this does not constitute a limitation on the embodiments of this disclosure. The steps in the task processing method can be executed serially or in parallel, depending on actual needs.
[0128] For example, compared to the above description, the task processing method provided in at least one embodiment of this disclosure may include more or fewer steps, and the embodiments of this disclosure do not limit this.
[0129] Figure 6 This is a schematic block diagram of an electronic device provided for at least one embodiment of the present disclosure.
[0130] For example, such as Figure 6As shown, the electronic device 600 provided in at least one embodiment of this disclosure includes a task processing device 601. This task processing device can be the task processing device provided in the at least one embodiment described above, for example... Figure 1 or Figure 2 The task processing device shown is an example of a computing chip or computing card. Electronic devices such as servers, high-performance computing devices, AI accelerators, data center accelerator cards, network processing devices, edge computing devices, or other computing systems suitable for high bandwidth, low latency, and high integration requirements can efficiently support a variety of heterogeneous workloads, covering scientific computing, AI training and inference, quantum computing, image rendering, and other computing tasks.
[0131] Figure 7 This is a schematic block diagram of another electronic device provided for at least one embodiment of the present disclosure.
[0132] For example, such as Figure 7 As shown, the electronic device 700 includes at least one processor 701 and at least one memory 702. The at least one memory 702 includes one or more computer program modules. These computer program modules are stored in the memory 702 and configured to be executed by the at least one processor 701. The one or more computer program modules include instructions for performing the task processing method described above. When executed by the at least one processor 701, they can perform one or more steps of the task processing method provided in at least one embodiment of this disclosure. The memory 702 and the processor 701 can be interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0133] For example, processor 701 can be a central processing unit (CPU), digital signal processor (DSP), graphics processing unit (GPU), general-purpose graphics processing unit (GPGPU), AI accelerator, or other processing unit with data processing and / or program execution capabilities, such as a field-programmable gate array (FPGA); for example, the central processing unit (CPU) can be an x86, ARM, or RISC-V architecture. Processor 701 can be a general-purpose processor or a special-purpose processor, capable of controlling other components in electronic device 700 to perform desired functions.
[0134] For example, memory 702 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc.
[0135] Figure 8 This is a schematic block diagram of another electronic device provided for at least one embodiment of the present disclosure.
[0136] The electronic devices in at least one embodiment of this disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, and fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0137] The electronic device includes at least one processor and a memory. The processor may be referred to as processing device 801 as described below, and the memory may include at least one of ROM 802, RAM 803, and storage device 808 as described below. The memory is used to store programs for performing the methods described in the various method embodiments above; the processor is configured to execute the programs stored in the memory. The processor may include a central processing unit (CPU) or other forms of processing unit having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0138] like Figure 8 As shown, the electronic device 800 may include a processing unit 801 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in ROM 802 or a program loaded from storage device 808 into RAM 803. RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interfaces are also connected to bus 804.
[0139] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, displays, speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0140] In particular, according to at least one embodiment of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, at least one embodiment of this disclosure includes a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of at least one embodiment of this disclosure.
[0141] It should be noted that the computer-readable medium described above in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In at least one embodiment of this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In at least one embodiment of this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, radio frequency (RF), etc., or any suitable combination thereof.
[0142] The aforementioned computer-readable medium may be included in the aforementioned electronic device 800; or it may exist independently and not assembled into the electronic device 800.
[0143] Figure 9 This is a schematic block diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure.
[0144] For example, such as Figure 9 As shown, a non-transitory computer-readable storage medium 900 stores computer-readable instructions 901, which, when executed by at least one processor, perform one or more steps of the task processing method described above.
[0145] For example, the storage medium may include a memory card for a smartphone, a storage component for a tablet computer, a hard drive for a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, or other suitable storage media. For example, the readable storage medium may also be... Figure 7 The memory 702 in the memory is described in the foregoing content and will not be repeated here.
[0146] Although the present disclosure has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to the embodiments of the present disclosure, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present disclosure are within the scope of protection claimed by the present disclosure.
[0147] The following points should be noted regarding this disclosure:
[0148] (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure. Other structures can be referred to the general design.
[0149] (2) For clarity, the thickness of layers or regions in the drawings used to describe embodiments of the present disclosure is enlarged or reduced, i.e., these drawings are not drawn to actual scale.
[0150] (3) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0151] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. The scope of protection of this disclosure should be determined by the scope of protection of the claims.
Claims
1. A task processing device, characterized in that, The task processing device includes computation scheduling cores and functional cores integrated in a single package. The computing scheduling core includes a hardware subsystem and a computing core. The hardware subsystem is configured to receive and parse task information to obtain at least one task, and to assign the at least one task to at least one of the functional cores and the computing cores according to the task type; The computing core is configured to perform computing operations according to the tasks assigned by the hardware subsystem. The functional cores include quantum computing management cores and ray tracing core cores. The quantum computing management chip is configured to control a quantum computing device to perform quantum computing operations according to a task assigned by the hardware subsystem, wherein the quantum computing device is connected to the quantum computing management chip; The ray tracing core is configured to perform ray tracing calculations according to the tasks assigned by the hardware subsystem.
2. The task processing device according to claim 1, characterized in that, The computing scheduling kernel also includes a bus interface, through which the computing scheduling kernel receives task information from external devices.
3. The task processing device according to claim 1, characterized in that, The computation scheduling kernel also includes an inter-kernel interconnection interface, through which the computation scheduling kernel is connected to the functional kernel.
4. The task processing device according to claim 1, characterized in that, The single package also integrates a memory. The computation scheduling kernel also includes a memory control interface, through which the computation scheduling kernel is connected to the memory.
5. The task processing apparatus according to claim 4, characterized in that, The hardware subsystem includes a command processor, a memory management unit, and a data transfer unit. The command processor is configured to receive and parse task information to obtain at least one task, and to assign the at least one task to at least one of the functional cores and the computing cores according to the task type; The memory management unit is configured to manage access to the memory; The data transfer unit is configured to perform data transfer operations between the computing scheduling core and the memory.
6. The task processing apparatus according to claim 1, characterized in that, The functional modules in the computing scheduling chip are interconnected through an on-chip interconnect structure.
7. The task processing apparatus according to claim 1, characterized in that, The computational core includes at least one of a tensor computational core and a vector computational core.
8. The task processing apparatus according to claim 1, characterized in that, The functional core also includes input / output cores configured to enable interconnection between the plurality of said task processing devices.
9. A task processing method, characterized in that, The task processing method is applied to the task processing apparatus according to any one of claims 1 to 8, comprising: The hardware subsystem receives and parses task information to obtain at least one task, and assigns the at least one task to at least one of the functional core and the computing core according to the task type; In response to receiving a first type of task assigned by the hardware subsystem, the computing core performs a computing operation according to the first type of task; In response to receiving a second type of task assigned by the hardware subsystem, the quantum computing management chip controls the quantum computing device to perform quantum computing operations according to the second type of task; In response to receiving a third type of task assigned by the hardware subsystem, the ray tracing core performs ray tracing computation operations according to the third type of task.
10. The task processing method according to claim 9, characterized in that, The at least one task includes at least one of the first type of task, the second type of task, and the third type of task. The hardware subsystem allocates the at least one task to at least one of the computing cores of the functional core and the computing scheduling core according to the task type, including: The hardware subsystem assigns the first type of task to the computing core; The hardware subsystem assigns the second type of task to the quantum computing management chip; or The hardware subsystem assigns the third type of task to the ray tracing core.
11. The task processing method according to claim 9, characterized in that, The task processing method further includes: The computing core accesses memory or input / output chips according to the first type of task; The quantum computing management chip accesses the memory or the input / output chip according to the second type of task; or The ray tracing core accesses the memory or the input / output core according to the third type of task.
12. The task processing method according to claim 9, characterized in that, After the quantum computing management chip controls the quantum computing device to perform quantum computing operations according to the second type of task, the task processing method further includes: In response to the completion of the quantum computing operation, the quantum computing management chip sends a second termination signal to the hardware subsystem.
13. The task processing method according to claim 12, characterized in that, The task processing method further includes: In response to receiving the second termination signal, the hardware subsystem assigns the first type of task to the computing core, so that the computing core performs a computational operation based on the result of the quantum computing operation.
14. The task processing method according to claim 9 or 13, characterized in that, After the computing core performs a computational operation based on the first type of task, the task processing method further includes: In response to the completion of the computation operation, the computing core sends a first termination signal to the hardware subsystem.
15. The task processing method according to claim 14, characterized in that, The task processing method further includes: In response to receiving the first termination signal, the hardware subsystem assigns the third type of task to the ray tracing core kernel, so that the ray tracing core kernel performs ray tracing computation based on the result of the computation operation.
16. An electronic device, characterized in that, The electronic device includes a task processing apparatus according to any one of claims 1-8.
17. An electronic device, characterized in that, The electronic device includes: At least one processor; At least one memory, including one or more computer program modules; The one or more computer program modules are stored in the at least one memory and configured to be executed by the at least one processor, and the one or more computer program modules are used to implement the task processing method according to any one of claims 9-15.
18. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer-readable instructions, wherein the computer-readable instructions, when executed by at least one processor, perform the task processing method according to any one of claims 9-15.
Citation Information
Patent Citations
Calculation module, wafer level processor and wafer level computer
CN120509502A
Firmware partitioning for GPU via virtual SoC
CN120631552A
Calculation method of AI chip based on RISC-V CPU Core and chip
CN121255709A