Processing device, chip system and computer equipment

By designing multiple processing units in the processing device to cooperate with an accelerator, and using the arbitration unit and the cache unit for resource management, the problem of low resource utilization of the accelerator is solved, and efficient data transmission and processing flow is achieved.

CN120218146APending Publication Date: 2025-06-27BEIJING ESWIN COMPUTING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510182419.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Because the accelerator's processing speed is faster than the CPU, the accelerator's input data may be insufficient and the resource utilization rate is low.

Method used

A processing device is designed, including a plurality of processing units and an accelerator, which is used to acquire operation instructions and send instructions and data of the second operation operation to the accelerator, and the accelerator executes and returns the operation result. Through the cooperation of the arbitration unit and the cache unit, efficient data transmission and resource management between the processing unit and the accelerator are realized.

Benefits of technology

The resource utilization rate of the accelerator is improved, the problem of accelerator calculation results being stacked in the processing unit is avoided, and the interaction matching between the processing unit and the accelerator is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218146A_ABST
    Figure CN120218146A_ABST
Patent Text Reader

Abstract

The invention discloses a processing device, a chip system and computer equipment, and belongs to the technical field of computers. The processing device comprises a plurality of processing units and accelerators, and the computing power of the accelerators is higher than that of the processing units; the processing unit is used for executing the first operation to obtain an operation result of the first operation; the processor is further used for providing instructions and data of a second arithmetic operation for the accelerator, and the computational complexity of the second arithmetic operation is higher than that of the first arithmetic operation; and the accelerator is used for executing a second arithmetic operation based on the instruction and the data under the condition that the instruction and the data provided by the processing unit are obtained, obtaining an arithmetic result of the second arithmetic operation, and sending the arithmetic result of the second arithmetic operation to the processing unit. Therefore, the accelerator can be used for accelerating the second operation of the plurality of processing units, the computing power of the accelerator is fully utilized, and the resource utilization rate of the accelerator is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and particularly to a processing device, a chip system, and a computer device. Background Art

[0002] In the field of computer technologies, a general-purpose processor can be combined with an accelerator to improve the performance and efficiency of the general-purpose processor. Among them, the accelerator is used to execute certain types of calculations and has high computing power. The general-purpose processor is, for example, a Central Processing Unit (CPU), etc.

[0003] In related technologies, one CPU is combined with one accelerator. However, since the processing speed of the accelerator is faster than that of the CPU, and the CPU's requirements for the accelerator change in different stages, the input data of the accelerator may be insufficient, resulting in low resource utilization of the accelerator. Summary of the Invention

[0004] This application provides a processing device, a chip system, and a computer device for improving the resource utilization of an accelerator.

[0005] In a first aspect, a processing device is provided. The processing device includes a plurality of processing units and an accelerator, and the computing power of the accelerator is higher than that of the processing units.

[0006] The processing unit is configured to obtain an operation instruction, and when the operation instruction indicates a first operation, execute the first operation to obtain an operation result of the first operation; when the operation instruction indicates a second operation, provide an instruction and data of the second operation to the accelerator, and receive an operation result of the second operation returned by the accelerator, where the computational complexity of the second operation is higher than that of the first operation.

[0007] The accelerator is configured to obtain the instruction and data provided by the processing unit, execute the second operation based on the instruction and the data to obtain an operation result of the second operation, and send the operation result of the second operation to the processing unit.

[0008] In a possible implementation, the processing device further includes an arbitration unit; the arbitration unit is configured to receive instructions and data provided by the plurality of processing units, select one of the plurality of processing units, and send the instructions and data corresponding to the selected processing unit to the accelerator.

[0009] In a possible implementation, the selected processing unit is the processing unit with the highest arbitration index among the multiple processing units, and the arbitration index is determined based on at least one of the instruction sending order, the operation type of the instruction, the data volume of the data, and the priority of the processing unit.

[0010] In a possible implementation, the processing device further includes a cache unit; the cache unit is used to cache the instructions and data provided by the processing unit to the accelerator; the arbitration unit is used to read the instructions and data provided by the processing unit from the cache unit.

[0011] In a possible implementation, the cache unit is further used to cache the operation results sent by the accelerator to the processing unit.

[0012] In a possible implementation, the number of cache units is multiple, and the multiple processing units correspond to the multiple cache units one by one.

[0013] In a possible implementation, the multiple processing units multiplex the accelerator separately in time.

[0014] In a possible implementation, the processing unit includes an extended instruction interface, and the processing unit provides the instructions and data of the second operation operation to the accelerator through the extended instruction interface.

[0015] In a possible implementation, the second operation operation is a matrix operation in a neural network algorithm.

[0016] In a second aspect, a chip system is further provided, and the chip system includes the processing device described in the first aspect above.

[0017] In a third aspect, a computer device is further provided, and the computer device includes the processing device described in the first aspect above.

[0018] The technical solution provided by this application can at least bring the following beneficial effects:

[0019] In the processing device provided by this application, by adopting the method of multiple processing units corresponding to one accelerator, the accelerator can be used to accelerate the second operation operations of multiple processing units, making full use of the computing power of the accelerator and improving the resource utilization rate of the accelerator. Moreover, for each processing unit, it can avoid the problem that the processing speed of the accelerator is too fast, resulting in the accumulation of the operation results of the second operation operations in the processing unit and not being processed in time by the processing unit, making the interaction between the processing unit and the accelerator more matched. Description of the Drawings

[0020] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0021] Figure 1 is a schematic structural diagram of a processing device provided by an embodiment of the present application;

[0022] Figure 2 is a schematic structural diagram of another processing device provided by an embodiment of the present application;

[0023] Figure 3 is a schematic structural diagram of another processing device provided by an embodiment of the present application;

[0024] Figure 4 is a schematic structural diagram of another processing device provided by an embodiment of the present application;

[0025] Figure 5 is a schematic structural diagram of a chip system provided by an embodiment of the present application;

[0026] Figure 6 is a schematic structural diagram of a computer device provided by an embodiment of the present application. Specific Embodiments

[0027] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0028] With the rapid development of artificial intelligence, neural network algorithms are also constantly evolving. For example, neural network algorithms have evolved from perceptrons to multi-layer perceptrons, to deep neural networks, and then to convolutional neural networks (CNNs) applied to the field of images and recurrent neural networks (RNNs) applied to the field of speech, etc.

[0029] In the process of the evolution of neural network algorithms, the complexity and computational amount of neural network algorithms have increased sharply, and the requirements for the computing performance of hardware have also become higher and higher. Although traditional general-purpose processors can handle the calculations of neural network algorithms, they often struggle to meet the requirements of specific application scenarios in terms of energy efficiency ratio and computing density. Among them, general-purpose processors are, for example, central processing units (CPUs) or graphics processing units (GPUs), etc.

[0030] Thus, the Neural Network Accelerator engine (NNA) came into being. Through optimizing the hardware architecture and algorithms, it is customized according to the characteristics of neural network calculations to achieve higher computing efficiency and lower power consumption. Among them, the neural network accelerator includes a dedicated computing unit for matrix multiplication operations, namely the matrix acceleration unit. The matrix acceleration unit usually includes a tensor computing array and data transceiver control, which jointly complete computing tasks such as matrix multiplication and convolution.

[0031] The basic building block of a neural network is the perceptron. Multiple perceptrons gather together to form a perceptron layer, and a neural network is composed of a large number of cascaded perceptron layers. In neural network algorithms, linear transformations are usually implemented through matrix multiplication. Therefore, accelerating matrix multiplication becomes the key to accelerating the neural network training and inference processes. That is to say, in neural network algorithms, the highest computing power demand is for matrix operations, which include but are not limited to matrix multiplication or convolution operations, etc. Matrix operations are usually dozens or even hundreds of times that of other types of operations except matrix operations, and the demand for matrix operations is relatively stable and has not undergone a subversive change during the evolution of neural network algorithms.

[0032] Therefore, matrix operations are the part of neural network algorithms that most require the design of targeted accelerators. However, there are many types of operations in neural network algorithms other than matrix operations, and it is difficult to process them through highly customized accelerators. For example, the operations of activation functions, etc. These operations other than matrix operations need to be processed by a general-purpose processor.

[0033] In this case, through the extended instructions of the general-purpose processor, the general-purpose processor and the matrix acceleration unit can be combined so that the general-purpose processor obtains the same computing power for matrix operations as the matrix acceleration unit. Among them, the general-purpose processor includes an extended instruction interface, which is used to provide efficient and convenient data interaction between the matrix acceleration unit and the general-purpose processor. Thus, the design based on extended instructions provides a strong guarantee for the overall optimization of neural network algorithms, enabling the neural network accelerator to better adapt to various neural network algorithms.

[0034] Exemplarily, the general-purpose processor may be a CPU, and the extended instruction interface of the CPU may be a Vector Co-processor Interface Extension (vcix) interface or a Scalar Interface Extension (scie). Through the scie / vcix interface, the instructions and general registers of the CPU can be directly opened to the neural network accelerator. Among them, the vcix interface allows users to add custom vector instructions to improve the vector operation performance of the processor. The scie interface allows users to add custom scalar instructions to enhance the computing power of the processor.

[0035] In a possible implementation manner, the process of implementing a neural network algorithm by combining a general-purpose processor and a matrix acceleration unit may include: the general-purpose processor provides instructions and data corresponding to matrix multiplication operations to the matrix acceleration unit through an extended instruction interface; the matrix acceleration unit performs matrix multiplication operations based on the instructions and data, and sends the operation results of the matrix multiplication operations to the general-purpose processor.

[0036] However, for a scenario where a general-purpose processor and a matrix acceleration unit are combined, since the processing speed of the matrix acceleration unit is faster than the speed at which the general-purpose processor provides instructions and data, the input data of the matrix acceleration unit may be insufficient, resulting in a low resource utilization rate of the matrix acceleration unit. Moreover, the general-purpose processor needs to manage a vector process unit (VPU) while managing matrix multiplication operations. The vector processing unit is a specially designed processor with a highly pipelined operation, which is used to perform operations other than matrix multiplication operations in neural network operations.

[0037] Therefore, for the operation results quickly processed by the matrix acceleration unit, the general-purpose processor may be difficult to process in a timely manner, resulting in data accumulation in the general-purpose processor, that is, the pipeline appears a bubble phenomenon, which may further lead to the exhaustion of memory or other resources, and may even cause system performance problems or failures. Here, the pipeline having a bubble usually means that in the data processing pipeline, the data processing speed at a certain stage is slower than that of the previous stage, resulting in data accumulation, forming a bubble, that is, the "hanging" point of data accumulation.

[0038] Taking neural network operations as an example of MobileNet, MobileNet is a lightweight convolutional neural network. MobileNet uses depthwise separable convolutions, which decompose a standard convolution into two steps: depthwise convolution and pointwise convolution. In a standard convolution, each convolution kernel performs convolution calculations on the feature maps of all input channels, and the number of convolution kernels is the same as the number of input channels, resulting in a relatively large computational cost for the convolution operation. In depthwise separable convolutions, each convolution kernel independently performs convolution calculations on the feature maps of only one input channel, and then performs pointwise convolution operations on the convolution results of each input channel to merge the feature maps of different channels, thereby generating the final output feature map. This decomposition can reduce the number of parameters and computational complexity, that is, it can significantly reduce the computational cost of the convolution operation, enabling MobileNet to operate efficiently on mobile devices.

[0039] In MobileNet, the computational cost of the convolution operation is significantly reduced, while the computational costs of operations such as batch normalization (BN) and activation operations increase significantly. That is, compared with the evolution of the standard convolutional neural network in MobileNet, the convolution computational cost decreases, and the non-convolution computational cost increases, resulting in the general-purpose processor not being able to well match the processing speed of the matrix acceleration unit, leading to the problem of pipeline bubbles. That is to say, with the continuous evolution of neural network algorithms, different neural network algorithms have different requirements for accelerators. How to combine the general-purpose processor and the matrix acceleration unit to be flexibly applicable to or compatible with more types of neural network algorithms is an urgent problem to be solved.

[0040] See Figure 1 , Figure 1 which is a schematic structural diagram of a processing device provided by an embodiment of the present application. As Figure 1 shown, the processing device includes multiple processing units and an accelerator, and the computing power of the accelerator is higher than that of the processing units. The embodiment of the present application does not limit the number of processing units included in the processing device, Figure 1 and 3 processing units are taken as an example for illustration. In actual applications, the number of processing units can be 2, 4, or any other arbitrary number.

[0041] Taking the first processing unit as an example, the first processing unit is any one of multiple processing units. The first processing unit is configured to obtain an operation instruction. When the operation instruction indicates a first operation, the first processing unit executes the first operation to obtain the operation result of the first operation. When the operation instruction indicates a second operation, the first processing unit provides the instruction and data of the second operation to the accelerator and receives the operation result of the second operation returned by the accelerator. Herein, the computational complexity of the second operation is higher than that of the first operation.

[0042] Optionally, the first processing unit includes but is not limited to a core, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an AI (Artificial Intelligence) processor, or a coprocessor, etc.

[0043] The accelerator is configured to obtain the instruction and data provided by the first processing unit, execute the second operation based on the instruction and data to obtain the operation result of the second operation, and send the operation result of the second operation to the first processing unit. Herein, the accelerator is used to accelerate the second operation of the first processing unit and has a higher computing power than the first processing unit. For example, the accelerator can be the above-mentioned neural network accelerator or matrix acceleration unit.

[0044] When the processing device is used to implement the execution of the neural network algorithm, the second operation is a matrix operation in the neural network algorithm, and the first operation is an operation other than the matrix operation in the neural network algorithm. Herein, multiple processing units can execute the operation of different network layers of the same neural network algorithm or the operations of different neural network algorithms.

[0045] The embodiments of the present application do not limit the data transmission manner between the processing unit and the accelerator. For example, the first processing unit includes an extended instruction interface, and the extended instruction interface is used to implement the interaction between the processing unit and the accelerator. Then, the first processing unit provides the instruction and data of the second operation to the accelerator through the extended instruction interface. When the first processing unit is a CPU, the extended instruction interface is the extended instruction interface of the CPU, and the instruction of the second operation provided by the first processing unit to the accelerator is the extended instruction of the CPU.

[0046] Among them, as the CPU instruction set is continuously improved, the general - purpose processor instruction set in the embodiments of the present application supports extended instructions, and users can customize the required instructions according to the extended instructions. The instructions output by the extended instruction interface in the embodiments of the present application can be understood as neural network acceleration instructions based on the CPU extended instructions, so as to implement a neural network accelerator in the chip system through the neural network acceleration instructions.

[0047] In the embodiments of the present application, the interaction mechanisms between multiple processing units and the accelerator include but are not limited to the following three.

[0048] First, multiple processing units multiplex the accelerator separately in time.

[0049] For example, multiple processing units interact with the accelerator in different time periods respectively. After one processing unit finishes interacting with the accelerator, another processing unit can interact with the accelerator. Optionally, multiple processing units interact with the accelerator in a polling manner, and the lengths of the time periods for different processing units to interact with the accelerator can be the same or different.

[0050] Second, multiple processing units interact with the accelerator simultaneously.

[0051] Thus, multiple processing units compete for one accelerator. Optionally, the processing device further includes an arbitration unit, which is used to receive the instructions and data provided by multiple processing units, and select one processing unit from the multiple processing units, and send the instructions and data corresponding to the selected processing unit to the accelerator.

[0052] Exemplarily, referring to Figure 2 the structural schematic diagram of another processing device shown, the arbitration unit is located between multiple processing units and the accelerator, so that multiple processing units are connected to the accelerator through the arbitration unit. Thus, the arbitration unit arbitrates the instructions and data provided by multiple processing units respectively, and sends the instructions and data provided by the second processing unit determined by the arbitration to the accelerator, and the second processing unit is the processing unit selected by the arbitration unit.

[0053] Among them, the present application does not limit the manner in which the arbitration unit arbitrates the instructions and data provided by multiple processing units respectively. In a possible implementation manner, the second processing unit is the processing unit with the highest arbitration index among multiple processing units, and the arbitration index is determined based on at least one of the instruction sending order, the instruction operation type, the data volume of the data, and the priority of the processing unit.

[0054] Exemplarily, the earlier the sending order of the instructions, the higher the arbitration index, that is, the earlier sent instructions are processed with higher priority; the operation types of the instructions may include matrix multiplication and convolution operations, and the arbitration index of the convolution operation is higher; the larger the data volume of the data, the higher the arbitration index; the higher the priority of the processing unit, the higher the arbitration index, and the priority of the processing unit can be configured in advance.

[0055] In a possible implementation manner, the processing device further includes a cache unit, which is used to cache the instructions and data provided by the processing unit to the accelerator. For example, it caches the instructions and data provided by multiple processing units to the accelerator respectively. Optionally, the cache unit is further used to cache the operation results sent by the accelerator to the processing unit. Exemplarily, the cache unit can be a queue.

[0056] In this case, the arbitration unit is used to read the instructions and data provided by the processing unit from the cache unit. For example, the arbitration unit arbitrates the instructions and data of multiple processing units cached in the cache unit. Among them, the cache unit can be located between multiple processing units and the arbitration unit, so that multiple processing units are connected to the arbitration unit through the cache unit.

[0057] Optionally, the number of cache units is multiple, and multiple processing units correspond to multiple cache units one by one. Refer to Figure 3 the structural schematic diagram of another processing device shown. The number of cache units is multiple. Among them, one processing unit corresponds to one cache unit, and the cache units corresponding to each processing unit are used to cache the instructions and data sent by the corresponding processing unit to the accelerator. Optionally, the cache unit corresponding to the first processing unit is used to cache the operation results sent by the accelerator to the first processing unit. The cache unit enables rate matching between the processing unit and the accelerator.

[0058] In summary, in the processing device provided in the embodiments of the present application, by the method that multiple processing units correspond to one accelerator, the accelerator can be used to accelerate the second operation operations of multiple processing units, making full use of the computing power of the accelerator and improving the resource utilization rate of the accelerator. Moreover, for each processing unit, it can avoid the problem that the processing speed of the accelerator is too fast, resulting in the accumulation of the operation results of the second operation operations in the processing unit and not being processed in time by the processing unit, making the interaction combination between the processing unit and the accelerator more matched.

[0059] Thirdly, at least two of the multiple processing units interact with the accelerator simultaneously in the first time period, so that the at least two processing units compete for one accelerator in the first time period; the other processing units except the at least two processing units among the multiple processing units interact with the accelerator at different time periods within the second time period. Among them, the first time period and the second time period do not overlap.

[0060] Among them, for the mechanism in which at least two processing units interact with the accelerator simultaneously, reference can be made to the mechanism in which multiple processing units interact with the accelerator simultaneously in the second method above. For the mechanism in which other processing units interact with the accelerator respectively, reference can be made to the mechanism in which multiple processing units interact with the accelerator at different time periods respectively in the first method above.

[0061] Optionally, at least two processing units are connected to the accelerator through an arbitration unit; the arbitration unit is configured to arbitrate the instructions and data respectively provided by at least two processing units, and send the instructions and data determined by the arbitration and provided by the second processing unit to the accelerator. The second processing unit is the processing unit with the highest arbitration index among at least two processing units, and the arbitration index is determined based on at least one of the sending order of the instructions, the operation type of the instructions, the data volume of the data, and the priority of the processing unit.

[0062] In a possible implementation manner, at least two processing units are connected to the arbitration unit through a cache unit; the cache unit is configured to cache the instructions and data respectively provided by at least two processing units to the accelerator; the arbitration unit is configured to arbitrate the instructions and data of at least two processing units cached in the cache unit. Optionally, the cache unit is further configured to cache the operation results sent by the accelerator to at least two processing units.

[0063] Exemplarily, the number of cache units is at least two, and one processing unit among at least two processing units corresponds to one cache unit. The cache unit corresponding to each processing unit is configured to cache the instructions and data sent by the corresponding processing unit to the accelerator.

[0064] Next, taking the processing unit as the CPU, the accelerator as the matrix unit, and the interaction mechanism between multiple CPUs and one matrix unit as the second method above as an example, the processing device provided by the embodiments of the present application will be described by way of example. Refer to Figure 4 As shown in the structural schematic diagram of the processing device, multiple CPUs are connected to one matrix unit through an arbiter. A queue is included on the path from each CPU among multiple CPUs to the arbiter. The arbiter corresponds to the above arbitration unit, and the queue corresponds to the above cache unit. Each CPU includes a vector unit, and the vector unit is configured to perform operations other than matrix operations.

[0065] Among them, the CPU provides instructions and data for matrix operations to the matrix unit through the extended instruction interface. Since the processing speed of the matrix unit is faster than that of the CPU in providing instructions and data, and the CPU needs to manage the vector unit while managing matrix operations, it is necessary to handle the load balance between the CPU and the matrix unit well.

[0066] The load balance between the CPU and the matrix is divided into two levels: ① Three CPUs compete for one matrix unit. By utilizing the characteristic that the computing requirements of matrix operations at different stages of the CPU are different, the three CPUs use the matrix unit at different times; ② A queue is added between the CPU and the arbiter to relieve the instantaneous load imbalance. That is to say, the arbiter between the CPU and the matrix unit is used to enable multiple CPUs to compete for one matrix unit, and the queue on the path from the CPU to the arbiter is used for rate matching.

[0067] The execution process of the CPU during the operation instruction can include the following steps. Instruction Fetch: The CPU determines the address of the instruction through the Program Counter (PC) register, fetches the instruction from the memory and loads it into the instruction register, and then the PC register is incremented for future execution of the next instruction. Among them, the register is used to store temporary data and instructions, including general-purpose registers and status registers.

[0068] Instruction Decode: According to the instruction in the instruction register, parse out the specific operation type (such as the first arithmetic operation or the second arithmetic operation, etc.), and determine the object of the operation (such as the data in the register or the memory address, etc.). Execute Instruction: According to the parsed operation type and object, the CPU executes the corresponding operation. If it is the first arithmetic operation, it is executed by the vector unit, and if it is the second arithmetic operation, it is executed by the matrix unit.

[0069] When the computing power requirements for matrix operations of each CPU remain unchanged, by increasing the number of CPUs, the throughput of the CPUs increases, the matrix unit has more sufficient data input, and at the same time, the operation results of the matrix unit can be quickly received and processed by the CPUs. For multiple CPUs, taking one layer of the neural network algorithm processed by each CPU as an example, since the proportion of matrix operations in different layers of the network is different, where the matrix operations are completed by the matrix unit and other operations are completed by the CPUs. When the proportion of matrix operations in a certain layer of the network is small, the corresponding CPU can release the matrix unit, enabling the matrix unit to be used by the CPUs processing other network layers, thus avoiding unnecessary idle time of the matrix unit.

[0070] Therefore, the processing device provided in the embodiment of the present application adopts the method of multiple CPUs competing for 1 matrix unit to make full use of the computing power of the matrix unit, improve the utilization rate of the matrix unit, and avoid bubbles in the pipeline. It provides high flexibility on the premise of ensuring sufficient throughput and can be compatible with more neural network algorithms.

[0071] Figure 5 It is a schematic structural diagram of a chip system provided in the embodiment of the present application. The chip system includes Figures 1-4 Any of the shown processing devices. Exemplarily, the chip system can be a SoC (System on Chip) chip.

[0072] Among them, the chip system includes multiple processing units and accelerators, belonging to a heterogeneous system-level chip. A heterogeneous system-level chip refers to a chip integrated with different types of processor cores on one chip. Different processor cores are used to optimize different computing tasks, and different processor cores work together to improve performance, reduce power consumption, save costs and space. Since heterogeneous system-level chips usually face complex and inefficient problems, therefore, the interaction mechanism between the multiple processing units and accelerators provided in the embodiment of the present application is more friendly to software, has better immediacy, and can reduce the development difficulty of heterogeneous system-level chips.

[0073] Figure 6 It is a schematic structural diagram of a computer device provided in the embodiment of the present application. The computer device 800 includes a processing device 801, and the processing device 801 is Figures 1-4 Any of the shown processing devices. As Figure 6 shown, the processing unit is connected to the accelerator through an extended instruction interface, Figure 6The number of processing units in [description] is only an example. In practical applications, the number of processing units included in the processing device 801 can be any number.

[0074] In Figure 6 [description], the processing device 801 is coupled to the memory 802, and it should be understood that the computer device 800 also supports other memory configurations known in the art. The memory 802 may include one or more computer-readable storage media, which may be non-transitory. At least one computer program is stored in the computer-readable storage media, and the at least one computer program is loaded and executed by the processing device 801.

[0075] The memory 802 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 is used to store at least one instruction for being executed by the processing device 801.

[0076] In one possible implementation, the above computer-readable storage media may be read-only memory (ROM), random access memory (RAM), compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage devices, etc.

[0077] Figure 6 A display 806 coupled to the processing device 801 through a display controller 804 is also shown. In some cases, the computer device 800 can be used for wireless communication. Figure 6 A speaker 809 and a microphone 810 coupled to the processing device 801 through an encoder / decoder 811 are also shown; and a wireless antenna 808 coupled to a wireless controller 805.

[0078] The display 806 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display 806 is a touch screen display, the display 806 also has the ability to collect touch signals on or above the surface of the display 806. At this time, the display 806 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display 806, which is disposed on the front panel of the terminal; in other embodiments, there may be at least two displays 806, which are respectively disposed on different surfaces of the terminal or are in a folding design; in other embodiments, the display 806 may be a flexible display screen, which is disposed on a curved surface or a folding surface of the terminal. Even, the display 806 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display 806 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0079] The microphone 810 is used to collect sound waves of the user and the environment, and input the sound waves to the processing device 801 for processing. For the purpose of stereo collection or noise reduction, there may be multiple microphones 810, which are respectively disposed at different parts of the terminal. The microphone 810 can also be an array microphone or an omnidirectional collection type microphone. The speaker 809 is used to convert an electrical signal from the processing device 801 into sound waves. The speaker 809 can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert an electrical signal into sound waves audible to humans, but also convert an electrical signal into sound waves inaudible to humans for uses such as ranging.

[0080] The processing device 801 and the memory 802 may be included in a system-in-package or a system-on-chip device.

[0081] The input device 807 and the power supply 803 are coupled to the system-on-chip device 812. Optionally, as Figure 6 shown, when one or more optional boxes exist, the display 806, the input device 807, the speaker 809, the microphone 810, the wireless antenna 808, and the power supply 803 are outside the system-on-chip device 812. However, each of the display 806, the input device 807, the speaker 809, the microphone 810, the wireless antenna 808, and the power supply 803 can be coupled to a component of the system-on-chip device 812, such as an interface or a controller.

[0082] The power supply 803 is used to supply power to each component in the terminal. The power supply 803 can be alternating current, direct current, a primary battery, or a rechargeable battery. When the power supply 803 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0083] In a possible implementation, the processing device 801 and the memory 802 can be integrated into the following: a set-top box, a server, a music player, a video player, an entertainment unit, a navigation device, a personal digital assistant (PDA), a fixed-location data unit, a computer, a laptop computer, a tablet computer, a communication device, a mobile phone, or other similar devices.

[0084] Those skilled in the art can understand that Figure 6 the structure shown in does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0085] The terms "first", "second", "third", "fourth", etc. in the description, claims, and drawings of this application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having", and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0086] The above are only optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the principles of this application shall be included within the protection scope of this application.

Claims

1. A processing device, characterized in that: The processing device includes a plurality of processing units and an accelerator, wherein the computing capability of the accelerator is higher than the computing capability of the processing unit; The processing unit is configured to obtain a computing instruction, and when the computing instruction indicates a first computing operation, execute the first computing operation to obtain a computing result of the first computing operation; when the computing instruction indicates a second computing operation, provide the accelerator with instructions and data for the second computing operation, and receive a computing result of the second computing operation returned by the accelerator, wherein the computing complexity of the second computing operation is higher than the computing complexity of the first computing operation; The accelerator is used to obtain instructions and data provided by the processing unit, perform the second computing operation based on the instructions and the data, obtain the computing result of the second computing operation, and send the computing result of the second computing operation to the processing unit.

2. The device according to claim 1, characterized in that The processing device also includes an arbitration unit; The arbitration unit is used to receive instructions and data provided by the multiple processing units, select one of the multiple processing units, and send the instructions and data corresponding to the selected processing unit to the accelerator.

3. The device according to claim 2, characterized in that The selected processing unit is a processing unit with the highest arbitration index among the multiple processing units, and the arbitration index is determined based on at least one of the sending order of instructions, the operation type of instructions, the data volume, and the priority of the processing unit.

4. The device according to claim 2, characterized in that The processing device also includes a cache unit; The cache unit is used to cache instructions and data provided by the processing unit to the accelerator; The arbitration unit is used to read the instructions and data provided by the processing unit from the cache unit.

5. The device according to claim 4, characterized in that The cache unit is also used to cache the calculation results sent by the accelerator to the processing unit.

6. The device according to claim 4, characterized in that There are multiple cache units, and the multiple processing units correspond to the multiple cache units one by one.

7. The device according to claim 1, characterized in that The plurality of processing units time-separately multiplex the accelerator.

8. The device according to any one of claims 1 to 7, characterized in that: The processing unit includes an extended instruction interface, and the processing unit provides the accelerator with instructions and data for the second computing operation through the extended instruction interface.

9. The device according to any one of claims 1 to 7, characterized in that: The second computing operation is a matrix operation in a neural network algorithm.

10. A chip system, characterized in that: The chip system comprises a processing device as described in any one of claims 1-9.

11. A computer device, characterized in that: The computer device comprises a processing apparatus as claimed in any one of claims 1 to 9.