Refresh processing method for computing core, computing device and computing system
Patent Information
- Application Number
- CN202610932095.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-08
AI Technical Summary
然而,对于用户已明确中止的任务,其执行结果已不再被关心,但计算核仍会继续执行已下发的指令包,导致刷新操作的时间开销显著增加,影响系统的响应速度和资源利用效率
[0018] In this embodiment, after receiving a refresh instruction for a target task, the task distribution unit introduces a refresh check mechanism before the computational core instruction packet is sent to the computational core. For computational core instruction packets belonging to the target task, the actual sending to the corresponding computational core is prevented, and the computational core is simulated to generate an instruction completion response, thereby skipping the actual execution of the instruction packet. In this way, while maintaining the pairing relationship between computational core instruction packets and instruction completion responses and ensuring system state consistency, the time overhead of the refresh operation is significantly reduced.
Smart Images

Figure CN122711318A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of chip technology, and more specifically, to a refresh processing method for computing cores, a computing device, and a computing system. Background Technology
[0002] In a computing system, a task distribution unit distributes instruction packets belonging to different tasks to the various computing cores. The computing cores then parse the instruction packets and execute the corresponding tasks. During task execution, if a task times out and the user wishes to abort the task, triggering a refresh operation, the computing system must wait for all instruction packets sent to the task distribution unit to complete execution before completely stopping the task and releasing related resources to ensure system stability and consistency. However, for tasks that the user has explicitly aborted, their execution results are no longer important, but the computing cores will still continue to execute the sent instruction packets, significantly increasing the time overhead of the refresh operation and impacting system response speed and resource utilization efficiency. Summary of the Invention
[0003] One object of this disclosure is to provide a new refresh processing method for computing cores, so as to at least reduce the time overhead of refresh operations.
[0004] According to a first aspect of this disclosure, a refresh processing method for computing cores is provided, comprising: Receive refresh instructions for the target task; Based on the refresh instruction, check whether the computing core instruction package to be sent to the computing core is the target instruction package corresponding to the target task; If the computing core instruction packet is the target instruction packet, prevent the computing core instruction packet from being sent to the corresponding computing core, and generate and return an instruction completion response for the blocked computing core instruction packet.
[0005] Optionally, after checking whether the computing core instruction package to be sent to the computing core is the target instruction package corresponding to the target task, the method further includes: If the computing core instruction package is not the target instruction package, the computing core instruction package is sent to the corresponding computing core.
[0006] Optionally, generating and returning an instruction completion response for the blocked computation kernel instruction packet includes: The simulated computing core generates and returns an instruction completion response for the instruction packet of the blocked computing core.
[0007] Optionally, the step of checking whether the computing core instruction package to be sent to the computing core is the target instruction package corresponding to the target task, based on the refresh instruction, includes: Based on the refresh instruction, intercept the computing core instruction packet from the instruction stream to be issued; Check whether the intercepted computing kernel instruction packet is the target instruction packet of the corresponding target task.
[0008] Optionally, intercepting the computation kernel instruction packet from the instruction stream to be issued includes: Select one of the multiple compute core instruction cache units that correspond one-to-one with multiple compute cores; Intercept computation core instructions read from the selected computation core instruction cache unit.
[0009] Optionally, the computing core instruction package carries the task identifier of the task to which it belongs; the step of checking whether the computing core instruction package to be sent to the computing core is the target instruction package corresponding to the target task includes: Extract the task identifier from the computational core instruction package; The extracted task identifier is compared with the target task identifier indicated by the refresh command; If the comparison results are consistent, the computational kernel instruction package is determined to be the target instruction package.
[0010] Optionally, the step of preventing the distribution of the computing core instruction packet to the corresponding computing core, and generating and returning an instruction completion response for the blocked computing core instruction packet, includes: The instruction packets of the blocked computing kernels will be temporarily stored in the refresh processing cache unit; Read the temporarily stored computation kernel instruction packet from the refresh processing cache unit; For the read computational core instruction packet, simulate the corresponding computational core to generate and return the corresponding instruction completion response.
[0011] Optionally, generating and returning an instruction completion response for the blocked computation kernel instruction packet includes: Generate corresponding instruction completion responses for the blocked computation kernel instruction packets; The instruction completion response is returned to the instruction channel corresponding to the blocked computing core instruction, triggering the instruction channel to release the resources occupied by the blocked computing core instruction based on the instruction completion response.
[0012] Optionally, the method further includes: When all computational core instruction packets of the target task generate corresponding instruction completion responses, the refresh operation for the target task is determined to be complete, and the check operation ends.
[0013] Optionally, there are multiple refresh instructions, each corresponding to a different target task; the step of checking whether the computing core instruction package to be sent to the computing core is the target instruction package corresponding to the target task, based on the refresh instructions, includes: For each refresh instruction, check whether the computing core instruction packet to be sent to the computing core is the target instruction packet corresponding to the target task. If the computation core instruction packet is the target instruction packet, prevent the computation core instruction packet from being sent to the corresponding computation core, including: If the computing core instruction package is a target instruction package corresponding to any target task, prevent the computing core instruction package from being sent to the corresponding computing core.
[0014] Optionally, the target task is a video processing task, and the computing core is a computing core that performs video encoding and decoding tasks.
[0015] According to a second aspect of this disclosure, a computing device is also provided, comprising: Multiple computational cores; and, A task distribution unit configured to perform the method according to the first aspect of this disclosure.
[0016] According to a second aspect of this disclosure, another computing device is also provided, comprising: The refresh instruction receiving unit is configured to receive refresh instructions for the target task; The instruction packet checking unit is configured to check, based on the refresh instruction, whether the computing core instruction packet to be sent to the computing core is the target instruction packet corresponding to the target task; and The instruction packet processing unit is configured to, when the computing core instruction packet is the target instruction packet, prevent the computing core instruction packet from being sent to the corresponding computing core, and generate and return an instruction completion response for the blocked computing core instruction packet.
[0017] According to a fourth aspect of this disclosure, a computing system is also provided, the computing system including a control device and a computing device, the computing device being the computing device described in the second or third aspect of this disclosure, the control device being configured to send computing core instruction packets to the computing device for a required processing task.
[0018] In this embodiment, after receiving a refresh instruction for a target task, the task distribution unit introduces a refresh check mechanism before the computational core instruction packet is sent to the computational core. For computational core instruction packets belonging to the target task, the actual sending to the corresponding computational core is prevented, and the computational core is simulated to generate an instruction completion response, thereby skipping the actual execution of the instruction packet. In this way, while maintaining the pairing relationship between computational core instruction packets and instruction completion responses and ensuring system state consistency, the time overhead of the refresh operation is significantly reduced.
[0019] The features and advantages of the embodiments of this specification will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0020] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments of this specification and, together with their description, serve to explain the principles of these embodiments.
[0021] Figure 1 This is a schematic diagram of the structure of a computing system to which the methods of the embodiments of this disclosure can be applied; Figure 2 This is a schematic diagram of the structure of a computing device; Figure 3 This is a flowchart illustrating a refresh processing method according to some embodiments; Figure 4 This is a flowchart illustrating a refresh processing method according to other embodiments; Figure 5 This is a schematic diagram of the structure of a computing device according to some embodiments; Figure 6 This is a schematic diagram of the composition of a computing device according to some embodiments. Detailed Implementation
[0022] Various exemplary embodiments of this specification will now be described in detail with reference to the accompanying drawings.
[0023] The following description of at least one exemplary embodiment is merely illustrative and is not intended to limit the embodiments of this specification or their application or use.
[0024] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0025] This disclosure relates to a refresh processing method in which a computing core performs the task. Figure 1 This is a schematic diagram of the composition structure of a computing system 1000 to which the refresh processing method provided in the embodiments of this disclosure can be applied.
[0026] like Figure 1 As shown, the computing system 1000 includes a computing device 1100, a control device 1200, and a storage device 1300. The computing device, control device, and storage device can communicate with each other via a bus network or other means.
[0027] Storage device 1300 is used to store instructions and / or data, which can be retrieved and used by a control device or a computing device. For example, the storage device can store program instructions executed by the control device or the computing device, or it can store data such as text, images, audio, and configuration parameters. Control device 1200 is used to control the computing device to perform related tasks to achieve corresponding system functions, such as artificial intelligence, scientific computing, image processing, and video processing based on user needs. Exemplarily, the control device can be a central processing unit (CPU); the computing device can be a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a neural network processing unit (NPU), or a tensor processing unit (TPU); the storage device can be static random access memory (SRAM), read-only memory (ROM), or erasable programmable read-only memory (EPROM). Storage devices can also be referred to as memory, and computing devices can also be referred to as chips, processors, or artificial intelligence chips. It is understood that computing system 1000 can also be referred to as an artificial intelligence chip or an artificial intelligence system, and this disclosure does not limit it in this way.
[0028] The aforementioned computing device 1100 is equipped with multiple computing cores for executing tasks, and a task distribution unit for these computing cores.
[0029] Taking video processing tasks as an example, the computing device 1100 has a video processing module (Video Codec), which is responsible for decoding the raw data used for training or inference. The decoded data is then used by the GPU computing cores for training or inference. Figure 2 As shown, the video processing module of the computing device 1100 may include a task distribution unit 1110 and multiple computing cores 1120 for performing video encoding and decoding tasks. In the video processing module, the task distribution unit 1110 is also referred to as a Video Command Processor Dispatch (VCPD).
[0030] In some embodiments, such as Figure 2As shown, the task distribution unit may include multiple instruction channels 1111, an instruction channel selection unit 1112, multiple computing core instruction cache units 1113, and a computing core instruction cache selection unit 1114.
[0031] Multiple instruction channels 1111 are configured to allocate computational core instructions. Each instruction channel can be configured with its own instruction allocation control state machine. Figure 2 Eight instruction channels are shown as an example, namely instruction channel 0 to instruction channel 7.
[0032] The instruction channel selection unit 1112 is configured to select from multiple instruction channels 1111 associated with the computing core instruction to be assigned.
[0033] When selecting an instruction channel, the instruction channel selection unit 1112 may employ the following method: If the instruction cache unit corresponding to the computational core pointed to by the instruction packet at the head of the queue in the current instruction channel is not full, then the current instruction channel is determined as a candidate instruction channel, and the instruction allocation control state machine of that instruction channel is set to a candidate allocation state; then, all candidate instruction channels are randomly selected to determine the selected instruction channel. For example, in the current clock cycle, the instruction allocation control state machines of instruction channel 1, instruction channel 2, and instruction channel 5 are all in a candidate allocation state, then the instruction channel selection unit 1112 randomly selects one instruction channel from instruction channel 1, instruction channel 2, and instruction channel 5.
[0034] The number of multiple core instruction cache units 1113 is the same as the number of multiple cores 1120. Each core instruction cache unit 1113 corresponds one-to-one with each core 1120 and is used to temporarily store core instruction packets pointing to the corresponding core. Core instruction packets in the selected instruction channel are cached in the core instruction cache unit corresponding to the core to which the instruction packet is destined.
[0035] The core instruction cache selection unit 1114 is configured to select a core instruction cache unit that is in a non-empty state and whose corresponding core is in an idle state, and to read the core instruction packets from the selected core instruction cache unit and send them to the corresponding core. The instruction channel selection operation and the core instruction cache unit selection operation can be performed in parallel within the same clock cycle.
[0036] Taking video processing tasks as an example, in related technologies, if a task execution timeout occurs or a user actively terminates the task, triggering a refresh operation, the control device 1200 will stop sending new computational core instruction packets to the target task. However, the computational core instruction packets already in the task distribution unit for this target task, including those already in the computational core instruction cache unit and those still queuing in the instruction channel waiting to enter the computational core instruction cache unit, must wait for the computational core to complete execution and return an instruction completion response before the refresh operation is considered complete. During this process, the residual instruction packets of the terminated task will still be executed completely, occupying the computational core's computation cycle, preventing system resources from being released quickly, and thus affecting the efficiency of multi-task concurrency and switching.
[0037] To address this, this disclosure proposes a refresh processing scheme that intercepts the computation core instruction packet belonging to the suspended target task before it is sent to the computation core and directly returns an instruction completion response. This eliminates the need to actually execute the residual instruction packet to complete the instruction loop, effectively reducing the time overhead of the refresh operation.
[0038] The following will combine Figure 1 The computing system shown and Figure 2 The computing device shown illustrates various embodiments of this disclosure.
[0039] <First Embodiment> Figure 3 A refresh processing method for computing cores according to some embodiments is illustrated, which can be implemented by the task distribution unit of a computing device. For example... Figure 3 As shown, the refresh processing method of this embodiment may include the following steps S310 to S340.
[0040] Step S310: Receive a refresh command for the target task.
[0041] A refresh command can be issued by the control device 1200 (e.g., CPU) to indicate the target task that needs to be stopped and cleaned up. This refresh command can be generated based on user triggering or automatically triggered based on system timeout detection, etc.
[0042] In some examples, the target task is a video processing task, and the computation kernel is a computation kernel that performs video encoding and decoding tasks. Video processing tasks include, but are not limited to, encoding and decoding formats such as H.264, H.265, AV1, AVS2, and VP9. For example, the target task might be to decode an H.264 format video stream used for training a large model to obtain the original data. Another example is to encode the feature map of the inference output in real time to reduce bandwidth. The target task can be a complete video processing task, or one or more sub-tasks within that task; this disclosure does not limit its scope.
[0043] In other examples, the target task may also be a general computational task, which is not limited in this disclosure.
[0044] Step S320: Based on the refresh instruction, check whether the computing core instruction package to be sent to the computing core is the target instruction package corresponding to the target task.
[0045] In this embodiment, the computational core instruction package to be sent to the computational core refers to: instruction packages that have entered the task distribution unit but have not yet been retrieved and executed by the computational core, including computational core instruction packages temporarily stored in the instruction channel and computational core instruction packages that have been written to the computational core instruction cache unit but have not yet been read out. The target instruction package refers to the computational core instruction package used to cause the computational core to perform operations related to the target task. That is, the target instruction package is the computational core instruction package to be sent, belonging to the target task that is the object of refresh. Since the target task has been triggered for refresh, the actual execution result of the target instruction package is no longer needed.
[0046] Combination Figure 2 As shown, an interception point for checking computational core instruction packets can be set between the read ports of multiple computational core instruction cache units 1113 and the receive ports of multiple computational cores 1120; or, the interception point can be set between the read ports of multiple instruction channels 1111 and the receive ports of multiple computational core instruction cache units 1113. In an example where the task distribution unit does not have multiple computational core instruction cache units 1113, the interception point can also be set between the read ports of multiple instruction channels 1111 and the receive ports of multiple computational cores 1120. In this way, the solution of this embodiment can be implemented by adding a refresh processing unit without significantly modifying the original hardware structure.
[0047] Therefore, in some examples, step S320, based on the refresh instruction, checks whether the computing core instruction packet to be sent to the computing core is the target instruction packet corresponding to the target task, may further include the following steps: based on the refresh instruction, intercepting the computing core instruction packet from the instruction stream to be sent; and checking whether the intercepted computing core instruction packet is the target instruction packet corresponding to the target task.
[0048] The instruction stream to be sent refers to the sequence of instruction packets being transmitted to the next level processing unit. The instruction stream to be sent can be a sequence of instruction packets from the instruction channel to the instruction cache unit of the computing core, or a sequence of instruction packets from the instruction cache unit of the computing core to the computing core.
[0049] In some examples, intercepting computation core instruction packets from the instruction stream to be issued may further include: selecting one of a plurality of computation core instruction cache units corresponding one-to-one with a plurality of computation cores; and intercepting and inspecting the computation core instruction packets read from the selected computation core instruction cache unit.
[0050] The criteria for selecting a compute core instruction cache unit can be: the compute core instruction cache unit is not empty, and the compute core pointed to by the compute core instruction packet at the head of the compute core instruction cache unit queue is free.
[0051] by Figure 2 Taking the video processing module as an example, the video processing module contains 4 computing cores, corresponding to 4 independent computing core instruction cache units. In one clock cycle, one computing core instruction packet cache unit is selected from the 4 computing core instruction cache units for reading and sending, and the read computing core instruction packet is intercepted and checked in the current clock cycle or the next clock cycle.
[0052] In this example, the computation core instruction packet read from the selected computation core instruction cache unit is intercepted, and a check is performed on the refresh instruction. That is, this check is performed at the end of the task distribution unit facing the computation core. This can reduce the impact of the refresh check on the intermediate links of task distribution, and at the same time, the existing arbitration result can be used without adding additional multiplexing logic.
[0053] In some examples, the checking operation in step S320 can be performed based on the target task identifier indicated by the refresh command and the task identifier of the task to which the computing core command package belongs, in order to quickly filter out the target command package that needs to be refreshed. This may include: extracting the task identifier from the computing core command package; comparing the extracted task identifier with the target task identifier indicated by the refresh command; and, if the comparison results match, determining that the computing core command package is the target command package, and if the comparison results do not match, determining that the computing core command package is not the target command package.
[0054] In this example, the task identifier can be any type of identifier that reflects the task it belongs to, such as the thread ID created by the task, the task number, or a combination of the thread ID and the indicator channel ID. The corresponding comparison operation can be a strict equality comparison for bit-width matching, or a comparison between specific bit fields.
[0055] The task identifier carried by the compute core instruction packet can be assigned by the driver running on the control device when the task is created, and remains unchanged in all related instruction packets throughout the task's lifecycle. The target task identifier indicated by the refresh instruction is the task identifier of the target task that is being refreshed. When the control device detects a refresh operation, it obtains the target task identifier of the target task corresponding to the refresh operation, generates a refresh instruction for the target task based on the target task identifier, and sends the generated refresh instruction to the compute device for execution. Therefore, when the compute device determines whether the compute core instruction packet to be sent to the compute core is the target instruction packet of the corresponding target task, it only needs to compare the task identifier carried by the compute core instruction packet with the target task identifier indicated by the refresh instruction to determine whether the compute core instruction packet belongs to the target task. Based on this correspondence, the compute device only needs to perform a simple numerical comparison to complete the instruction packet attribution determination, without needing to track other contextual information of the task, thus achieving fast and accurate refresh target identification with low hardware overhead.
[0056] Step S330: If the computing core instruction packet is the target instruction packet, prevent the distribution of the computing core instruction packet to the corresponding computing core, and generate and return an instruction completion response for the blocked computing core instruction packet.
[0057] In this embodiment, preventing the target instruction packet from being sent to the corresponding computing core means that the instruction packet will not enter the computing core's execution pipeline. Generating an instruction completion response for the blocked computing core instruction packet maintains the correctness of the system instruction channel's state machine and ensures normal resource reclamation without actually executing the computing core instruction packet.
[0058] The generation logic can generate a valid response message based on the instruction channel ID and instruction sequence number of the target instruction packet. The format of the instruction completion response generated for the target instruction packet or the instruction packet of the blocked computing core can be consistent with the instruction completion response returned by the real computing core. That is, it can simulate the computing core generating the corresponding instruction completion response for the instruction packet of the blocked computing core. By maintaining a consistent response format, the refresh function can be seamlessly integrated without modifying the design of the existing instruction channel, reducing verification complexity.
[0059] In some examples, to reduce the impact of checking refresh instructions on the incoming instruction stream, the refresh processing unit can set up a refresh processing cache unit and temporarily store the detected computation kernel instruction packets belonging to the target task in the refresh processing cache unit. This allows for concurrent checking and simulation operations within one clock cycle, reducing the time occupied and avoiding blocking the reading of subsequent instruction packets due to the delay in generating simulation responses.
[0060] In these examples, step S330 may further include the following steps: temporarily storing the blocked computation kernel instruction packet in a refresh processing cache unit; reading the temporarily stored computation kernel instruction packet from the refresh processing cache unit; and, for the read computation kernel instruction packet, simulating the corresponding computation kernel to generate and return the corresponding instruction completion response.
[0061] The refresh processing cache unit can be a first-in-first-out (FIFO) queue, used to temporarily store intercepted computation kernel instruction packets, waiting for refresh processing.
[0062] In another example, the refresh processing unit can also generate a response and discard the instruction packet through combinational logic.
[0063] The refresh processing unit can return the simulated instruction completion response according to the original return path of the hardware design to achieve closed-loop control. For example, the instruction completion response can be returned to the instruction channel to which the target instruction packet is assigned. This instruction channel updates the instruction packet count value for the target task. The instruction packet count value represents the number of instruction packets that have been issued from the instruction channel but for which no response has been received. When the instruction packet count value becomes 0, a target task completion message is reported, thereby ending the refresh of the target task. Alternatively, the instruction completion response can be returned to the command processor upstream of the instruction channel. The command processor updates the total instruction packet count value for all issued computing core instruction packets. When the total instruction packet count value becomes 0, a task completion message is reported to the control device.
[0064] In some examples, step S330, generating and returning the instruction completion response for the computing core instruction packet, may include the following steps: generating a corresponding instruction completion response for the blocked computing core instruction packet; returning the generated instruction completion response to the instruction channel corresponding to the computing core instruction packet, triggering the instruction channel to quickly release the occupied resources based on the instruction completion response. The resources that can be released include, for example, the corresponding storage entries in the internal queue of the instruction channel.
[0065] In some examples, when all computational core instruction packets of the target task have generated instruction completion responses, the refresh operation for the target task is determined to be complete, and the check operation in step S320 ends, releasing the resources occupied by the check operation and resuming the normal instruction packet delivery process. In this example, timely termination of the refresh process by checking the instruction completion responses can avoid the impact of invalid check operations on the normal instruction packet delivery process. For example, the instruction packet count value maintained by the instruction channel can be used to determine whether instruction completion responses for all computational core instruction packets of the target task have been received. When the instruction packet count value becomes 0, it means that the processing of all computational core instruction packets of the target task has been completed.
[0066] In some examples, the task distribution unit supports parallel processing of multiple refresh instructions. In this case, checks can be performed synchronously on multiple refresh instructions to further improve refresh processing efficiency. In this example, there are multiple refresh instructions, each corresponding to a different target task. Therefore, step S320, based on the refresh instruction, checks whether the computation core instruction packet to be sent to the computation core is the target instruction packet for the corresponding target task. This includes checking whether the instruction packet to be sent is the target instruction packet for each target task indicated by the refresh instruction. Step S330, if the computation core instruction packet is a target instruction packet, prevents the sending of the computation core instruction packet to the corresponding computation core. This includes preventing the sending of the computation core instruction packet to the corresponding computation core if the computation core instruction packet is the target instruction packet corresponding to any target task. In this example, the sending is prevented when the checked computation core instruction packet matches any refresh task.
[0067] In this example, the refresh processing unit can also employ time-division multiplexing, using a single comparator to sequentially compare each target task identifier. Alternatively, the refresh processing unit can maintain a list of target task identifiers and use multiple comparators to compare the task identifier of the instruction packet being inspected with all target task identifiers in the list in parallel, thereby reducing inspection latency.
[0068] Based on steps S310 to S330 above, this embodiment actively checks the computational core instruction packets belonging to the aborted task, preventing them from entering the execution pipeline, and replaces actual execution with simulated instruction completion responses. This skips the actual execution process of invalid instructions while maintaining the integrity of the hardware handshake protocol. Compared to traditional solutions that must wait for all instruction packets to complete naturally, this embodiment significantly reduces the time overhead of refresh operations, improves system response speed and resource utilization efficiency, and is particularly suitable for high-concurrency scenarios such as real-time multimedia processing and AI inference services that require high-frequency task switching.
[0069] In other embodiments, such as Figure 4 As shown, after step S320 above, the refresh processing method may further include the following step S340: if the computing core instruction packet is not the target instruction packet, the computing core instruction packet is sent to the corresponding computing core.
[0070] In this embodiment, non-target tasks can remain unaffected by the refresh and maintain their normal execution flow.
[0071] Taking the video processing module as an example again, the improved video processing module is as follows: Figure 5As shown, the video processing module adds a refresh processing unit 1115, which can be located between the multiple computing core instruction packet cache selection unit 1114 and the multiple computing cores 1120. The computing core instruction packet read from the multiple computing core instruction packet cache unit 1114 is checked by the refresh processing unit 1115 to see if it is a target instruction packet. If it is determined to be a target instruction packet, it prevents the packet from being sent to the computing core, simulates the computing core to generate an instruction completion response, and returns the simulated instruction completion response to the instruction channel allocated to the computing core instruction packet. If it is not a target instruction packet, it releases the computing core instruction packet and allows it to be processed by the corresponding computing core.
[0072] <Second Embodiment> This disclosure also provides a computing device, such as a computing device that is Figure 1 The computing device 1100 is configured, in some embodiments, to perform a refresh processing method according to any embodiment of the present disclosure.
[0073] In other embodiments, such as Figure 6 As shown, the computing device 1100 may include a refresh instruction receiving unit 610, an instruction packet checking unit 620, and an instruction packet processing unit 630. The refresh instruction receiving unit is configured to receive refresh instructions for a target task. The instruction packet checking unit 620 is configured to check, based on the refresh instruction received by the instruction receiving unit 610, whether a computing core instruction packet to be sent to the computing core is a target instruction packet. The instruction packet processing unit 630 is configured to, when the checked computing core instruction packet is a target instruction packet, prevent the sending of the computing core instruction packet to the computing core, and simulate the computing core to generate an instruction completion response.
[0074] In some examples, the instruction packet processing unit 630 may also be configured to send the instruction to the corresponding computing core for execution when the instruction packet being checked is not the target instruction packet.
[0075] The refresh instruction receiving unit 610, instruction packet checking unit 620, and instruction packet processing unit 630 in this embodiment can be components of the refresh processing unit. By adding this refresh processing unit to the task distribution unit of the computing device, a fast response to refresh operations can be achieved.
[0076] <Third Embodiment> This disclosure also provides a computing system, which includes a control device and a computing device. The computing device can be the computing device according to the second embodiment. The control device is used to generate a sequence of computing core instruction packets corresponding to a task oriented towards the computing device, and to sequentially send these computing core instruction packets to a task distribution unit, which then allocates the computing core instruction packets to the corresponding computing cores for execution.
[0077] The computing device of the computing system is, for example, Figure 1 The computing device 1100 and the control device are, for example, Figure 1 The control device 1200 in the middle.
[0078] The form of the computing system can be a chip, module, terminal device, workstation or server, etc., without limitation.
[0079] This disclosure also provides a computing system, which includes a control device and a computing device. The computing device may be the computing device described in the second embodiment. The control device generates a corresponding sequence of computing core instruction packages for each task targeting the computing device, and sequentially distributes the sequence of computing core instruction packages to a task distribution unit, which then allocates each computing core instruction package to the corresponding computing core for execution.
[0080] As an example, the computing device in this computing system could be Figure 1 The computing device 1100 and the control device can be Figure 1 The control device 1200 in the middle.
[0081] The computing system can be any type of physical entity, such as, but not limited to, a chip, module, user terminal, workstation, or server, etc., and this disclosure does not limit it.
[0082] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and apparatuses according to various embodiments of this specification. In this regard, each block in a flowchart or block diagram may represent a module, unit, or part of a circuit. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using hardware that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. Unless otherwise specified, implementation in hardware, implementation in software, and implementation using a combination of software and hardware can be equivalent.
[0084] Various embodiments of this specification have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A refresh processing method for computing cores, characterized in that, include: Receive refresh instructions for the target task; Based on the refresh instruction, check whether the computing core instruction package to be sent to the computing core is the target instruction package corresponding to the target task; If the computing core instruction packet is the target instruction packet, prevent the computing core instruction packet from being sent to the corresponding computing core, and generate and return an instruction completion response for the blocked computing core instruction packet.
2. The method according to claim 1, characterized in that, After checking whether the computing core instruction package to be sent to the computing core is the target instruction package corresponding to the target task, the method further includes: If the computing core instruction package is not the target instruction package, the computing core instruction package is sent to the corresponding computing core.
3. The method according to claim 1, characterized in that, The generation and return of the instruction completion response for the blocked computation kernel instruction packet includes: The simulated computing core generates and returns an instruction completion response for the instruction packet of the blocked computing core.
4. The method according to claim 1, characterized in that, The step of checking whether the computing core instruction package to be sent to the computing core is the target instruction package corresponding to the target task, based on the refresh instruction, includes: Based on the refresh instruction, intercept the computing core instruction packet from the instruction stream to be issued; Check whether the intercepted computing kernel instruction packet is the target instruction packet of the corresponding target task.
5. The method according to claim 4, characterized in that, The interception of computation kernel instruction packets from the instruction stream to be issued includes: Select one of the multiple compute core instruction cache units that correspond one-to-one with multiple compute cores; Intercept computation core instructions read from the selected computation core instruction cache unit.
6. The method according to claim 1, characterized in that, The computation kernel instruction package carries the task identifier of the task to which it belongs; The step of checking whether the computing core instruction package to be sent to the computing core is the target instruction package corresponding to the target task includes: Extract the task identifier from the computational core instruction package; The extracted task identifier is compared with the target task identifier indicated by the refresh command; If the comparison results are consistent, the computational kernel instruction package is determined to be the target instruction package.
7. The method according to any one of claims 1 to 6, characterized in that, The step of preventing the distribution of the computing core instruction packet to the corresponding computing core, and generating and returning an instruction completion response for the blocked computing core instruction packet, includes: The instruction packets of the blocked computing kernels will be temporarily stored in the refresh processing cache unit; Read the temporarily stored computation kernel instruction packet from the refresh processing cache unit; For the read computational core instruction packet, generate and return the corresponding instruction completion response.
8. The method according to any one of claims 1 to 6, characterized in that, The generation and return of the instruction completion response for the blocked computation kernel instruction packet includes: Generate corresponding instruction completion responses for the blocked computation kernel instruction packets; The generated instruction completion response is returned to the instruction channel corresponding to the blocked computing core instruction, triggering the instruction channel to release the resources occupied by the blocked computing core instruction package based on the instruction completion response.
9. The method according to any one of claims 1 to 6, characterized in that, The method further includes: When all computational core instruction packets of the target task generate corresponding instruction completion responses, the refresh operation for the target task is determined to be complete, and the check operation ends.
10. The method according to any one of claims 1 to 6, characterized in that, The refresh instructions are multiple, and different refresh instructions correspond to different target tasks; the step of checking whether the computing core instruction package to be sent to the computing core is the target instruction package corresponding to the target task based on the refresh instructions includes: For each refresh instruction, check whether the computing core instruction packet to be sent to the computing core is the target instruction packet corresponding to the target task. If the computation core instruction packet is the target instruction packet, prevent the computation core instruction packet from being sent to the corresponding computation core, including: If the computing core instruction package is a target instruction package corresponding to any target task, prevent the computing core instruction package from being sent to the corresponding computing core.
11. The method according to any one of claims 1 to 6, characterized in that, The target task is a video processing task, and the computing core is a computing core that performs video encoding and decoding tasks.
12. A computing device, characterized in that, include: Multiple computing cores; as well as, A task distribution unit configured to perform the method according to any one of claims 1 to 11.
13. A computing device, characterized in that, include: The refresh instruction receiving unit is configured to receive refresh instructions for the target task; The instruction packet checking unit is configured to check, based on the refresh instruction, whether the computing core instruction packet to be sent to the computing core is the target instruction packet corresponding to the target task; as well as, The instruction packet processing unit is configured to, when the computing core instruction packet is the target instruction packet, prevent the computing core instruction packet from being sent to the corresponding computing core, and generate and return an instruction completion response for the blocked computing core instruction packet.
14. A computing system, characterized in that, include: A control device and a computing device, wherein the computing device is the computing device according to claim 12 or 13, and the control device is configured to send computing core instruction packets to the computing device for the required processing task.