Command processing apparatus, method, electronic device, and computer-readable storage medium

By dividing command processing operations into multiple processing blocks with fine granularity and utilizing time-division multiplexing technology, the problem of low efficiency in synchronous command execution between multiple processes is solved, achieving more efficient command processing.

CN114816777BActive Publication Date: 2025-11-21SHANGHAI POWERTENSORS INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110130200.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-29
Publication Date
2025-11-21
Estimated Expiration
2041-01-29

AI Technical Summary

Technical Problem

Existing technologies that use time-division multiplexing to achieve synchronous execution of commands between multiple processes suffer from low processing efficiency, especially the latency caused by waiting for incomplete commands during context switching.

Method used

The command processing operation is divided into multiple processing blocks, and the starting processing block in the current time slice is determined by the first processing block identifier. Time-division multiplexing is used for context switching to reduce waiting time and improve processing efficiency.

Benefits of technology

By dividing processing blocks into fine-grained segments and using time-division multiplexing, context switching waiting time is reduced, command processing efficiency is improved, and the continuity and efficiency of command processing are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114816777B_ABST
    Figure CN114816777B_ABST
Patent Text Reader

Abstract

The present disclosure provides a command processing device, method, electronic equipment and computer readable storage medium, wherein the device comprises: a microcontroller and an operation unit; the microcontroller is configured to obtain a to-be-executed command after a current time slice arrives; the to-be-executed command carries a first processing block identifier used to indicate a processing block corresponding to the current time slice in a plurality of processing blocks corresponding to the to-be-executed command; and the operation unit is configured to obtain the to-be-executed command and execute a processing task corresponding to the to-be-executed command based on the first processing block identifier. With the command processing device, the efficiency of command processing can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer science, and in particular, to a command processing apparatus and method, an electronic device, and a computer readable storage medium. BACKGROUND

[0002] In the field of computer science, by means of virtualization, physical resources can be converted into logically manageable resources to improve the utilization of physical resources of a server. Currently, when implementing virtualization by deploying a graphics processing unit (GPU) or an artificial intelligence (AI) chip, a time division multiplexing method is usually used to realize the synchronous execution of commands in multiple processes. The current command processing method has the problem of low processing efficiency. SUMMARY

[0003] The embodiments of the present disclosure at least provide a command processing apparatus, method, electronic device, and computer readable storage medium.

[0004] In a first aspect, the embodiments of the present disclosure provide a command processing apparatus, comprising: a microcontroller and an operation unit; wherein the microcontroller is configured to obtain a to-be-executed command after a current time slice arrives; the to-be-executed command carries a first processing block identifier indicating a processing block corresponding to the current time slice among a plurality of processing blocks corresponding to the to-be-executed command; and the operation unit is configured to obtain the to-be-executed command and execute a processing task corresponding to the to-be-executed command based on the first processing block identifier.

[0005] In this way, by dividing a plurality of processing operations corresponding to a command into a plurality of processing blocks, the starting processing block to be executed in the current time slice can be determined based on the first processing block identifier determined based on the processing result of the to-be-executed command in the previous corresponding time slice, and the processing task of the to-be-executed command can be executed based on the starting processing block. Since the plurality of processing operations are divided into a plurality of processing blocks of finer granularity, when context switching is performed by means of time division multiplexing, the waiting time caused by the to-be-executed command not being processed can be reduced, thereby improving the efficiency of command processing.

[0006] In an alternative implementation, the processing device further comprises a command distributor; the microcontroller is configured to acquire a to-be-executed command after a current time slice arrives, and store the to-be-executed command in a command queue; the command distributor is configured to acquire the to-be-executed command from the command queue, and distribute the to-be-executed command to the operation unit; and the operation unit is configured to acquire the to-be-executed command, and execute a processing task corresponding to the to-be-executed command based on the first processing block identifier after receiving the to-be-executed command distributed by the command distributor.

[0007] In an alternative implementation, the operation unit is configured to determine a starting processing block to be executed in the current time slice from the plurality of processing blocks based on the first processing block identifier, and execute the processing task corresponding to the to-be-executed command based on the starting processing block.

[0008] In this way, since the first processing block identifier can indicate the first processing block that has not been processed from the plurality of processing blocks included in the to-be-executed command, the operation unit can focus on the first processing block that needs to be processed based on the first processing block identifier, without searching the plurality of processing blocks in the to-be-executed command to determine the processing block that needs to be processed, thereby effectively improving the efficiency of command processing.

[0009] In an alternative implementation, the microcontroller is configured to read the to-be-executed command from a target cache corresponding to the current time slice, or acquire the to-be-executed command from a host.

[0010] In this way, the microcontroller can flexibly acquire the to-be-executed command from the target cache or the host according to different time slices and execution conditions of different to-be-executed commands.

[0011] In an alternative implementation, the microcontroller is configured to determine whether the command queue is idle, listen to a buffer in a case where the command queue is idle, the buffer being configured to store the to-be-executed command by the host, and read the to-be-executed command from the buffer in a case where the to-be-executed command exists in the buffer.

[0012] In this way, by determining whether the command queue is idle and whether the to-be-executed command exists in the buffer in a case where the command queue is idle, the to-be-executed command can be continuously issued to the command queue in a case where the buffer continuously receives the to-be-executed command, so that the processing device can maintain high efficiency in processing the continuously issued to-be-executed command.

[0013] In an alternative implementation, the buffer comprises a ring buffer; the ring buffer has a plurality of entries; different entries are used to store to-be-executed commands of different command streams; and the microcontroller is configured to: determine a target entry in the buffer based on a command stream corresponding to a current time slice; and listen to whether the to-be-executed command is stored in the buffer based on the determined target entry.

[0014] In this way, the commands of different command streams are stored in the ring buffer by using the target entries corresponding to the command streams, so that the microcontroller can synchronously listen to a plurality of entries of the ring buffer, and the commands in different command streams can be pulled down in one processing period, thereby improving the efficiency of command acquisition. Meanwhile, the ring buffer can provide mutual exclusion access to the buffer for the communication program, which is conducive to avoiding the increased system overhead of the storage queue when frequent command allocation is used.

[0015] In an alternative implementation, the microcontroller is configured to: determine whether a to-be-executed command exists in a target cache corresponding to a current time slice; and acquire the to-be-executed command from the host in a case where the to-be-executed command does not exist in the target cache corresponding to the current time slice.

[0016] In this way, after the to-be-executed command in the target cache is executed, a new to-be-executed command can be acquired from the host, so that the command processing device can work dynamically and process the newly issued to-be-executed command more smoothly, thereby improving the efficiency of the command processing device when processing the commands.

[0017] In an alternative implementation, the microcontroller is further configured to: read the to-be-executed command from the target cache in a case where the to-be-executed command exists in the target cache corresponding to the current time slice.

[0018] In this way, since the to-be-executed command that has not been executed is stored in the target cache, the microcontroller can read the to-be-executed command from the target cache in a case where the to-be-executed command exists in the target cache corresponding to the current time slice, thereby continuing to process the to-be-executed command, so that the command processing device can process the to-be-processed commands in order according to the order of the to-be-processed commands, and can avoid overloading of the target cache caused by accumulation of to-be-executed commands that have not been executed corresponding to the same user.

[0019] In an alternative implementation, the operation unit is further configured to report a second processing block identifier corresponding to a processing block that has been executed most recently to the microcontroller after the current time slice ends; and the microcontroller is further configured to update the to-be-executed command in the command queue based on the second processing block identifier reported by the operation unit after receiving the second processing block identifier.

[0020] In an alternative implementation, the operation unit is configured to: after a processing task corresponding to a currently-executing processing block is executed, take the currently-executing processing block as the most-recently-executed processing block, and report a second processing block identifier corresponding to the most-recently-executed processing block to the microcontroller.

[0021] In this way, the operation unit can quickly determine the second processing block identifier corresponding to the most-recently-executed processing block after processing the to-be-executed command in the current time slice, and thus the operation unit can more conveniently report the second processing block identifier to the microcontroller. Meanwhile, the microcontroller can update the to-be-executed command in the command queue after receiving the second processing block identifier reported by the operation unit, so that the to-be-executed command after the update can be directly processed according to the updated to-be-executed command when the to-be-executed command is processed in the next corresponding time slice.

[0022] In an alternative implementation, the operation unit is configured to: send the second processing block identifier to the command distributor; and the command distributor is further configured to send the second processing block identifier to the microcontroller.

[0023] In this way, the operation unit can send the second processing block identifier to the microcontroller by using the communication channel between the operation unit and the command distributor, and the communication channel between the command distributor and the microcontroller, so that the communication channel between the operation unit and the microcontroller can not be established, and the data interface of the microcontroller can be less occupied.

[0024] In an alternative implementation, the microcontroller is configured to: in a case where the second processing block identifier is a processing block identifier of a last processing block in the plurality of processing blocks, delete the to-be-executed command from the command queue; and in a case where the second processing block identifier is not a processing block identifier of a last processing block in the plurality of processing blocks, determine a target processing block identifier based on the second processing block identifier, replace a first processing block identifier in the to-be-executed command with the target processing block identifier, and generate a new to-be-executed command; wherein the target processing block identifier is a processing block identifier of a next processing block of the most-recently-executed processing block.

[0025] In this way, the microcontroller can more accurately and easily learn the processing status of the to-be-executed command according to the second processing block identifier. In a case where the second processing block identifier is a processing block identifier of a last processing block in the plurality of processing blocks, the microcontroller can determine that the to-be-executed command has been processed, and delete the to-be-executed command from the command queue, so that a possible running error of repeatedly processing the to-be-executed command in the command queue can be effectively avoided.

[0026] In an alternative implementation, the microcontroller is further configured to store the new to-be-executed command into a target cache corresponding to the to-be-executed command after generating the new to-be-executed command.

[0027] In this way, the new to-be-executed command can be stored into the target cache, and a processing scheme corresponding to the new to-be-executed command does not need to be generated again, but the new to-be-executed command can be processed by using the method for processing any to-be-executed command in the target cache, so that the work of the microcontroller is more simple and stable, and thus the command processing device is more stable when processing commands, and the probability of running errors is reduced.

[0028] In a second aspect, the embodiments of the present disclosure further provide another command processing device, which comprises a microcontroller and an operation unit; the operation unit is configured to report a first processing block identifier of a current processing block of a target command executed in a current time slice to the microcontroller in response to the end of the current time slice; the current processing block is any one of at least one processing block in the target command; and the microcontroller is configured to update the target command by using the first processing block identifier after receiving the first processing block identifier reported by the operation unit.

[0029] In a third aspect, the embodiments of the present disclosure further provide a command processing method applied to a command processing device, which comprises a microcontroller and an operation unit; the command processing method comprises the following steps: the microcontroller acquires a to-be-executed command after a current time slice arrives; the to-be-executed command carries a first processing block identifier used to indicate a processing block corresponding to the current time slice in a plurality of processing blocks corresponding to the to-be-executed command; and the operation unit acquires the to-be-executed command, and executes a processing task corresponding to the to-be-executed command based on the first processing block identifier.

[0030] In an alternative implementation, the command processing device further comprises a command distributor; the microcontroller acquires a to-be-executed command after a current time slice arrives, which comprises the following steps: the microcontroller acquires a to-be-executed command after a current time slice arrives, and stores the to-be-executed command into a command queue; the command processing method further comprises the following steps: the command distributor acquires the to-be-executed command from the command queue, and distributes the to-be-executed command to the operation unit; and the operation unit acquires the to-be-executed command, and executes a processing task corresponding to the to-be-executed command based on the first processing block identifier after receiving the to-be-executed command distributed by the command distributor.

[0031] In an alternative implementation, the operation unit determines, based on the first processing block identifier, a starting processing block to be executed in the current time slice from the plurality of processing blocks, and executes the processing task corresponding to the to-be-executed command based on the starting processing block.

[0032] In an alternative implementation, the microcontroller obtains the to-be-executed command after the current time slice arrives, including: the microcontroller reads the to-be-executed command from a target buffer corresponding to the current time slice; or obtains the to-be-executed command from the host.

[0033] In an alternative implementation, the microcontroller obtains the to-be-executed command after the current time slice arrives, including: the microcontroller determines whether the command queue is idle; in the case that the command queue is idle, the microcontroller listens to a buffer; the buffer is used by the host to store the to-be-executed command; in the case that the to-be-executed command is listened to in the buffer, the microcontroller reads the to-be-executed command from the buffer.

[0034] In an alternative implementation, the buffer includes a ring buffer; the ring buffer has a plurality of entries; different entries are used by the host to store to-be-executed commands of different command streams; the microcontroller obtains the to-be-executed command after the current time slice arrives, including: the microcontroller determines a target entry from the buffer based on a command stream corresponding to the current time slice; and listens to whether the to-be-executed command is stored in the buffer based on the determined target entry.

[0035] In an alternative implementation, the microcontroller obtains the to-be-executed command after the current time slice arrives, including: the microcontroller determines whether the to-be-executed command exists in a target buffer corresponding to the current time slice; in the case that the to-be-executed command does not exist in the target buffer corresponding to the current time slice, the microcontroller obtains the to-be-executed command from the host.

[0036] In an alternative implementation, the microcontroller obtains the to-be-executed command after the current time slice arrives, including: the microcontroller reads the to-be-executed command from a target buffer corresponding to the current time slice in the case that the to-be-executed command exists in the target buffer corresponding to the current time slice.

[0037] In an alternative implementation, further including: the operation unit reports a second processing block identifier corresponding to a processing block that has been executed most recently to the microcontroller after the current time slice ends; and the microcontroller updates the to-be-executed command in the command queue based on the second processing block identifier reported by the operation unit after receiving the second processing block identifier.

[0038] In an alternative implementation, the operation unit reports a second processing block identifier corresponding to a most recently executed processing block to the microcontroller after the end of the current time slice, including: after the operation unit finishes executing a processing task corresponding to a currently executing processing block, the operation unit takes the currently executing processing block as the most recently executed processing block, and reports a second processing block identifier corresponding to the most recently executed processing block to the microcontroller.

[0039] In an alternative implementation, the operation unit sends the second processing block identifier to the command distributor, and the command distributor sends the second processing block identifier to the microcontroller.

[0040] In an alternative implementation, the updating of the to-be-executed command in the command queue based on the second processing block identifier includes: in a case where the second processing block identifier is a processing block identifier of a last processing block in the plurality of processing blocks, the microcontroller deletes the to-be-executed command from the command queue; in a case where the second processing block identifier is not a processing block identifier of a last processing block in the plurality of processing blocks, the microcontroller determines a target processing block identifier based on the second processing block identifier, replaces a first processing block identifier in the to-be-executed command with the target processing block identifier, and generates a new to-be-executed command; the target processing block identifier is a processing block identifier of a next processing block of the most recently executed processing block.

[0041] In an alternative implementation, the microcontroller stores the new to-be-executed command in a target buffer corresponding to the to-be-executed command after generating the new to-be-executed command.

[0042] In a fourth aspect, the embodiments of the present disclosure further provide a command processing method applied to a command processing device, the command processing device including a microcontroller and an operation unit, and the command processing method including: the operation unit reporting a first processing block identifier of a current processing block of a target command executed in a current time slice to the microcontroller in response to the end of the current time slice; the current processing block being any one of at least one processing block in the target command; and the microcontroller updating the target command by using the first processing block identifier after receiving the first processing block identifier reported by the operation unit.

[0043] In a fifth aspect, the embodiments of the present disclosure further provide an electronic device including a host, a buffer, and a command processing device; the host is configured to issue a to-be-executed command and store the to-be-executed command in the buffer; and the command processing device is configured to execute the method in any one of the third aspect or the fourth aspect.

[0044] In a sixth aspect, the embodiments of the present disclosure further provide a computer readable storage medium, which stores a computer program, and the program is executed by a microcontroller and an operation unit to implement the method in any of the embodiments of the third aspect or the fourth aspect.

[0045] The effects of the command processing method are described in the description of the command device, and will not be repeated here.

[0046] In order to make the above objectives, characteristics and advantages of the present disclosure more obvious and easy to understand, the following preferred embodiments are specifically described below, and the accompanying drawings are referred to for a detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings needed to be used in the embodiments. The drawings incorporated into the description and form a part of the description, which show the embodiments consistent with the present disclosure, and are used to explain the technical solutions of the present disclosure together with the description. It should be understood that the following drawings only show some of the embodiments of the present disclosure, and therefore should not be considered as a limitation to the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor.

[0048] Figure 1 A command processing device provided by the embodiments of the present disclosure is shown;

[0049] Figure 2 A state diagram of a command queue provided by the embodiments of the present disclosure is shown;

[0050] Figure 3 A schematic diagram of a command processing device corresponding to a specific process example of command processing provided by the embodiments of the present disclosure is shown;

[0051] Figure 4 A flowchart of a command processing method provided by the embodiments of the present disclosure is shown;

[0052] Figure 5 A flowchart of another command processing method provided by the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0053] To make the purposes, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. The components of the embodiments of the present disclosure described and shown herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the claimed present disclosure, but only represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present disclosure.

[0054] It is found through research that when processing commands in the command streams of different users in a time division multiplexing manner, a time slice is usually set, commands in the command stream of a user are executed in the current time slice, and after the end of the time slice, the context of different users is switched, that is, from the current user to the next user, and the next time slice executes the commands in the command stream corresponding to the next user. Since at the end of the time slice, there may be a case that the command corresponding to the current user is being executed, when the context is switched, the executing command needs to be executed completely before switching to the next user, resulting in a long time delay when the context is switched, so that the command corresponding to the next user needs to wait for a long time before being processed, causing the problem of low efficiency in processing commands.

[0055] Based on the above research, the present disclosure provides a command processing apparatus and method, an electronic device, and a computer readable storage medium, which divide a plurality of processing operations corresponding to a command into a plurality of processing blocks, when it is the turn of a to-be-executed command to be executed, a starting processing block to be executed in the current execution period is determined according to a first processing block identifier determined in the last execution period when the to-be-executed command is executed, and a processing task of the to-be-executed command is executed based on the starting processing block, so as to divide the command into a finer granularity, and perform time division multiplexing according to the finer granularity when time division multiplexing, thereby reducing the waiting time required for context switching and improving the efficiency of processing commands.

[0056] The defects of the above solutions are the results of the inventors after practice and careful research, therefore, the discovery process of the above problems and the solutions proposed by the present disclosure to solve the above problems in the following should be the contributions of the inventors to the present disclosure in the process of the present disclosure.

[0057] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.

[0058] For the convenience of understanding the present embodiment, first, a command processing method disclosed by the present embodiment is introduced in detail.

[0059] The command processing device provided by the present embodiment can be applied to a GPU, an artificial intelligence chip, or other command processing devices including a command processor and an execution unit.

[0060] The command processing device provided by the present embodiment is described below by taking the application of the command processing device to a GPU as an example.

[0061] Referring to Figure 1 The command processing device provided by the present embodiment includes a microcontroller 10 and an operation unit 20.

[0062] The microcontroller 10 is configured to acquire a to-be-executed command after a current time slice arrives.

[0063] The operation unit 20 is configured to acquire the to-be-executed command and execute a processing task corresponding to the to-be-executed command based on the first processing block identifier.

[0064] In the present embodiment, the microcontroller 10 acquires a to-be-executed command after a current time slice arrives. In the to-be-executed command, a first processing block identifier is carried to indicate a processing block corresponding to the current time slice in a plurality of processing blocks corresponding to the to-be-executed command. In the process, a plurality of processing operations corresponding to the command are divided into a plurality of processing blocks, and the first processing block identifier is used to indicate the processing block to be executed. Therefore, in the execution of the processing task of the to-be-executed command, the command can be divided in a finer granularity, and time division multiplexing can be performed in a finer granularity in time division multiplexing. In one embodiment, the division of the plurality of processing blocks can refer to the division of a thread block in a unified computing device architecture (CUDA, Compute Unified Device Architecture). The processing block can be equivalent to the thread block.

[0065] In another embodiment of the present disclosure, the command distributor 30 is further included; and the microcontroller 10, after obtaining the to-be-executed command, can further store the to-be-executed command into a command queue. The command distributor 30 obtains the to-be-executed command from the command queue and distributes the to-be-executed command to the operation unit. Here, when it is the turn of a to-be-executed command to be executed, the command distributor 30 obtains the to-be-executed command stored into the command queue by the microcontroller 10 from the command queue and distributes the to-be-executed command to the operation unit, and the operation unit executes the to-be-executed command based on the first processing block identifier in the to-be-executed command.

[0066] The microcontroller 10, the operation unit 20 and the command distributor 30 are described in detail below.

[0067] The user in the embodiment of the present disclosure may, for example, include any one of the following: a virtual machine, a computer container, an application program, or different functions in an application program.

[0068] Taking the user as an application program for example, the application program can generate a main process and multiple sub-processes for executing different processing tasks when the application program is run; each process generates a corresponding command stream when executing the corresponding processing task. Further, an application program generates multiple command streams when the application program is run; each command stream includes at least one command; and the command processing apparatus provided by the embodiment of the present disclosure performs time-sliced processing on different command streams of different application programs based on time division multiplexing technology.

[0069] Taking the user as different functions in an application program for example, the application program can generate multiple processes when the application program is run; each process can include multiple threads; each thread can execute a corresponding processing task and generate a command stream corresponding to the thread; and further, different functions of the application program can be respectively taken as different users, and the command processing apparatus provided by the embodiment of the present disclosure performs time-sliced processing on different functions in the same application program based on time division multiplexing technology.

[0070] Taking the user as a virtual machine for example, the application layer of the virtual machine can simultaneously exist one or more command streams stream; each stream includes at least one command, and the commands in the stream are, for example, issued from a host to a buffer corresponding to the user.

[0071] Taking a user (i.e. a virtual machine, which can be represented as VM1 for example) as an example, when a software is running in the user VM1, one or more software functions of the software can be run; when the user is running only one software function of the software, the host can generate a thread for running the software function, the thread generates a command stream stream when performing a processing task, and the CPU of the user stores the commands in the stream into a buffer corresponding to the user VM1.

[0072] In a specific implementation, a software function corresponding to the virtual machine VM1 can include, for example, an image processing function of implementing target recognition of a target object in an image or implementing pose determination of the target object in the image, or a general data processing task of storing or calling an image. The software function implemented by the virtual machine can be determined according to actual conditions, and will not be described here.

[0073] The to-be-executed command is any command in the command stream. Taking the virtual machine VM1 as an example, the virtual machine VM1 can perform image processing on different images such as Pic1, Pic2, etc. At this time, one of the multiple command streams stream corresponding to the virtual machine VM1 can include, for example, multiple to-be-executed commands (Kernel) for performing image processing on the image Pic1, such as convolution operation, pooling operation, full connection operation, etc.

[0074] The to-be-executed command includes address information for pointing to an operand, indication information for indicating a processing task of a specific operator, and a first processing block identifier for indicating any processing block of multiple processing blocks corresponding to the to-be-executed command.

[0075] The address information can be used by the operation unit 20 to obtain the operand when executing the to-be-executed command, and the operand can be processed according to the indication information. When the operand is processed, the process can be divided into multiple processing operations. Different processing operations correspond to different data (such as different pixel points in an image) in the operand; or different processing operations include different operation processes (such as first calculating the product of the pixel value of a pixel point and a weight parameter, and then summing the product results of different pixel points and the weight parameter) on the operand.

[0076] The multiple processing operations included in a to-be-executed command are divided into multiple processing blocks (Block). In an implementation, the to-be-executed command is executed through a thread block, each thread block includes multiple threads, and after the processing operations in all processing blocks are executed, the processing of the to-be-executed command is completed.

[0077] The first processing block identifier included in the to-be-executed command is used to point to any one of the processing blocks in the to-be-executed command which have not been processed. In an embodiment, the first processing block identifier can point to any one of the processing blocks in the to-be-executed command which have not been processed, without any requirement on the execution order of the processing blocks for the execution process. For the execution process which has a requirement on the execution order of the processing blocks, the first processing block identifier can point to the first processing block which needs to be executed among the processing blocks in the to-be-executed command which have not been processed.

[0078] For example, the to-be-executed command K1 can include a convolution operation on the image Pic1. In the case that the size of the image Pic1 is 128*128, the convolution operation on each pixel in the image Pic1 can be regarded as a processing operation, i.e., the to-be-executed command K1 includes 128*128 processing operations, and each 32*32 processing operations is divided into a processing block, i.e., the to-be-executed command includes 4*4 processing blocks. The 16 processing blocks can be sequentially numbered, e.g., numbered as 0-15. In an embodiment, each 32*32 processing operations is assigned to a thread block, and the threads in the thread block perform the operation, i.e., the to-be-executed command is assigned to 4*4 thread blocks to perform the operation.

[0079] In the first time slice for processing the execution command K1, the microcontroller 10 can obtain the to-be-executed command from the command stream issued by the host. At this time, the first processing block identifier in the to-be-executed command can be in a default state, indicating that the processing of the to-be-executed command K1 needs to start from the first processing block, i.e., the processing block numbered 0.

[0080] In the time slice other than the first time slice for processing the to-be-executed command K1, since part of the processing blocks in the to-be-executed command K1 have been processed, the first processing block identifier corresponding to the to-be-executed command K1 at this time points to any one of the processing blocks in the to-be-executed command which have not been processed. For example, if the processing blocks numbered 0-12 have been processed before the current time slice, the first processing block identifier corresponding to the to-be-executed command K1 in the current time slice is 13.

[0081] Here, the processing tasks corresponding to different virtual machines can be the same or different. For the same virtual machine, when performing the same processing task on different images, the image attributes such as the pixel size of the image can not be consistently limited, i.e., different images Pic1, Pic2, etc. can have different pixel sizes.

[0082] When acquiring a command to be executed, the microcontroller 10 reads the command from the target cache corresponding to the current time slice, or acquires the command from the host. Specifically, the methods by which the microcontroller 10 acquires the command to be executed include the following (a1) and (a2):

[0083] (a1): When the microcontroller obtains commands to be executed from the host, the microcontroller 10 determines whether the command queue is idle. If the command queue is idle, the microcontroller 10 listens to the buffer and determines the target entry from the buffer based on the command stream corresponding to the current time slice; based on the determined target entry, it listens to whether the buffer stores commands to be executed.

[0084] There are multiple command queues, each corresponding to a command stream. Different command streams from different users can share a single command queue at different times.

[0085] Here, the command queue can be a software queue or a hardware queue. When the command queue is a software queue, the microcontroller 10 can control the generation and deletion of the command queue. When the microcontroller 10 obtains a command to be executed, it can determine the number of command streams corresponding to the user. If the number of command queues already generated is less than the number of command streams corresponding to the user, the microcontroller 10 generates a new command queue, making the number of command queues greater than or equal to the number of command streams for that user. If a command queue has not been used for a long time, the microcontroller 10 can delete the command queue.

[0086] For example, when the maximum number of command streams corresponding to virtual machines VM1 and VM2 is the same, such as when the maximum number of command streams corresponding to virtual machines VM1 and VM2 is 6, it can be determined that there are 6 command queues.

[0087] When the maximum number of command streams corresponding to virtual machines VM1 and VM2 is different, for example, the maximum number of command streams corresponding to VM1 is 10 and the maximum number of command streams corresponding to VM2 is 6, then based on the comparison result of the maximum number of command streams corresponding to VM1 and VM2, it can be determined that there are 10 command queues. In this case, for virtual machine VM2, the corresponding command may only occupy 6 command queues in the command queue, and the remaining command queues can be left idle during the time slice corresponding to virtual machine VM2. For example, the microcontroller 10 may not need to detect the controlled command queues. Here, the time slice may include, for example, the time allocated to different virtual machines for processing commands.

[0088] When the command queue is a hardware queue, the microcontroller 10 can determine the command queue to use when processing the user's commands based on the number of user command streams.

[0089] When storing the command stream corresponding to a virtual machine into the command queue, taking virtual machine VM1 as an example, if the command stream includes image processing of image Pic1, the command queue corresponding to the command stream in the time slice of virtual machine VM1 stores the command stream corresponding to that virtual machine VM1. For example, this includes K1 and K2. Since virtual machine VM1 can correspond to multiple command streams, the virtual machine stores the command streams into the corresponding command queues.

[0090] For example, when a microprocessor determines whether a command queue is idle, it can do so by using the read pointers and write pointers corresponding to multiple command queues in the command queue. When the read pointer and write pointer in any command queue point to the same position, the command queue is idle.

[0091] If the command queue is idle, it is assumed that the command to be executed obtained from the task flow corresponding to the idle command queue has been executed, and it is necessary to obtain a new command to be executed from the task flow.

[0092] If the command queue is not idle, it is assumed that the commands to be executed obtained from the task flow corresponding to the command queue have not been completed. The commands to be executed that have already been obtained need to be processed before the commands to be executed are obtained from the host.

[0093] When retrieving a command to be executed from the host, the command can be retrieved from a buffer, for example. The buffer stores commands issued by the host.

[0094] See Figure 2 The diagram shows a state illustration of a command queue provided in an embodiment of this disclosure. In 2a, when the command queue 21 is idle, it waits for the host 22 to send an execution command to the buffer 23, and the controller 10 sends the execution command received in the buffer 23 to the command queue 21. The dashed box 211 in the command queue 21 indicates that the command queue is idle. In 2b, when the command queue 24 is not idle, for example, if it contains the execution command 241 represented by the solid box 241, the execution command 241 can be processed first, and then a new execution command sent by the host 22 can be received.

[0095] Here, the buffer may include, for example, a ring buffer or a storage queue. Since a ring buffer is a first-in-first-out circular buffer, it can provide mutual exclusion access to the buffer to the communication program, and can avoid the system overhead added by the storage queue during frequent command allocation. Therefore, in this embodiment of the disclosure, a ring buffer is selected as the command queue.

[0096] For any ring buffer, there are multiple entries for receiving different command streams of to-be-executed commands issued from a host. For example, the host issues two command streams of target recognition processing on a group of images Pic1 and Pic2 to a virtual machine VM1, which can be represented as s1 and s2. At this time, taking the command stream s1 as an example, the command stream s1 includes commands K1 and K2 for target recognition processing on the image Pic1, and the specific description of the commands K1 and K2 has been made in the foregoing, which will not be repeated here.

[0097] At this time, the commands can be issued into the buffer in the following manner: determining at least one command stream corresponding to a user; determining at least one to-be-executed command corresponding to each thread based on each to-be-executed command in the at least one command stream; and storing the at least one to-be-executed command corresponding to each command stream in the buffer by using the storage entry corresponding to each command stream in the buffer.

[0098] When storing the to-be-executed commands in the buffer, the read pointer corresponding to the buffer can be used to query a plurality of target entries according to a pre-set manner, and the write pointer corresponding to the buffer is used to write the to-be-executed commands issued by the host at the target entry corresponding to the to-be-executed commands; or the read pointer corresponding to the buffer can be used to poll a plurality of command storage spaces in the buffer, and the entry corresponding to the command storage space capable of storing the to-be-executed commands is used as the target entry, and then the write pointer corresponding to the buffer is used to write the to-be-executed commands. The specific method for storing the to-be-executed commands in the buffer can be determined according to actual conditions, which will not be repeated here.

[0099] When the microcontroller 10 listens to the buffer, the host continuously issues new command streams into the buffer, and the operation unit 20 continuously processes the commands in the buffer issued into the command queue (Stream Queue), so that when the microcontroller 10 listens to the buffer, there can be no corresponding to-be-executed command when the read pointer of the buffer polls the position of the target entry. At this time, the microcontroller 10 continues to listen to the buffer until the to-be-executed command at the corresponding position is listened to, and then the to-be-executed command is read from the buffer.

[0100] (a2): for the case that the microcontroller 10 reads the to-be-executed command from the target cache corresponding to the current time slice, the microcontroller 10 is configured to read the to-be-executed command from the target cache.

[0101] The target cache can be, for example, a device memory or a double data rate synchronous dynamic random access memory (DDR SDRAM), and is used to store the to-be-executed command read by the microcontroller 10 in the time slice corresponding to the non-virtual machine, or to store the to-be-executed command read by the microcontroller 10 from the command queue.

[0102] Here, since the microcontroller 10 transmits the to-be-executed commands of different users to the command queue in different time slices, time division multiplexing of different users to the command processing device is achieved. Therefore, for the to-be-executed command that is not executed in a time slice, the microcontroller 10 temporarily stores the to-be-executed command in the corresponding target cache after the end of the time slice in which the to-be-executed command is executed. When the next time slice for processing the to-be-executed command arrives, the microcontroller 10 retransmits the to-be-executed command in the target cache to the command queue, so that the command distributor 30 can re-distribute the to-be-executed command to the operation unit 20 for subsequent processing.

[0103] Specifically, when the microcontroller 10 obtains the to-be-executed command, it is used to determine whether there is a to-be-executed command in the target cache corresponding to the current time slice. Taking the example of two virtual machines VM1 and VM2, when the microcontroller 10 determines whether there is a to-be-executed command in the target cache corresponding to the current time slice, it includes the following (b1) or (b2) two cases:

[0104] (b1) In the case where there is a to-be-executed command in the target cache corresponding to the current time slice, the to-be-executed command in the target cache is transmitted to the command queue.

[0105] In implementation, the context at the end of execution of the last time slice of the current time slice can be saved to the target cache corresponding to the last time slice, and the context information includes the processing block identifier corresponding to the last processing block executed at the end of execution, and the context information saved in the target cache corresponding to the current time slice is obtained, and the current time slice is continued to execute.

[0106] For example, if the i-th (i is a positive integer) time slice is allocated to the virtual machine VM1, the i+1-th time slice is allocated to the virtual machine VM2, and the i+2-th time slice is also allocated to the virtual machine VM1, in the i-th time slice, in the case where the command A corresponding to the virtual machine VM1 is not executed, the command A corresponding to the virtual machine VM1 is temporarily stored in the target cache corresponding to A; after waiting for the i+1-th time slice to end, the microcontroller 10 will pull the command A from the corresponding target buffer to the command queue in the i+2-th time slice. The specific processing process is not described here.

[0107] (b2) if there is no to-be-executed command in the target cache corresponding to the current time slice, obtaining the to-be-executed command from the host.

[0108] Specifically, after the end of any time slice, if the commands in the command queue are processed, there is no command that needs to be placed in the target cache to wait for processing in the next corresponding time slice, and when the next corresponding time slice of the corresponding user arrives, the newly issued to-be-executed command is obtained from the host.

[0109] For example, still taking the user as a virtual machine, if the i-th (i is a positive integer) time slice is allocated to the virtual machine VM1, the i+1-th time slice is allocated to the virtual machine VM2, and the i+2-th time slice is also allocated to the virtual machine VM1; in the i-th time slice, if the execution of the command A corresponding to the virtual machine VM1 is completed, after waiting for the end of the i+1-th time slice, the microcontroller 10 detects that there is no to-be-executed command in the target cache corresponding to the current time slice, and then the microcontroller 10 obtains the newly issued to-be-executed command from the host. For the to-be-executed command obtained from the host, it can be stored in the command queue, and in the execution, the plurality of processing operations corresponding to the to-be-executed command are divided into a plurality of processing blocks, and the processing manner provided in the embodiment of the present application is used for processing.

[0110] In addition, for the case that there are multiple virtual machines, for example, including N (N is a positive integer greater than 1) virtual machines, corresponding target caches are allocated as needed, and in an embodiment, at most N corresponding target caches are allocated, so that in the case that the to-be-executed command corresponding to any virtual machine is not executed completely in the time slice corresponding to the virtual machine, the to-be-executed command corresponding to the virtual machine which is not executed completely is stored in the target cache area corresponding to the virtual machine in the next time slice which does not correspond to the virtual machine.

[0111] For the case that the target cache includes the device memory, since the data of the device memory is unidirectional transmission, that is, in the process of data transmission, the data transmission channel only allows unidirectional data transmission, therefore, N memory units can be allocated to N virtual machines to place the to-be-executed commands corresponding to different virtual machines which are not executed completely. For the case that the target cache includes the double data rate synchronous dynamic random access memory, since the data of the double data rate synchronous dynamic random access memory is bidirectional transmission, that is, in the process of data transmission, the data transmission channel allows bidirectional data transmission, for example, direct memory access (Direct Memory Access, DMA), therefore, there can be a case that two virtual machines share one target cache, for example, the same target cache can also be set for the two virtual machines. The specific setting method is not described here.

[0112] After the microcontroller stores the to-be-executed command into the command queue, the command distributor 30 can obtain the to-be-executed command from the command queue and distribute the to-be-executed command to the operation unit 20.

[0113] In the to-be-executed command, a first processing block identifier corresponding to any one of the multiple processing blocks corresponding to the to-be-executed command is carried. After receiving the to-be-executed command issued by the command distributor 30, the operation unit 20 can parse the first processing block identifier from the to-be-executed command, and take the processing block indicated by the first processing block identifier as a starting processing block to be executed at the current time slice, and execute the processing task corresponding to the to-be-executed command based on the starting processing block.

[0114] For example, a scheduler is deployed in the operation unit 20; the to-be-executed command is issued to the scheduler; the scheduler determines the first processing block identifier from the to-be-executed command, and then determines the starting processing block to be executed first from the multiple processing blocks corresponding to the to-be-executed command by using the first processing block identifier, and then assigns each processing operation in the processing block corresponding to the processing block identifier to each thread running in the operation unit 20 as a task, and executes the calculation task corresponding to the processing operation by each thread.

[0115] For example, in the case of processing an image task, when any thread processes any processing operation, the address information of the pixel point to be processed is carried in the task issued by the scheduler, and the thread can obtain the corresponding operand from the storage space storing the image based on the address information, and then process the operand.

[0116] In the embodiment implemented by the GPU, each operator (kernel) can be composed of multiple threads (thread). When the hardware is executed, a certain number of threads can be combined into a processing block (block) or a thread block, as the smallest scheduling granularity.

[0117] The block scheduler in the command distributor 30 schedules the processing blocks in sequence, and when the time slice (slot) ends, the scheduler enters a stop mode (stop mode), and after the block is completed, the block scheduler reports the block identifier (block_id) that has been run to the microcontroller (MCU); when the operator is called again, the reported block identifier is transmitted; in this way, the switching at the processing block level is realized, and the context switching no longer needs to wait for the completion of the current operator, but only needs to wait for the completion of the executed processing block, thereby shortening the waiting time of the current task and improving the context switching efficiency.

[0118] For example, for the to-be-executed command K1, in the case that the to-be-executed command K1 contains 128*128 processing operations, each 32*32 processing operation is determined as a processing block, and thus the to-be-executed command K1 is divided into 16 processing blocks. In this case, if the total processing time of the operation unit 20 for the to-be-executed command K1 is 0.16 s, the processing time for each processing block is 0.01 s, and one time slice is 0.1 s, the processing of the first 10 processing blocks in the to-be-executed command K1 can be completed in 0.1 s by using the processing blocks, and then the context switching between virtual machines is performed.

[0119] In the case that the to-be-executed command K1 is not divided in a finer granularity, that is, in the prior art, after the end of one time slice, the to-be-executed command K1 still needs to be processed, that is, the operation unit 20 still needs to wait for 0.06 s before the to-be-executed command K1 is processed; after the to-be-executed command K1 is processed, the operation unit 20 reports the processing completion information to the microcontroller 10, and then the microcontroller 10 performs the context switching, that is, switches to the next virtual machine to execute the to-be-executed command corresponding to the next virtual machine. In this way, a time delay of 0.06 s is caused for the command processing of the next virtual machine; in the case that there are multiple virtual machines, when the virtual machine context is switched, the Nth virtual machine far away from the first virtual machine may have a phenomenon such as freezing due to the continuously accumulated time delay when the context switching is performed from the first virtual machine to the N-1th virtual machine.

[0120] In the embodiment of the present disclosure, the to-be-executed command is divided in a finer granularity, so that after the end of the current time slice, the context switching of the virtual machine can be performed only after the processing of the processing block being processed is completed, thereby reducing the time delay caused by the context switching after the to-be-executed command being processed in the current time slice is completely processed.

[0121] In the above example, for the to-be-executed command K1, the corresponding first processing block identifiers may be represented as K1_0, K1_2, …, K1_15, and when the operation unit 20 is processing the processing block with the identifier K1-5 at the end of the current time slice, the operation unit 20 reports the processing completion information to the microcontroller 10 after the processing of the processing block with the identifier K1-5 is completed, and the microcontroller 10 performs the context switching; in this case, the time delay is at most 0.01 s.

[0122] For example, when the operation unit 20 reports the processing completion information to the microcontroller 10, the operation unit 20 reports the second processing block identifier corresponding to the processing block that is recently executed to the microcontroller 10.

[0123] The operation unit 20, when reporting the second processing block identifier corresponding to the last processed processing block to the microcontroller 10, is specifically configured to: after the processing task corresponding to the currently executed processing block is executed, take the currently executed processing block as the last executed processing block, and report the second processing block identifier corresponding to the last executed processing block to the microcontroller 10.

[0124] For example, when the to-be-executed command K1 is not processed by the operation unit 20, the corresponding first processing block identifier is K1_0; at the end of a time slice, the operation unit 20 is executing the 10th processing block in the to-be-executed command K1. The operation unit 20 will continue to execute the 10th processing block in the to-be-executed command K1, and take the 10th processing block identifier K1_9 as the second processing block identifier corresponding to the last executed processing block and report it to the microcontroller 10.

[0125] Here, when the operation unit 20 reports the second processing block identifier to the microcontroller 10, the operation unit 20 can send the second processing block identifier to the command distributor; and the command distributor sends the second processing block identifier to the microcontroller 10. In this way, the microcontroller 10 does not need to set an interface for data transmission with the operation unit 20, which can reduce the interface overhead and achieve hardware isolation between the operation unit 20 and the microcontroller 10. In addition, this method can directly use the data transmission channel between the operation unit 20 and the command distributor and the data transmission channel between the command distributor and the microcontroller 10, further improving the multiplexing rate of the data channel.

[0126] After receiving the second processing block identifier reported by the operation unit 20, the microcontroller 10 updates the to-be-executed command in the command queue based on the second processing block identifier.

[0127] Here, when the microcontroller 10 updates the to-be-executed command in the command queue, in a case where the second processing block identifier is the processing block identifier of the last processing block in the plurality of processing blocks, the microcontroller 10 deletes the to-be-executed command from the command queue; in a case where the second processing block identifier is not the processing block identifier of the last processing block in the plurality of processing blocks, the microcontroller 10 determines a target processing block identifier based on the second processing block identifier, replaces the first processing block identifier in the to-be-executed command with the target processing block identifier, and generates a new to-be-executed command.

[0128] The target processing block identifier is the processing block identifier of the next processing block of the last executed processing block.

[0129] After updating the to-be-executed command in the command queue, if the microcontroller 10 replaces the first processing block identifier in the to-be-executed command with the target processing block identifier, the microcontroller 10 stores the new to-be-executed command in a target cache corresponding to the to-be-executed command.

[0130] After the microcontroller 10 stores the new to-be-executed command into the target buffer corresponding to the to-be-executed command, the microcontroller 10 takes the next time slice as a new current time slice, and reads a to-be-executed command corresponding to another virtual machine from the target buffer corresponding to the new current time slice, or obtains the to-be-executed command corresponding to the another virtual machine from the host, so as to complete the context switching of the virtual machine.

[0131] In addition, for multiple commands issued from the host, the multiple commands can also be multiple command streams containing the execution sequence associated with each other. In an embodiment, the to-be-executed command carries a start processing block identifier of a next processing block to be executed in a plurality of processing modules corresponding to the to-be-executed command. In another embodiment, the processing result of a previous command can be set as an operand of a next command. For example, the command corresponding to the first command stream includes target recognition of an image, and the command corresponding to the first command stream includes determining pose information of a plurality of target objects in the image after obtaining the result of the target recognition. At this time, in the second command stream, the address of a pixel point of the image contained in the processing block corresponding to the to-be-executed command can be set as the address of the pixel point of the image obtained after the end of the first command stream, that is, the command processing device can not only complete the processing of one virtual machine corresponding to a plurality of images, but also complete a plurality of continuous processing tasks of the same image by a plurality of virtual machines.

[0132] The embodiment of the disclosure further provides a specific process example of command processing by using the command processing device provided by the embodiment of the disclosure. In the example, a virtual machine VM1 and a virtual machine VM2 are included; one command stream s1 in a plurality of command streams corresponding to the VM1 includes a command K1; one command stream s2 in a plurality of command streams corresponding to the virtual machine VM2 includes a command K2. The command K1 includes 128*128 processing operations, and the 128*128 processing operations are divided into 4*4 processing blocks, and the processing block identifiers of the respective processing blocks are K1_0 to K1_15. The command K2 includes 64*128 processing operations, and the 64*128 processing operations are divided into 2*4 processing blocks, and the processing block identifiers of the respective processing blocks are K2_0 to K2_7.

[0133] (1): When the first time slice t1 arrives, the microcontroller pulls down the command K1 from the buffer RBUF1 corresponding to the VM1, and issues the command K1 to the command queue SQ1 corresponding to the VM1; at this time, the first command block identifier carried in the K1 is K1_0.

[0134] The command distributor obtains the command K1 from the SQ1, and distributes the command K1 to the operation unit ALU1;

[0135] The operation unit ALU1, after parsing the first processing block identifier K1_0 from K1, takes the first processing block in 4*4 as the initial processing block and processes the initial processing block.

[0136] (2) After the end of the time slice t1 and the start of the second time slice t2, the operation unit processes K1_0 to K1_5 and is executing K1_6. At this time, the operation unit continues to process K1_6 and reports the second processing block identifier K1_6 to the command distributor.

[0137] The command distributor reports the second processing block identifier K1_6 to the microcontroller.

[0138] The microcontroller determines K1_7 as the target processing block identifier based on the second processing block identifier K1_6, replaces the first processing block identifier K1_0 carried in the command K1 with K1_7, generates a new command K1', and stores the command K1' in the target buffer corresponding to the command stream s1.

[0139] After storing the command K1' in the target buffer corresponding to the command stream s1, the microcontroller pulls down the command K2 from the buffer RBUF2 corresponding to VM2 and issues the command K2 to the command queue SQ1 corresponding to VM2; at this time, the first command block identifier carried in K2 is K2_0.

[0140] The command distributor obtains the command K2 from SQ1 and distributes the command K2 to the operation unit ALU1;

[0141] The operation unit ALU1, after parsing the first processing block identifier K2_0 from K2, takes the first processing block in 2*4 as the initial processing block and processes the initial processing block.

[0142] (3) After the end of the time slice t2 and the start of the third time slice t3, the operation unit processes K2_0 to K2_3 and is executing K2_4. At this time, the operation unit continues to process K2_4 and reports the second processing block identifier K2_4 to the command distributor.

[0143] The command distributor reports the second processing block identifier K2_4 to the microcontroller.

[0144] The microcontroller determines K2_5 as the target processing block identifier based on the second processing block identifier K2_4, replaces the first processing block identifier K2_0 carried in the command K2 with K2_5, generates a new command K2', and stores the command K2' in the target buffer corresponding to the command stream s2.

[0145] The microcontroller reads the command K1' from the target cache corresponding to the command stream s1, and issues the command K1' to the command queue SQ1 corresponding to the VM1; at this time, the first command block identification carried in K1' is K1_7.

[0146] The command distributor obtains the command K1' from SQ1, and distributes the command K1' to the operation unit ALU1.

[0147] The operation unit ALU1, after parsing the first processing block identification K1_7 from K1', takes the eighth processing block in 4*4 as the initial processing block, and processes the initial processing block.

[0148] (4) At the end of the time slice t3, after the start of the third time slice t4, the operation unit processes K1_7-K1_12 and is executing K1_13. At this time, the operation unit continues to process K1_13 and reports K1_13 to the command distributor as the second processing block identification.

[0149] The command distributor reports the second processing block identification K1_13 to the microcontroller.

[0150] The microcontroller determines K1_14 as the target processing block identification based on the second processing block identification K1_13, replaces the first processing block identification K1_7 carried in the command K1' with K1_14, generates a new command K1'', and stores the command K1'' in the target cache corresponding to the command stream s1.

[0151] The microcontroller reads the command K2' from the target cache corresponding to the command stream s2, and issues the command K2' to the command queue SQ1 corresponding to the VM2; at this time, the first command block identification carried in K2' is K2_5.

[0152] The command distributor obtains the command K2' from SQ1, and distributes the command K2' to the operation unit ALU1.

[0153] The operation unit ALU1, after parsing the first processing block identification K2_5 from K2', takes the sixth processing block in 2*4 as the initial processing block, and processes the initial processing block.

[0154] (5) After the end of the time slice t4, the operation unit processes K2_5-K2_7. At this time, K2_7 is reported to the command distributor as the second processing block identification.

[0155] The command distributor reports the second processing block identification K2_7 to the microcontroller.

[0156] The microcontroller determines that the command K2 is executed based on the second processing block identification K2_7, and deletes it from the command queue Q1.

[0157] At this time, the microcontroller can monitor the RBUF2 corresponding to s2, and if there is a new command, continue to pull down to the command queue Q1 or the target cache corresponding to s2.

[0158] (6) At the end of the time slice t5, after the start of the third time slice t6, the microcontroller reads the command K1" from the target cache corresponding to the command stream s1, and issues the command K1" to the command queue SQ1 corresponding to VM1; at this time, the first command block carried in K1" is K1_14.

[0159] The command distributor obtains the command K1" from SQ1, and distributes the command K1" to the operation unit ALU1.

[0160] The operation unit ALU1, after parsing the first processing block identifier K1_14 from K1", takes the 15th processing block in 4*4 as the initial processing block, and processes the initial processing block.

[0161] (7) At the end of the time slice t6, after the start of the third time slice t7, the operation unit processes K1_14~K1_15. At this time, K1_15 is taken as the second processing block identifier and reported to the command distributor.

[0162] The command distributor reports the second processing block identifier K1_15 to the microcontroller.

[0163] The microcontroller determines that the command K1 is executed based on the second processing block identifier K1_15, and deletes it from the command queue Q1.

[0164] At this time, the microcontroller can monitor the RBUF1 corresponding to s1, and if there is a new command, continue to pull down to the command queue Q1 or the target cache corresponding to s1.

[0165] Through the above process, time division multiplexing of VM1 and VM2 on the command processing device is realized.

[0166] Referring to Figure 3 The microcontroller 31, a plurality of stream queues 32, a command distributor 33, and a plurality of operation units 34 are provided. The plurality of stream queues 32 includes a command queue SQ1 321 and a command queue SQ2 322. The plurality of operation units 34 includes an operation unit 0, an operation unit 1, …, and an operation unit n. The plurality of command streams respectively correspond to a plurality of target caches 35. The plurality of target caches 35 includes a target cache corresponding to a command stream s1 351 and a target cache corresponding to a command stream s2 352. For any command stream, the commands stored in the corresponding target cache are respectively represented as eSQ1, eSQ2, …, and eSQm.

[0167] Referring to Figure 4 FIG. 4 shows a flowchart of a command processing method according to an embodiment of the present disclosure, which includes the following steps:

[0168] S401: The microcontroller acquires a to-be-executed command after a current time slice arrives; the to-be-executed command carries a first processing block identifier used to indicate a processing block corresponding to the current time slice in a plurality of processing blocks corresponding to the to-be-executed command.

[0169] S402: The operation unit acquires the to-be-executed command and executes a processing task corresponding to the to-be-executed command based on the first processing block identifier.

[0170] In an optional implementation, the command processing apparatus further includes a command distributor.

[0171] The microcontroller acquires the to-be-executed command after the current time slice arrives, including:

[0172] The microcontroller acquires the to-be-executed command after the current time slice arrives, and stores the to-be-executed command in a command queue.

[0173] The command processing method further includes:

[0174] The command distributor acquires the to-be-executed command from the command queue and distributes the to-be-executed command to the operation unit.

[0175] The operation unit acquires the to-be-executed command and executes a processing task corresponding to the to-be-executed command based on the first processing block identifier, including:

[0176] The operation unit acquires the to-be-executed command and, after receiving the to-be-executed command distributed by the command distributor, executes a processing task corresponding to the to-be-executed command based on the first processing block identifier.

[0177] In an optional implementation, the operation unit executes a processing task corresponding to the to-be-executed command based on the first processing block identifier, including: the operation unit determines a starting processing block to be executed in the current time slice from the plurality of processing blocks based on the first processing block identifier, and executes a processing task corresponding to the to-be-executed command based on the starting processing block.

[0178] In an optional implementation, the microcontroller acquires the to-be-executed command after the current time slice arrives, including: the microcontroller reads the to-be-executed command from a target cache corresponding to the current time slice.

[0179] Or, acquires the to-be-executed command from a host.

[0180] In an alternative implementation, the microcontroller acquires the command to be executed after the current time slice arrives, including that the microcontroller determines whether the command queue is idle;

[0181] In the case that the command queue is idle, the microcontroller monitors the buffer; the buffer is used by the host to store the command to be executed;

[0182] In the case that the command to be executed exists in the buffer, the microcontroller reads the command to be executed from the buffer.

[0183] In an alternative implementation, the buffer includes a ring buffer; the ring buffer has multiple entries; different entries are used by the host to store the command to be executed of different command streams;

[0184] The microcontroller acquires the command to be executed after the current time slice arrives, including that the microcontroller determines whether the command queue is idle;

[0185] The microcontroller determines the target entry from the buffer based on the command stream corresponding to the current time slice; based on the determined target entry, the microcontroller monitors whether the command to be executed exists in the buffer.

[0186] In an alternative implementation, the microcontroller acquires the command to be executed after the current time slice arrives, including that the microcontroller determines whether the command to be executed exists in the target cache corresponding to the current time slice;

[0187] In the case that the command to be executed does not exist in the target cache corresponding to the current time slice, the microcontroller acquires the command to be executed from the host.

[0188] In an alternative implementation, the microcontroller acquires the command to be executed after the current time slice arrives, including that the microcontroller reads the command to be executed from the target cache in the case that the command to be executed exists in the target cache corresponding to the current time slice.

[0189] In an alternative implementation, further including that the operation unit reports the second processing block identifier corresponding to the processing block that has been executed most recently to the microcontroller after the current time slice ends;

[0190] The microcontroller updates the command to be executed in the command queue based on the second processing block identifier after receiving the second processing block identifier reported by the operation unit.

[0191] In an alternative implementation, the operation unit reports the second processing block identifier corresponding to the last executed processing block to the microcontroller after the end of the current time slice, including: the operation unit, after the execution of the processing task corresponding to the currently executed processing block is completed, takes the currently executed processing block as the last executed processing block, and reports the second processing block identifier corresponding to the last executed processing block to the microcontroller.

[0192] In an alternative implementation, the operation unit sends the second processing block identifier to the command distributor.

[0193] The command distributor sends the second processing block identifier to the microcontroller.

[0194] In an alternative implementation, the updating of the to-be-executed command in the command queue based on the second processing block identifier includes: in the case where the second processing block identifier is the processing block identifier of the last processing block in the plurality of processing blocks, the microcontroller deletes the to-be-executed command from the command queue.

[0195] In the case where the second processing block identifier is not the processing block identifier of the last command block in the plurality of processing blocks, a target processing block identifier is determined based on the second processing block identifier, and the first processing block identifier in the to-be-executed command is replaced with the target processing block identifier to generate a new to-be-executed command.

[0196] The target processing block identifier is the processing block identifier of the next processing block of the last executed processing block.

[0197] In an alternative implementation, the microcontroller stores the new to-be-executed command in the target cache corresponding to the to-be-executed command after generating the new to-be-executed command.

[0198] Referring to Figure 5 FIG. 1 is a flowchart of another command processing method provided by an embodiment of the present disclosure, including:

[0199] S501: The operation unit reports a first processing block identifier of a current processing block of a target command executed in a current time slice to a microcontroller in response to the end of the current time slice; the current processing block is any one of at least one processing block in the target command.

[0200] S502: After the microcontroller receives the first processing block identifier reported by the operation unit, the microcontroller updates the target command by using the first processing block identifier.

[0201] The electronic device provided by the embodiments of the present disclosure can include a host, a buffer, and a command processing apparatus. The host is configured to issue a command to be executed and store the command in the buffer. The command processing apparatus is configured to perform the method described in any of the embodiments of the command processing method of the present disclosure.

[0202] The command processing apparatus provided by the embodiments of the present disclosure can include a chip, an AI chip, or the like. The electronic device provided by the embodiments of the present disclosure can include a smart terminal such as a mobile phone, or can be another device having a camera and capable of image processing, a server, or the like, which is not limited herein.

[0203] The computer-readable storage medium provided by the embodiments of the present disclosure has a computer program stored thereon. The program is executed by a microcontroller or an arithmetic unit to implement the method described in any of the embodiments of the command processing method of the present disclosure.

[0204] Those skilled in the art can understand that, in the above method of the specific embodiments, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0205] Based on the same inventive concept, the embodiments of the present disclosure also provide a command processing method corresponding to the command processing apparatus. Since the principle of the device in the embodiments of the present disclosure solves the problem, the implementation of the method can refer to the implementation of the device, and the repeated parts will not be described herein.

[0206] The embodiments of the present disclosure also provide a computer program product carrying a program code. The program code includes commands that can be used to execute the steps of the command processing method described in the above method embodiments. For details, refer to the above method embodiments, which will not be described herein.

[0207] The computer program product can be specifically implemented by hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied as a computer storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), and the like.

[0208] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the foregoing method embodiment, and will not be repeated here. In several embodiments provided in the present disclosure, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and there can be another division in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some communication interfaces, devices or units, and can be electrical, mechanical or other forms.

[0209] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0210] In addition, the functional units in each embodiment of the present disclosure can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.

[0211] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer readable storage medium executable by a processor. Based on this understanding, the technical solutions of the present disclosure essentially or the part of the prior art or the part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of commands for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present disclosure. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program codes that can be stored in the medium.

[0212] Finally, it should be noted that the above-described embodiments are merely specific embodiments of the present disclosure, used to illustrate the technical solutions of the present disclosure, and are not intended to limit the present disclosure. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can make modifications or easy changes to the technical solutions described in the foregoing embodiments, or easily think of changes or equivalent replacements for some of the technical features; and these modifications, changes or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A command processing apparatus characterized by comprising: The processing device comprises: a microcontroller and an operation unit; the microcontroller is configured to acquire a to-be-executed command after a current time slice arrives, the to-be-executed command carrying a first processing block identifier indicating a processing block corresponding to the current time slice in a plurality of processing blocks corresponding to the to-be-executed command, the plurality of processing blocks being divided by a plurality of processing operations corresponding to the to-be-executed command, and the first processing block identifier indicating the processing block to be executed in the current time slice; the operation unit is configured to acquire the to-be-executed command and execute a processing task corresponding to the to-be-executed command based on the first processing block identifier; the processing device further comprises a command distributor; the microcontroller is configured to acquire a to-be-executed command after a current time slice arrives and store the to-be-executed command in a command queue; the command distributor is configured to acquire the to-be-executed command from the command queue and distribute the to-be-executed command to the operation unit; the operation unit is configured to acquire the to-be-executed command and execute a processing task corresponding to the to-be-executed command based on the first processing block identifier after receiving the to-be-executed command distributed by the command distributor; the operation unit is further configured to report a second processing block identifier corresponding to a processing block that has been executed most recently to the microcontroller after the current time slice ends; the microcontroller is further configured to update the to-be-executed command in the command queue based on the second processing block identifier after receiving the second processing block identifier reported by the operation unit.

2. The command processing device according to claim 1, characterized by the operation unit is configured to: determine a starting processing block to be executed in the current time slice from the plurality of processing blocks based on the first processing block identifier, and execute the processing task corresponding to the to-be-executed command based on the starting processing block.

3. The command processing device according to claim 1, wherein the microcontroller is configured to: read the to-be-executed command from a target buffer corresponding to the current time slice, or acquire the to-be-executed command from a host. the microcontroller is configured to:

4. The command processing device according to claim 3, wherein determine whether the command queue is idle; in the case that the command queue is idle, listen to a buffer, the buffer being used by the host to store the to-be-executed command; in the case that the to-be-executed command is found in the buffer, read the to-be-executed command from the buffer. the buffer comprises a ring buffer, the ring buffer having a plurality of entrances, and different entrances being used by the host to store to-be-executed commands of different command streams; 5. The command processing device according to claim 4, wherein the microcontroller is configured to: determine a target entrance from the buffer based on a command stream corresponding to the current time slice; based on the determined target entrance, listen to whether the to-be-executed command is stored in the buffer. the microcontroller is configured to:

6. The command processing apparatus according to any one of claims 1 to 5, characterized by determine whether the to-be-executed command exists in a target buffer corresponding to the current time slice; in the case that the to-be-executed command does not exist in the target buffer corresponding to the current time slice, acquire the to-be-executed command from the host. the microcontroller is further configured to, in the case that the to-be-executed command exists in the target buffer corresponding to the current time slice, read the to-be-executed command from the target buffer.

7. The command processing device according to claim 6, wherein the operation unit is configured to:

8. The command processing apparatus according to claim 1, wherein ​ After a processing task corresponding to a currently executed processing block is executed, the currently executed processing block is taken as a last executed processing block, and a second processing block identifier corresponding to the last executed processing block is reported to the microcontroller.

9. The command processing apparatus according to claim 1 or 8, wherein The operation unit is configured to send the second processing block identifier to a command distributor. The command distributor is further configured to send the second processing block identifier to the microcontroller.

10. The command processing apparatus according to claim 1 or 8, wherein The microcontroller is configured to: in a case where the second processing block identifier is a processing block identifier of a last processing block in the plurality of processing blocks, delete the to-be-executed command from the command queue; in a case where the second processing block identifier is not a processing block identifier corresponding to a last command block in the plurality of processing blocks, determine a target processing block identifier based on the second processing block identifier, replace a first processing block identifier in the to-be-executed command with the target processing block identifier, and generate a new to-be-executed command; wherein the target processing block identifier is a processing block identifier corresponding to a next processing block of the last executed processing block.

11. The command processing device according to claim 10, wherein The microcontroller is further configured to, after the new to-be-executed command is generated, store the new to-be-executed command into a target cache corresponding to the to-be-executed command.

12. A command processing device, characterized by comprising: The command processing device comprises: a microcontroller and an operation unit. The operation unit is configured to, in response to the end of a current time slice, report a first processing block identifier of a current processing block of a target command executed in the current time slice to the microcontroller; wherein the current processing block is any one of at least one processing block in the target command, the at least one processing block is divided by a plurality of processing operations corresponding to the target command, and the first processing block identifier indicates a processing block to be executed in the current time slice. The microcontroller is configured to, after receiving the first processing block identifier reported by the operation unit, update the target command by using the first processing block identifier.

13. A command processing method characterized by comprising: The command processing method is applied to a command processing device, and the command processing device comprises a microcontroller and an operation unit. The microcontroller obtains a to-be-executed command after a current time slice arrives; the to-be-executed command carries a first processing block identifier indicating a processing block corresponding to the current time slice in a plurality of processing blocks corresponding to the to-be-executed command; wherein the plurality of processing blocks are divided by a plurality of processing operations corresponding to the to-be-executed command, and the first processing block identifier indicates a processing block to be executed in the current time slice. The operation unit obtains the to-be-executed command, and executes a processing task corresponding to the to-be-executed command based on the first processing block identifier. The command processing device further comprises a command distributor. The microcontroller obtains a to-be-executed command after a current time slice arrives, which comprises: The microcontroller obtains a to-be-executed command after a current time slice arrives, and stores the to-be-executed command into a command queue. The command processing method further comprises: The command distributor obtains the to-be-executed command from the command queue, and distributes the to-be-executed command to the operation unit. The operation unit obtains the to-be-executed command, and executes a processing task corresponding to the to-be-executed command based on the first processing block identifier. The operation unit obtains the to-be-executed command, and executes a processing task corresponding to the to-be-executed command based on the first processing block identifier. Further comprising: the operation unit reports a second processing block identifier corresponding to a processing block that has been executed most recently to the microcontroller after the current time slice ends. The microcontroller updates the to-be-executed command in the command queue based on the second processing block identifier after receiving the second processing block identifier reported by the operation unit.

14. The command processing method of claim 13, wherein, The operation unit determines a starting processing block to be executed in the current time slice from the plurality of processing blocks based on the first processing block identifier, and executes a processing task corresponding to the to-be-executed command based on the starting processing block.

15. The command processing method of claim 13, wherein, The microcontroller obtains the to-be-executed command after the current time slice arrives, including: the microcontroller reads the to-be-executed command from a target buffer corresponding to the current time slice. Alternatively, the microcontroller obtains the to-be-executed command from a host.

16. The command processing method of claim 15, wherein, The microcontroller obtains the to-be-executed command after the current time slice arrives, including: the microcontroller determines whether the command queue is idle. In the case that the command queue is idle, the buffer is listened to; the buffer is used by the host to store the to-be-executed command. In the case that the to-be-executed command exists in the buffer, the to-be-executed command is read from the buffer.

17. The command processing method of claim 16, wherein The buffer includes a ring buffer; the ring buffer has a plurality of entrances. Different entrances are used by the host to store to-be-executed commands of different command streams. The microcontroller obtains the to-be-executed command after the current time slice arrives, including: The microcontroller determines a target entrance from the buffer based on a command stream corresponding to the current time slice. Based on the determined target entrance, it is listened whether the to-be-executed command is stored in the buffer.

18. The command processing method according to any one of claims 13-17, wherein, The microcontroller obtains the to-be-executed command after the current time slice arrives, including: the microcontroller determines whether there is a to-be-executed command in a target buffer corresponding to the current time slice. In the case that there is no to-be-executed command in the target buffer corresponding to the current time slice, the to-be-executed command is obtained from a host.

19. The command processing method of claim 18, wherein, The microcontroller obtains the to-be-executed command after the current time slice arrives, including: in the case that there is the to-be-executed command in the target buffer corresponding to the current time slice, the to-be-executed command is read from the target buffer.

20. The command processing method of claim 13, wherein, The operation unit reports a second processing block identifier corresponding to a processing block that has been executed most recently to the microcontroller after the current time slice ends, including: after a processing task corresponding to a processing block that is being executed currently is executed, the processing block that is being executed currently is taken as a processing block that has been executed most recently, and a second processing block identifier corresponding to the processing block that has been executed most recently is reported to the microcontroller.

21. The command processing method according to claim 13 or 20, characterized by, Further comprising: The operation unit sends the second processing block identifier to the command distributor; The command distributor sends the second processing block identifier to the microcontroller.

22. The command processing method according to claim 13 or 20, wherein The microcontroller updates the to-be-executed command in the command queue based on the second processing block identifier, including: in a case where the second processing block identifier is a processing block identifier of a last processing block in the plurality of processing blocks, the microcontroller deletes the to-be-executed command from the command queue; In a case where the second processing block identifier is not a processing block identifier of a last command block in the plurality of processing blocks, the microcontroller determines a target processing block identifier based on the second processing block identifier, and replaces a first processing block identifier in the to-be-executed command with the target processing block identifier to generate a new to-be-executed command; The target processing block identifier is a processing block identifier of a next processing block of the last executed processing block.

23. The command processing method of claim 22, wherein Further comprising: The microcontroller stores the new to-be-executed command in a target cache corresponding to the to-be-executed command after generating the new to-be-executed command.

24. A command processing method characterized by comprising: The command processing method is applied to a command processing device, and the command processing device includes a microcontroller and an operation unit. The command processing method includes: The operation unit reports a first processing block identifier of a current processing block of a target command executed in a current time slice to the microcontroller in response to the end of the current time slice; the current processing block is any one of at least one processing block in the target command; the at least one processing block is divided by a plurality of processing operations corresponding to the target command, and the first processing block identifier indicates a processing block to be executed in the current time slice. The microcontroller updates the target command by using the first processing block identifier after receiving the first processing block identifier reported by the operation unit.

25. An electronic device, comprising: The command processing device includes a host, a buffer, and a command processing device. The host is configured to issue a to-be-executed command and store the to-be-executed command in the buffer. The command processing device is configured to execute the command processing method in any one of claims 13 to 23 or the command processing method in claim 24.

26. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the microcontroller and the operation unit to implement the command processing method in any one of claims 13 to 23 or the command processing method in claim 24.

Citation Information

Patent Citations

  • Message processing method and device

    CN105337896A

  • Multi-thread task scheduling method, apparatus and device, and readable medium

    CN110795222A