Data processing device and method and electronic equipment

By managing different types of instruction queues and utilizing the data processing devices and methods of the instruction input module and execution module, the problem of low efficiency of CPU in high-concurrency data comparison and large-scale data transportation is solved, and efficient parallel processing of instruction execution is achieved.

CN121523737AActive Publication Date: 2026-02-13MOORE THREADS TECH CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202511695698.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-13
Estimated Expiration
2045-11-18

AI Technical Summary

Technical Problem

Existing technologies cannot simultaneously meet the demands of high-concurrency data comparison (small data, low latency) and large-scale data transfer (high throughput), resulting in low CPU execution efficiency.

Method used

A data processing apparatus and method are provided, which manages different types of instruction queues through an instruction input module and an instruction execution module, reads target instructions from the instruction queues using the idle instruction entries and non-idle instruction categories of the internal storage module, and returns an interrupt signal when the instructions are successfully executed, thereby realizing the scheduling of different instruction queues.

Benefits of technology

It enables parallel processing of different instruction queues, reduces system processing overhead, improves system instruction execution efficiency, and meets the needs of high-concurrency data comparison and large-scale data transfer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121523737A_ABST
    Figure CN121523737A_ABST
Patent Text Reader

Abstract

The invention provides a data processing device and method and electronic equipment. The device comprises an instruction input module, an internal storage module and an instruction execution module, the instruction input module is configured to read an unexecuted target instruction from a first instruction queue and / or a second instruction queue according to an item identifier of an idle first item in the internal storage module and an instruction category in a non-idle second item; storing the instruction information of the target instruction to the first entry; the instruction execution module is configured to execute a target instruction according to an instruction category under the condition that the instruction information is stored in the first entry; and under the condition that the execution result is successful execution, an interrupt signal of the target instruction is sent to the instruction input module. According to the embodiment of the invention, the instruction execution efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, and in particular, to a data processing apparatus and method, and an electronic device. BACKGROUND

[0002] In modern computer systems, a system on chip (SoC) integrated in a chip has been widely applied as a high-performance computing architecture technology. The chip is usually provided with a central processing unit (CPU), an on-chip memory and the like for performing corresponding tasks and realizing data storage and the like. Among them, the CPU needs to read data from the memory and carry it to the internal cache (for example, static random-access memory (SRAM)) when working normally, and also needs to perform data comparison and other types of tasks.

[0003] However, the efficiency of the CPU in performing different types of data processing tasks is different. The related art cannot simultaneously meet the needs of high-concurrency data comparison (small data, low latency) and large-scale data carrying (high throughput). SUMMARY

[0004] The present disclosure provides a data processing apparatus and method, and an electronic device.

[0005] In a first aspect, the present disclosure provides a data processing apparatus, comprising: an instruction input module, an internal storage module and an instruction execution module. The instruction input module is configured to read an unexecuted target instruction from a first instruction queue and / or a second instruction queue according to a target identifier of a first idle instruction entry and an instruction category of an instruction in a second non-idle instruction entry in a plurality of instruction entries of the internal storage module, and store instruction information of the target instruction to the first instruction entry; wherein the first instruction queue comprises instructions of a first category with an instruction execution number of 1, and the second instruction queue comprises instructions of a second category with an instruction execution number greater than or equal to 1. The instruction execution module is configured to execute the target instruction according to the instruction category of the target instruction in the case that the instruction information of the target instruction is stored in the first instruction entry, and send an interrupt signal of the target instruction to the instruction input module in the case that the execution result of the target instruction is execution success.

[0006] In a second aspect, the present disclosure provides a data processing method, comprising: reading an unexecuted target instruction from a first instruction queue and / or a second instruction queue according to a target identifier of a first idle instruction entry and an instruction category of an instruction in a second non-idle instruction entry in a plurality of instruction entries; and storing instruction information of the target instruction into the first instruction entry; wherein the first instruction queue comprises instructions of a first category with an execution number of 1, and the second instruction queue comprises instructions of a second category with an execution number greater than or equal to 1; executing the target instruction according to the instruction category of the target instruction to obtain an execution result of the target instruction; and reporting an interrupt signal of the target instruction in a case that the execution result of the target instruction is successful.

[0007] In a third aspect, the present disclosure provides an electronic device comprising the data processing apparatus described above.

[0008] The embodiments provided by the present disclosure can simultaneously manage different categories of instruction queues, the instruction input module reads a target instruction from an instruction queue and stores it into an idle instruction entry according to an identifier of an idle instruction entry in an internal storage module and an instruction category of an instruction in a non-idle instruction entry, the instruction execution module executes the target instruction according to the instruction category, and returns an interrupt signal of the target instruction in a case that the execution is successful, thereby realizing scheduling of different instruction queues, simultaneously executing instructions of different categories in the execution queue, reducing processing overhead of the system, and improving instruction execution efficiency of the system.

[0009] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The accompanying drawings are included to provide a further understanding of the present disclosure and constitute a part of the specification, which together with the embodiments of the present disclosure serve to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent from the detailed description of the specific example embodiments, with reference to the accompanying drawings, in which:

[0011] Figure 1 A structural schematic diagram of a data processing apparatus provided by an embodiment of the present disclosure.

[0012] Figure 2 A flowchart of a data processing method provided by an embodiment of the present disclosure.

[0013] Figure 3 A block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0014] In order to better understand the technical solutions of the present disclosure, the exemplary embodiments of the present disclosure are described below in conjunction with the drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.

[0015] In the case of no conflict, each embodiment of the present disclosure and each feature in the embodiments can be combined with each other.

[0016] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0017] The terms used herein are only used to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the terms "comprise" and / or "consist of," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The terms "connected" or "coupled" and / or similar terms are not limited to a physical or mechanical connection, but can include an electrical connection, whether direct or indirect.

[0018] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0019] In a SoC chip, the host side (Host) usually refers to a software program running on the upper layer of the SoC chip, which is used to receive external tasks and issue them to the chip for execution. The firmware (Firmware, FW) usually refers to a hardware program running in the chip, which is used to drive the hardware (such as CPU) to execute the tasks and instructions issued by the host side.

[0020] In actual task processing, the tasks to be processed can include data carrying tasks and data comparison tasks. In the data comparison task, semaphore is a mechanism for synchronizing access of multiple execution threads (or tasks) to shared resources, which is essentially a counter. If the host side wants the firmware to perform a task, but the task depends on the completion of some other pre-task, the host side will quantify this dependency into a numerical value, i.e., the dependency value of the semaphore, and write it into the instruction queue of the firmware. At the same time, there is a shared variable in the on-chip memory, i.e., the current value of the semaphore, and each time a pre-task is completed, there will be a release operation to atomically increase the current value of the semaphore by 1. The CPU on the chip needs to frequently read the current value from the memory and compare it with the dependency value sent by the host side. If the condition is met, the firmware can execute the task, otherwise, the check needs to be performed again later, i.e., the polling mechanism.

[0021] In the related art, the CPU is usually controlled by the firmware to sequentially initiate data reading and comparison operations. This operation is a sequential operation, and the CPU cannot be controlled to simultaneously initiate a large number of data reading operations, resulting in low efficiency of the CPU in executing data comparison instructions, and a dedicated value comparison unit (VCE) is needed to implement comparison processing.

[0022] In the data carrying task, the CPU needs to read data from the memory and carry it to the internal cache. It can be considered that the essence is a direct memory access (DMA) operation. In the related art, a dedicated DMA module is designed to accelerate the data carrying process.

[0023] Among them, the number of DMA instructions is usually small and the execution period is short, the number of VCE instructions is usually large but the execution period is long, and repeated comparison needs to be performed by constant polling, but the operation of the CPU to read the current value of the semaphore and the implementation of the data carrying operation are similar. The DMA engine in the related art is difficult to simultaneously meet the needs of high-concurrency data comparison (small data, low delay) and large-scale data carrying (high throughput). Therefore, a new type of data management architecture capable of simultaneously and efficiently supporting long and short cycle queues is needed, which can also be called a high-speed data management unit (DMU)

[0024] According to the data processing apparatus provided by the embodiment of the present disclosure, different categories of instruction queues can be managed simultaneously, the instruction input module reads target instructions from the instruction queue and stores them into idle instruction entries according to the identification of the idle instruction entries and the instruction category in the non-idle instruction entries in the internal storage module, the instruction execution module executes the target instructions according to the instruction category and returns an interrupt signal of the target instructions when the execution is successful, thereby realizing the scheduling of different instruction queues, executing instructions of different categories in the execution queue simultaneously, reducing the processing overhead of the system and improving the instruction execution efficiency of the system.

[0025] The data processing apparatus according to the embodiment of the present disclosure, also referred to as a data management unit, can be applied to a chip system, which can be any chip system itself or a device or component in the chip system. The chip system can be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a data processing unit (DPU), etc., and the present disclosure does not limit the specific type of chip system to which the data processing apparatus corresponds.

[0026] Figure 1 A structural schematic diagram of the data processing apparatus provided by the embodiment of the present disclosure is shown in FIG. 1. Figure 1 The data processing apparatus according to the embodiment of the present disclosure can include an instruction input module 11, an internal storage module 12, an instruction execution module 13, and a register control (REG_CTRL) module 14. The instruction input module 11 is connected to the on-chip cache of the firmware FM through a preset communication interface, so as to obtain the instructions in the instruction queue of the on-chip cache. The host side can send the instructions corresponding to the task to be executed by the firmware into the instruction queue.

[0027] In an example, when the communication interface adopts an Advanced eXtensible Interface (AXI) and a Static Random-Access Memory (SRAM) is used as the internal storage, the instruction input module 11 can be referred to as AXI_I, the internal storage module 12 can be referred to as DESC_SRAM, and the instruction execution module 13 can be referred to as AXI_D. The present disclosure does not limit the specific type of communication interface and the specific storage type of the internal storage.

[0028] In some possible implementation manners, the instruction queue can include a first instruction queue and a second instruction queue, the first instruction queue includes instructions of a first category with an execution frequency of 1, and the second instruction queue includes instructions of a second category with an execution frequency greater than or equal to 1.

[0029] That is, the first instruction queue is a short loop queue, the life cycle of the instruction in the queue is 1, that is, the execution times of the instruction is 1, and the instruction is executed in sequence. The second instruction queue is a long loop queue, the life cycle of the instruction in the queue depends on the execution times, and polling is needed until the instruction is executed successfully, that is, the execution times of the instruction is greater than or equal to 1, and the instruction is executed out of order.

[0030] In some possible implementation ways, the first type of instruction includes a direct memory access (DMA) based data transfer instruction; and the second type of instruction includes a data comparison unit (VCE) based data comparison instruction.

[0031] In an exemplary use scenario, the direct memory access (DMA) based data transfer instruction can also be directly referred to as a DMA instruction, and is executed in the short loop queue; and the data comparison unit (VCE) based data comparison instruction can also be directly referred to as a VCE instruction, and is executed in the long loop queue.

[0032] In some possible implementation ways, the instruction input module 11 is configured to: read an unexecuted target instruction from the first instruction queue and / or the second instruction queue according to a first entry identifier of an idle first entry and an instruction category of an instruction in a non-idle second entry in the plurality of instruction entries of the internal storage module 12; and store instruction information of the target instruction into the first entry.

[0033] The instruction execution module 13 is configured to: in the case that the instruction information of the target instruction is stored in the first entry, execute the target instruction according to the instruction category of the target instruction; and in the case that the execution result of the target instruction is execution success, send an interrupt signal of the target instruction to the instruction input module.

[0034] For example, the instruction input module 11 is responsible for instruction reading and instruction identification (ID) maintenance; the internal storage module 12 is provided with a plurality of instruction entries, each instruction entry is used for storing instruction information of an instruction, and each instruction entry has an entry identifier, which is used for indicating a storage address corresponding to the instruction entry; and the instruction execution module 13 is used for reading instruction information from the instruction entry and performing corresponding processing.

[0035] In the case of using the AXI interface, the bus identifier can be an identifier of the AXI interface, the AXI interface supports concurrent processing of outstanding transactions, and allows the master device (for example, a CPU) to continuously initiate multiple bus transactions without receiving a response to a previous transaction. In an example, the number of instruction entries (or the depth of the internal storage module) can be 32, which can store up to 1 instruction of the first type and 31 instructions of the second type. It should be understood that a person skilled in the art can set the number of instruction entries and the maximum number of instructions of each type according to actual conditions, and the present disclosure does not limit this.

[0036] In some possible implementation manners, if there is a free first entry in the instruction entries of the internal storage module, the internal storage module sends the bus identifier of the first entry to the instruction input module; after receiving the bus identifier of the first entry, the instruction input module can read instructions from the first instruction queue and the second instruction queue and arbitrate.

[0037] In some possible implementation manners, the instructions in the first instruction queue and the second instruction queue each have a flag bit for indicating the instruction type, for example, the flag bit being 1 indicates that the instruction is an instruction of the first type, that is, a DMA instruction; and the flag bit being 0 indicates that the instruction is an instruction of the second type, that is, a VCE instruction. The flag bit is stored in the non-free second entry in the internal storage module.

[0038] In some possible implementation manners, the instruction input module arbitrates according to the instruction type of the instruction in the non-free second entry in the instruction entries of the internal storage module; if the second entry includes an instruction of the first type, the arbitration result is to read an instruction of the second type this time; if the second entry does not include an instruction of the first type, the arbitration result is to read at least one instruction of the first type this time; and the total number of instructions read this time is determined according to the number of first entries, so as to obtain the target instructions read this time.

[0039] In an example, the second entry does not include an instruction of the first type, and the number of first entries is 5, so that 1 instruction of the first type and 4 instructions of the second type can be read, and 5 target instructions are obtained.

[0040] In some possible implementation manners, the instruction input module can read the target instructions respectively, delete the reserved bits in the target instructions, and write the remaining contents as instruction information into the corresponding first entries.

[0041] In this way, the instruction entries of the internal storage module can store instructions of two types at the same time, parallel processing of instructions of different types can be implemented, and the processing efficiency can be improved.

[0042] In some possible implementation manners, after the instruction information of the target instruction is stored in the first entry of the internal storage module 12, the instruction execution module 13 is notified to perform data reading. After receiving the notification from the internal storage module 12, the instruction execution module 13 can read the instruction information of the target instruction from the internal storage module 12, determine the category of the target instruction, and execute the target instruction according to the instruction category, for example, perform data transfer or data comparison of the corresponding data. If the execution result of the target instruction is successful, the target instruction corresponding to the entry identifier and an interrupt signal are sent to the instruction input module, so as to be reported.

[0043] In this way, the scheduling of different instruction queues can be implemented, the parallel execution of instructions of different categories in the queues can be implemented, the processing overhead of the system can be reduced, and the execution efficiency of the system can be improved.

[0044] The data processing apparatus according to the embodiments of the present disclosure is described below.

[0045] After the system is started, the host side can send instructions corresponding to tasks to be executed by the firmware to the instruction queues in the on-chip cache. The instruction input module 11 connects the on-chip cache through a preset communication interface, so as to read the instructions in the instruction queues. Different categories of instructions are sent to different instruction queues. The first instruction queue includes instructions of a first category with an execution frequency of 1, and the second instruction queue includes instructions of a second category with an execution frequency greater than or equal to 1.

[0046] In some possible implementation manners, the data processing apparatus according to the embodiments of the present disclosure further includes a register control module 14 configured to maintain the execution state of the instructions.

[0047] In some possible implementation manners, the instruction input module is further configured to, in a case where there is a new instruction in the first instruction queue and / or the second instruction queue, set an execution state bit corresponding to the new instruction in the register control module, and set the execution state bit to a first value.

[0048] That is, whenever the first instruction queue and / or the second instruction queue receives a new instruction, referred to as a new instruction, the instruction ID of the new instruction is recorded in the register control module, and the execution state bit corresponding to the new instruction is set, and the value of the execution state bit is set to a first value. The first value is used to indicate that the instruction is not executed, and correspondingly, the second value is used to indicate that the instruction is executed. In an example, the first value can be 0, and the second value can be 1. In this way, the maintenance of the execution state of the instructions can be implemented.

[0049] In some possible implementation manners, the instruction input module can detect the value of the execution state bit of the instruction in the register control module, and when there is an instruction with the first value of the execution state bit in the first instruction queue and / or the second instruction queue, that is, when there is an unexecuted instruction, the instruction input module starts the processing of polling the instruction.

[0050] At the start, the instruction entries of the internal storage module 12 are all idle, and the instruction input module can directly start reading the instruction from the head position of the first instruction queue and the second instruction queue for the first time; during the processing, when the instruction corresponding to the instruction entry of the internal storage module 12 is executed, the instruction input module is sent the instruction identifier of the first idle entry, so that the instruction input module starts reading the instruction again.

[0051] In some possible implementation manners, in each processing of reading the instruction, the instruction input module reads the unexecuted target instruction from the first instruction queue and / or the second instruction queue according to the instruction identifier of the first idle entry and the instruction category of the instruction in the second non-idle entry of the plurality of instruction entries of the internal storage module; and stores the instruction information of the target instruction into the first entry.

[0052] In some possible implementation manners, the processing process includes: in the case that the instruction in the second entry includes the instruction of the first category, reading a first number of target instructions from the second instruction queue according to the first number of the first entry; in the case that the instruction in the second entry does not include the instruction of the first category, reading a second number of target instructions from the first instruction queue and reading a third number of target instructions from the second instruction queue according to the first number of the first entry, and the first number is the sum of the second number and the third number.

[0053] That is, the instruction input module arbitrates according to the instruction category of the instruction in the second non-idle entry of the instruction entry of the internal storage module; if the second entry includes the instruction of the first category, that is, there is a short cycle instruction in execution, then the long cycle instruction can be read this time, and the arbitration result is to read the instruction of the second category this time. In this way, the execution efficiency of the long cycle instruction can be improved.

[0054] In some possible implementation manners, if the second entry does not include the instruction of the first category, that is, there is no short cycle instruction in execution, then the instruction of the first category needs to be read preferentially in order to ensure the parallel execution of the short cycle instruction, and the arbitration result is to read the instruction of the first category preferentially this time. In this way, the parallel processing of the two categories of instructions can be maintained.

[0055] In some possible implementation manners, the number of instructions read this time can be determined according to the first number of the first entries. If the instructions of the first category are included in the second entries, the first number of target instructions are read from the second instruction queue; and the reserved bits in each target instruction are deleted respectively, and the remaining contents are written into the corresponding first entries as instruction information.

[0056] In some possible implementation manners, if the instructions of the first category are not included in the second entries, the second number of target instructions are read from the first instruction queue, and the third number of target instructions are read from the second instruction queue; and the reserved bits in each target instruction are deleted respectively, and the remaining contents are written into the corresponding first entries as instruction information.

[0057] In some possible implementation manners, the first number is the sum of the second number and the third number. The second number is usually a preset value, for example, 1 or 2, and the third number is the difference between the first number and the second number. In an example, the instructions of the first category are not included in the second entries, and the number of the first entries is 5, and then 1 instruction of the first category and 4 instructions of the second category can be read to obtain 5 target instructions.

[0058] In this way, the instructions of two categories can be stored in the instruction entries of the internal storage module at the same time, parallel processing of instructions of different categories is implemented, and the processing efficiency is improved.

[0059] In some possible implementation manners, reading the first number of target instructions from the second instruction queue includes: starting from a target position of the second instruction queue, selecting the first number of instructions with the execution state bit being the first value as target instructions, and reading the target instructions, and the target position is a next position of a tail position of the last selected instruction.

[0060] That is, the instruction input module reads instructions in a polling manner, that is, starting from a next position of a last read position in the instruction queue, and when the tail position of the instruction queue is read, reading continues from the head position of the instruction queue.

[0061] For the case of reading the first number of target instructions from the second instruction queue, the next position of the tail position of the last selected instruction can be used as the target position; starting from the target position, the first number of instructions with the execution state bit being the first value (not executed) are selected as target instructions and the target instructions are read.

[0062] Correspondingly, the processing manner of reading the second number of target instructions from the first instruction queue and reading the third number of target instructions from the second instruction queue is consistent with the processing manner described above, and details are not described herein again.

[0063] In this way, the polling can be performed on the selected non-executed instruction, the occupation of the instruction entry by the multiple unsuccessful execution of the instruction can be reduced, and the overall execution efficiency of the instruction can be improved.

[0064] In some possible implementation manners, the internal storage module 12 is configured to, in the case where the instruction information of the target instruction is stored in the first entry, send the entry identifier of the first entry to the instruction execution module, so that the instruction execution module initiates a read request for the instruction information of the target instruction.

[0065] That is, after the instruction information of the target instruction is stored in the first entry of the internal storage module 12, the instruction execution module 13 is notified to perform data reading. After receiving the notification of the internal storage module 12, the instruction execution module 13 can send a read request for the instruction information of the target instruction to the internal storage module 12, and the read request includes the entry identifier. The internal storage module 12 responds to the read request and sends the instruction information of the target instruction in the corresponding first entry to the instruction execution module 13. The instruction execution module 13 analyzes the received instruction information, determines the instruction category according to the flag bit of the instruction category in the instruction information, and executes the target instruction according to the instruction category, for example, performs the corresponding data transfer or data comparison.

[0066] In some possible implementation manners, in the case where the target instruction is of the first category, the instruction information includes the data amount of the data to be transferred, the first source address and the destination address of the data to be transferred, and can further include the instruction identifier, the entry identifier, the identifier ID of the address translation page table, and the like, which are not limited in the present disclosure.

[0067] In the case where the target instruction is of the first category, the step of executing the target instruction according to the instruction category of the target instruction includes the following steps. According to the data amount of the data to be transferred, the first source address, and the preset data amount of the data read request, at least one first read request of the target instruction is determined, and the first read request is sent respectively. In the case where the return data of the first read request is received, the return data is cached. In the case where the cached return data is the complete data to be transferred, the data to be transferred is written to the destination address.

[0068] For example, for any target instruction, the instruction execution module 13 analyzes the received instruction information. If the flag bit of the instruction category indicates that the target instruction is of the first category, that is, a DMA instruction, the data can be read according to the source address and the data amount. According to the data amount of the data to be transferred, the first source address, and the preset data amount of the data read request, at least one first read request of the target instruction can be determined.

[0069] In an example, in the case of employing AXI interface, the data volume of the data to be carried by each data carrying instruction is 512 Bytes at most, which can be split into 8 first read requests, that is, 8 read requests are supported as unfinished transactions. The preset data volume of the data read request is 64 Bytes, that is, 64 Bytes are carried by each first read request, which is sent according to 4 time taps, and each time tap is 16 Bytes, that is, 128 bits. Each time tap corresponds to a clock cycle. In this way, 512 Bytes of data to be carried requires a total of 32 time taps.

[0070] In this way, according to the data volume of the data to be carried, the first source address and the preset data volume of the data read request, at least one first read request of the target instruction can be determined respectively. For example, the data volume of the data to be carried is 256 Bytes, and the preset data volume of the data read request is 64 Bytes, then 4 first read requests can be determined, and according to the first source address, the request addresses of the 4 first read requests are determined respectively. Further, each first read request is sent respectively.

[0071] In some possible implementation manners, when the instruction execution module receives the return data of the first read request, the return data can be cached in a cache space (DATA_SRAM) inside the instruction execution module, and the size of the cache space can be 128 bits x 32, so as to be able to accommodate the maximum size of the data to be carried.

[0072] It should be understood that the maximum data volume of the data to be carried and the preset data volume of the data read request can be set by the person skilled in the art according to the actual conditions such as the type of the interface, and the present disclosure does not limit this.

[0073] In some possible implementation manners, when the instruction execution module receives the return data of all first read requests, that is, the cached return data is the complete data to be carried, the instruction execution module can write the data to be carried to the destination address through the AXI interface to complete the execution process of the data carrying instruction.

[0074] In some possible implementation manners, after the instruction execution module receives the information that the write reply of the AXI interface is valid and no error occurs, the instruction execution module sends the interrupt information and the entry identifier of the target instruction to the instruction input module to inform that the execution of the target instruction is completed. In addition, the instruction execution module can also send the entry identifier to the internal storage module, so that the internal storage module sets the corresponding entry as idle to read subsequent instructions.

[0075] In this way, the execution of the data carrying instruction can be realized, the execution demand of high-throughput large-scale data carrying can be met, and the execution efficiency can be improved.

[0076] In some possible implementation manners, in the case where the target instruction is of the second category, the instruction information comprises the data comparison logic, the second source address of the current value, and the dependent value, and can further comprise information such as the instruction identifier, the entry identifier, and the identifier ID of the address translation page table, which are not limited by the present disclosure.

[0077] The data comparison logic can comprise that the current value is greater than the dependent value, the current value is less than the dependent value, the current value is equal to the dependent value, the current value is greater than or equal to the dependent value, the current value is less than or equal to the dependent value, the current value is not equal to the dependent value, and the like, and the present disclosure does not limit the specific data comparison logic.

[0078] In the case where the target instruction is of the second category, the step of executing the target instruction according to the instruction category of the target instruction comprises: reading the current value according to the second source address of the current value; in the case where the current value is read, reading the data comparison logic and the dependent value of the target instruction; and performing data comparison on the current value and the dependent value according to the data comparison logic to obtain the execution result of the target instruction.

[0079] For example, for any target instruction, the instruction execution module 13 parses the received instruction information, and if the flag bit of the instruction category indicates that the target instruction is a VCE instruction of the second category, the instruction execution module can send a read request according to the second source address of the current value, read the current value of the semaphore from the second source address; after the current value is read, the data comparison logic and the dependent value of the target instruction are read from the internal storage module; the current value and the dependent value are compared according to the data comparison logic to obtain the execution result of the target instruction; and then the execution result is sent to the instruction input module 11. The execution result can comprise execution success or execution failure, and if the execution is successful, a corresponding interrupt signal will be sent.

[0080] In some possible implementation manners, the instruction execution module can further send the entry identifier to the internal storage module, so that the internal storage module sets the corresponding entry as idle, so as to read subsequent instructions.

[0081] In this way, parallel processing of data comparison instructions in multiple first entries can be implemented, high-concurrency data comparison can be met, low-latency parallel processing can be implemented, and the execution efficiency of the instructions can be improved.

[0082] In some possible implementation manners, the instruction execution module reads data from the internal storage module twice when executing instructions of the second category, and in the case where multiple instructions are executed in parallel, arbitration is needed, and the second reading is prior to the first reading.

[0083] In some possible implementation manners, the internal storage module is configured to: in a case where the received read request simultaneously includes a second read request of instruction information for the first target instruction and a third read request of data comparison logic and the dependent value for the second target instruction, respond to the third read request.

[0084] That is, for the internal storage module, if the received multiple read requests simultaneously include the second read request of instruction information for the first target instruction and the third read request of data comparison logic and the dependent value for the second target instruction, the third read request can be responded preferentially, and response information of the third read request is sent to the instruction execution module, so that the instruction execution module completes the data comparison process; after the third read request is completed, the second read request is executed, so that the instruction execution module judges the instruction category and processes accordingly.

[0085] In this way, the second read can be preferentially processed, so that the processing efficiency of the executing instruction is quickly completed, and the overall instruction processing efficiency of the system is improved.

[0086] In some possible implementation manners, the instruction input module is further configured to: in a case where an interrupt signal of the target instruction is received, update an execution state bit of the target instruction in the register control module from a first value to a second value.

[0087] That is, if the instruction input module receives an interrupt signal of the target instruction, the execution state bit of the target instruction in the register control module can be updated from the first value to the second value, for example, from 0 to 1, indicating that the target instruction has been executed, and the target instruction will be skipped when the next instruction is selected by polling; conversely, if information that the target instruction fails to execute is received, the value of the corresponding execution state bit is not updated, or the corresponding execution state bit is set to the first value, indicating that the target instruction has not been executed or has failed to execute, and the target instruction will be selected when the next instruction is selected by polling.

[0088] In this way, the execution state can be controlled, so that the situation of repeated selection of instructions and omission of instructions is reduced, and the processing logic is more perfect.

[0089] In some possible implementation manners, the register control module is configured to: in a case where the execution state bit of the target instruction is updated from the first value to the second value, report an interrupt signal of the target instruction.

[0090] That is, if the execution status bit of the target instruction is updated from the first value to the second value, that is, indicates that the target instruction has been executed, the register control module can report an interrupt signal of the target instruction to the firmware, to inform the firmware that the target instruction has been executed and completed, so that the firmware controls the CPU to execute subsequent processing, such as mathematical operation or graphic processing based on the data carried.

[0091] In this way, the reporting of instruction interruption can be realized to execute subsequent processing.

[0092] In some possible implementation manners, the instruction input module is further configured to stop reading instructions when the execution status bits of the instructions in the first instruction queue and the second instruction queue are both the second value.

[0093] That is, if the execution status bits of the instructions in the first instruction queue and the second instruction queue in the register control module are both the second value, that is, the instructions in the two instruction queues have both been executed, the instruction input module does not need to continue to poll and select instructions. In this case, the instruction input module can also stop reading instructions. In this way, the hardware can automatically exit the loop without software intervention, and the system efficiency is improved.

[0094] The data processing apparatus according to the embodiments of the present disclosure can distinguish between different types of instructions, such as short-period VCE instructions and long-period DMA instructions, through double-circulation queues, schedule different categories of instruction queues, improve overall processing efficiency, realize priority arbitration of instructions in different instruction queues through the instruction input module, and poll and select instructions to avoid instruction execution conflicts and resource conflicts, and automatically control the start, exit and status update of the instruction queue through hardware to reduce software intervention.

[0095] In the related art, there is no dedicated hardware resource to poll and compare Semaphore signals, and software control is needed to initiate data reading and comparison, only sequential operation can be performed, a large number of data reading operations cannot be initiated in parallel, the efficiency is low, and there is a similarity between the demand of the on-chip CPU for the DMA engine, which causes partial function redundancy.

[0096] The data processing apparatus according to the embodiments of the present disclosure can realize concurrent control based on hardware, execute instructions of different life cycles at the same time, VCE instructions and DMA instructions can be executed in an out-of-order and concurrent manner, the expansibility is strong, a plurality of instructions are supported, for example, VCE instructions support more complex comparison logic, or DMA instructions carry larger-scale data, and the like; the hardware can support loop self-start and exit, realize automatic control, reduce software overhead on the host side, and reduce software burden on the host side.

[0097] According to the data processing apparatus of the embodiments of the present disclosure, the concurrent processing of multiple unfinished transactions can be supported by identifying the maintenance, and the throughput of data reading and writing and the efficiency of data comparison can be improved. According to the embodiments of the present disclosure, the process of reading memory values and comparison on-chip CPU can be accelerated, and the DMA function can be combined to avoid module function redundancy and improve efficiency.

[0098] Figure 2 A flowchart of a data processing method according to an embodiment of the present disclosure is provided. Referring to Figure 2 The method includes the following steps S21-S23.

[0099] In step S21, according to the tag of the first idle entry in the plurality of instruction entries and the instruction category of the instruction in the second non-idle entry, an unexecuted target instruction is read from the first instruction queue and / or the second instruction queue, and the instruction information of the target instruction is stored in the first entry. The first instruction queue includes instructions of a first category with an execution number of 1, and the second instruction queue includes instructions of a second category with an execution number greater than or equal to 1.

[0100] In step S22, the target instruction is executed according to the instruction category of the target instruction, and the execution result of the target instruction is obtained.

[0101] In step S23, if the execution result of the target instruction is successful, an interrupt signal of the target instruction is reported.

[0102] In some possible implementation manners, the method further includes: in a case where the interrupt signal of the target instruction is received, updating an execution state bit of the target instruction from a first value to a second value, the first value being used to indicate that the instruction is not executed, and the second value being used to indicate that the instruction is executed.

[0103] In some possible implementation manners, step S21 includes: in a case where the instruction in the second entry includes the instruction of the first category, reading a first number of target instructions from the second instruction queue according to the first number of the first entry; in a case where the instruction in the second entry does not include the instruction of the first category, reading a second number of target instructions from the first instruction queue and a third number of target instructions from the second instruction queue according to the first number of the first entry, the first number being a sum of the second number and the third number.

[0104] In some possible implementation manners, step S21 includes: selecting, as a target instruction, an instruction of the first quantity and having the execution state bit being the first value from a target position of the second instruction queue, the target position being a next position of a tail position of a last selected instruction, and reading the target instruction.

[0105] In some possible implementation manners, in a case where the target instruction is of the first category, the instruction information includes a data quantity of to-be-carried data, a first source address of the to-be-carried data, and a destination address, and step S22 includes: determining at least one first read request of the target instruction according to the data quantity of the to-be-carried data, the first source address, and a preset data quantity of a data read request, and respectively sending the first read request; in a case where return data of the first read request is received, caching the return data; in a case where the cached return data is complete to-be-carried data, writing the to-be-carried data to the destination address.

[0106] In some possible implementation manners, in a case where the target instruction is of the second category, the instruction information includes data comparison logic, a second source address of a current value, and a dependent value, and step S22 includes: reading the current value according to the second source address of the current value; in a case where the current value is read, reading the data comparison logic and the dependent value of the target instruction; and performing data comparison on the current value and the dependent value according to the data comparison logic, to obtain an execution result of the target instruction.

[0107] In some possible implementation manners, the method further includes: in a case where the instruction information of the target instruction is stored in the first entry, sending an entry identifier of the first entry; and in a case where a read request received simultaneously includes a second read request for instruction information of a first target instruction and a third read request for data comparison logic and a dependent value of a second target instruction, responding to the third read request.

[0108] In some possible implementation manners, the method further includes: in a case where there is a newly added instruction in the first instruction queue and / or the second instruction queue, setting an execution state bit corresponding to the newly added instruction, and setting the execution state bit to the first value.

[0109] In some possible implementation manners, the method further includes: in a case where the execution state bit of the target instruction is updated from the first value to a second value, reporting an interrupt signal of the target instruction.

[0110] In some possible implementation manners, the method further includes: in a case where the execution state bits of the instructions in the first instruction queue and the second instruction queue are both the second value, stopping reading instructions.

[0111] In some possible implementation manners, the first type of instructions includes direct memory access (DMA) based data transfer instructions; and the second type of instructions includes data comparison unit (VCE) based data comparison instructions.

[0112] It can be understood that the above-mentioned various method embodiments mentioned in the disclosure can be combined with each other to form combined embodiments without deviating from the principle logic. Limited by the length of the disclosure, the disclosure will not be repeated. Those skilled in the art can understand that in the above-mentioned method of the specific embodiment, the specific execution order of each step should be determined according to its function and possible internal logic.

[0113] Figure 3 A block diagram of an electronic device provided by an embodiment of the disclosure.

[0114] With reference to Figure 3 The embodiment of the disclosure provides an electronic device, which comprises: at least one processor 701; at least one memory 702, and one or more I / O interfaces 703 connected between the processor 701 and the memory 702; wherein the memory 702 stores one or more computer programs executable by the at least one processor 701, and the one or more computer programs are executed by the at least one processor 701 to enable the at least one processor 701 to include the above-mentioned data processing apparatus or execute the above-mentioned data processing method.

[0115] The embodiment of the disclosure also provides a computer readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned data processing method when executed by a processor. The computer readable storage medium can be a volatile or non-volatile computer readable storage medium.

[0116] The embodiment of the disclosure also provides a computer program product comprising computer readable code or a non-volatile computer readable storage medium carrying computer readable code, when the computer readable code is executed in the processor of the electronic device, the processor in the electronic device executes the above-mentioned data processing method.

[0117] Those of ordinary skill in the art will realize and understand, all or some of the steps in the methods disclosed above and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation. Certain physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable storage medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media).

[0118] Example embodiments have been disclosed herein and, although the use of specific terms is expressly used herein, they are intended in the sense only of general descriptive purpose and should not be taken as limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics or / and elements described in connection with a particular embodiment can be used in conjunction with other embodiments unless otherwise explicitly stated. As such, those skilled in the art will appreciate that various changes can be made in form and detail without departing from the scope of the disclosure as set forth in the appended claims.

Claims

1. A data processing apparatus, characterized by, The device comprises: an instruction input module, an internal storage module and an instruction execution module; the instruction input module is configured to read an unexecuted target instruction from a first instruction queue and / or a second instruction queue according to a tag of a first entry and an instruction category of an instruction in a second entry in a plurality of instruction entries of the internal storage module, and store instruction information of the target instruction into the first entry; the first instruction queue comprises instructions of a first category with an execution frequency of 1, and the second instruction queue comprises instructions of a second category with an execution frequency greater than or equal to 1; the instruction execution module is configured to execute the target instruction according to the instruction category of the target instruction in the case that the instruction information of the target instruction is stored in the first entry, and send an interrupt signal of the target instruction to the instruction input module in the case that an execution result of the target instruction is successful.

2. The apparatus of claim 1, wherein, The device further comprises a register control module, and the instruction input module is further configured to: update an execution state bit of the target instruction in the register control module from a first value to a second value in the case that the interrupt signal of the target instruction is received, the first value being used to indicate that an instruction is not executed, and the second value being used to indicate that an instruction is executed.

3. The apparatus of claim 2, wherein, The instruction input module reads an unexecuted target instruction from a first instruction queue and / or a second instruction queue according to a tag of a first entry and an instruction category of an instruction in a second entry in a plurality of instruction entries of the internal storage module, comprising: in the case that the instruction in the second entry comprises the instruction of the first category, reading a first number of target instructions from the second instruction queue according to the first number of the first entry; in the case that the instruction in the second entry does not comprise the instruction of the first category, reading a second number of target instructions from the first instruction queue and a third number of target instructions from the second instruction queue according to a first number of the first entry, the first number being a sum of the second number and the third number.

4. The apparatus of claim 3, wherein, The reading of the first number of target instructions from the second instruction queue comprises: starting from a target position of the second instruction queue, selecting the first number of instructions with a first value of an execution state bit as target instructions, and reading the target instructions, the target position being a next position of a last position of a previously selected instruction.

5. The apparatus of claim 1, wherein, in the case that the target instruction is of the first category, the instruction information comprises a data amount of to-be-carried data, a first source address of the to-be-carried data and a destination address, wherein the instruction execution module executes the target instruction according to the instruction category of the target instruction, comprising: determining at least one first read request of the target instruction according to the data amount of the to-be-carried data, the first source address and a preset data amount of a data read request, and sending the first read request respectively; in the case that return data of the first read request is received, caching the return data; In a case where the cached return data is complete data to be carried, the data to be carried is written to the destination address.

6. The apparatus of claim 1, wherein, In a case where the target instruction is of a second type, the instruction information includes data comparison logic, a second source address of a current value, and a dependent value, and the instruction execution module executes the target instruction according to the instruction type of the target instruction, including: reading the current value according to the second source address of the current value; in a case where the current value is read, reading the data comparison logic and the dependent value of the target instruction; performing data comparison on the current value and the dependent value according to the data comparison logic to obtain an execution result of the target instruction.

7. The apparatus of claim 6, wherein, The internal storage module is configured to: in a case where the instruction information of the target instruction is stored in the first entry, sending an entry identifier of the first entry to the instruction execution module to enable the instruction execution module to initiate a read request for the instruction information of the target instruction; in a case where the received read request includes a second read request for instruction information of a first target instruction and a third read request for data comparison logic and a dependent value of a second target instruction, responding to the third read request.

8. The apparatus of claim 2, wherein, The instruction input module is further configured to, in a case where there is a new instruction in the first instruction queue and / or the second instruction queue, set an execution state bit corresponding to the new instruction in the register control module, and set the execution state bit to the first value.

9. The apparatus of claim 2, wherein, The register control module is configured to, in a case where the execution state bit of the target instruction is updated from a first value to a second value, report an interrupt signal of the target instruction.

10. The apparatus of claim 3, wherein, The instruction input module is further configured to, in a case where the execution state bits of instructions in the first instruction queue and the second instruction queue are both the second value, stop reading instructions.

11. The apparatus of claim 1, wherein, The first type of instruction includes a direct memory access (DMA) based data carrying instruction, and the second type of instruction includes a data comparison unit (VCE) based data comparison instruction.

12. A data processing method, characterized by, including: reading an unexecuted target instruction from the first instruction queue and / or the second instruction queue according to an entry identifier of a first entry that is free and an instruction type of an instruction in a second entry that is not free, and storing instruction information of the target instruction to the first entry; wherein the first instruction queue includes a first type of instruction with an execution number of 1, and the second instruction queue includes a second type of instruction with an execution number greater than or equal to 1; executing the target instruction according to the instruction type of the target instruction to obtain an execution result of the target instruction; in a case where the execution result of the target instruction is execution success, reporting an interrupt signal of the target instruction.

13. An electronic device, comprising: The data processing apparatus of any one of claims 1-11. The data processing apparatus of any one of claims 1-11.

Citation Information

Patent Citations

  • Data processing method and electronic device

    CN105243033A

  • Processor and method with function of reducing pipeline stall caused by full load of loading queue

    CN106951215A

  • Method and system for accessing multiple different types of Internet of Things devices to Internet of Things

    CN110417629A

  • Instruction cache filling and filtering device

    CN110737475A

  • Instruction allocation method and device, electronic equipment and readable storage medium

    CN114579187A