Co-processor Based on Data-Driven and Cooperative Work of Software and Hardware and Its Control Method

By designing a coprocessor based on data-driven and software and hardware collaboration, the design complexity and resource waste of existing coprocessors are solved, efficient data storage and transmission and reception are achieved, the work efficiency of the multi-core computing platform is improved, and it is suitable for communication signal processing and image recognition fields.

CN115113937BActive Publication Date: 2025-07-22UNIV OF ELECTRONICS SCI & TECH OF CHINA +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210621088.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2025-07-22
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

Existing coprocessors have problems such as complex design, high resource usage, insufficient configurability and scalability, poor data operation flexibility, and waste of storage resources.

Method used

A coprocessor based on data driver and software and hardware collaboration is designed, including a control unit, a data receiving module, a data sending module and a data storage module. The workflow is controlled through a state machine, and multiple rotational result storage areas and token tag matching mechanisms are used to realize data storage and transmission and reception.

Benefits of technology

It improves the computing performance of the processing core and the overall working efficiency of the system, has strong versatility and scalability, and is suitable for multi-core computing platforms, especially in the fields of communication signal processing and image recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115113937B_ABST
    Figure CN115113937B_ABST
Patent Text Reader

Abstract

The present invention discloses a coprocessor based on data-driven and software-hardware collaborative work and its control method. The coprocessor includes a control unit, a data receiving module, a data sending module, and a data storage module. The data receiving module is used to receive data from outside the coprocessor and store it in the data storage module. The data sending module is used to read out data from the data storage module and send it outside the coprocessor. The control unit is used to control the working process of the coprocessor. The coprocessor of the present invention works in collaboration with a vector processor to complete the storage and transceiver of data, and can improve the computing performance of the processing core and the overall working efficiency of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of coprocessors, and particularly relates to a coprocessor based on data driving and software-hardware cooperation and a control method thereof. Background Art

[0002] Due to the limitations of factors such as power consumption, cost, and process, multi-core processors on a chip have replaced high-frequency single-core processors as the mainstream in the market. With the increase in the number of processor cores, the need for multi-core cooperation and multi-thread parallelism has emerged. Coprocessors, due to their high-performance and low-power computing structures for specific applications, have gradually become an indispensable part of the multi-core computing platform on a chip, which can help the processor better complete computing tasks and improve hardware efficiency.

[0003] Efficient coprocessors have a wide range of applications in various fields. For example, for the existing problems in the current mainstream communication coprocessor architecture, such as large interconnection network power consumption and frequent scheduling, a new type of two-dimensional configurable coprocessor architecture for communication processors is proposed. By configuring the working mode and parameters, the power consumption of the bus interconnection network is reduced to 1 / 3 of the mainstream architecture. The existing packet classification coprocessor includes modules such as a multi-thread buffer module, a keyword processing unit module, and a result buffer. A hardware packet classification coprocessor that improves the packet classification technology at the data level and can assist devices such as network processors in packet classification further reduces the impact of range expansion. However, the disadvantage of this coprocessor is that it can only implement basic read / write and packet classification functions, and more functions need to be expanded. In addition, the existing research on important modules in the 10G HIMAC coprocessor helps to achieve the goal of "triple play". The 10G MAC core module, lookup table module, and FPGA high-speed serial interface based on the RocketIO module are designed and implemented. The designed coprocessor can support a data rate processing capacity of 10Gbps and functions such as unicast, multicast, and broadcast. The defect of this coprocessor is that the logic design has a certain complexity, the resource occupancy is slightly high, and the configurability and scalability are insufficient.

[0004] Coprocessors are also applied in the field of image recognition. For example, for the characteristics of large data volume, large amount of computation, and variable processing flow in image target detection, a scalable coprocessor architecture for target detection and a command packet format for implementing this architecture are proposed, enabling the microprocessor to call the hardware acceleration circuit by referring to the method of software function call and realizing the parallel operation of multi-functional IPs. It has the characteristics of strong scalability and versatility, and at the same time has the advantages of low power consumption and small area. However, in order to reduce complexity, the single data bit width of this coprocessor is set to 256bit, which limits the flexibility of data operation and also causes a certain waste of storage resources. Summary of the Invention

[0005] Therefore, to solve the problems of complex design, high resource occupation, or waste of storage resources existing in existing coprocessors, the present invention provides a coprocessor based on data-driven and software-hardware collaborative work. The coprocessor of the present invention works in cooperation with a vector processor to complete data storage, reception, and transmission, and can improve the computing performance of the processing core and the overall working efficiency of the system.

[0006] The present invention is implemented through the following technical solutions:

[0007] A coprocessor based on data-driven and software-hardware collaborative work includes a control unit, a data reception module, a data transmission module, and a data storage module;

[0008] Among them, the data reception module is used to receive data from outside the coprocessor and store it in the data storage module;

[0009] The data transmission module is used to read out data from the data storage module and send it outside the coprocessor;

[0010] The control unit is used to control the working process of the coprocessor.

[0011] As a preferred embodiment, the data storage module of the present invention is composed of 16 sub-RAMs, and one scalar data is stored in one address space of each sub-RAM.

[0012] As a preferred embodiment, each sub-RAM of the present invention includes a data stack area, a task data storage area, and an operation result storage area;

[0013] Among them, the data stack area is used to store temporary data generated by the vector processor during the operation process;

[0014] The task data storage area is used to store various variables during the operation process;

[0015] The operation result storage area adopts a form of multiple rotations.

[0016] As a preferred embodiment, the form of multiple rotations of the present invention is specifically:

[0017] Multiple identical result storage areas form a rotation queue;

[0018] After the result of the current operation is generated, another result storage area is rotated up to store the result of the next operation;

[0019] The result storage area of the current operation is called by the data transmission module. After the data transmission module completes the data transmission operation, the result storage area of the current operation will enter the rotation queue and wait for a new operation result to be written.

[0020] As a preferred embodiment, the control unit of the present invention controls the working process of the coprocessor through a state machine;

[0021] The state machine includes an IDLE state, an INIT state, a LOAD state, and a WORK state;

[0022] Among them, the IDLE state is the initial state, and the coprocessor is in a standby state;

[0023] INIT is the initialization state of the coprocessor, and the coprocessor loads the working parameter settings;

[0024] After the initialization is completed, the state machine is in the LOAD state. In this state, the coprocessor downloads the original data to be processed from the host computer, and at the same time uploads the operation results back to the host computer;

[0025] After the data download is completed, the state machine enters the WORK state. In this state, the coprocessor sends a control signal to the vector processor to drive the vector processor to start task operations. While the vector processor is performing operations, the coprocessor also needs to receive the operation data from other vector processors, store it in the data storage module, and complete the task of data transmission according to the instructions given by the vector processor.

[0026] In a second aspect, the present invention also proposes a control method based on the above coprocessor, including:

[0027] By checking data completeness, establish corresponding token tags to achieve data driving.

[0028] As a preferred embodiment, the data driving of the present invention specifically includes:

[0029] The external data packets received by the coprocessor are matched with the matching component for token tags. When the matching target value is consistent with the algorithm requirement target value of the vector processor, the coprocessor gives an instruction to start the operation of the vector processor, driving the vector processor to enter the working state and start the operation.

[0030] As a preferred embodiment, the token tag matching of the present invention specifically includes:

[0031] When the coprocessor receives data, adjust the value in the corresponding token according to the task information carried by the data itself. When the token is consistent with the token value of the vector processor, the data correlation matching is completed, indicating that the operation required data received is complete, and the operation starts.

[0032] As a preferred embodiment, the tokens of the present invention are sorted according to the task order of the required operations; and according to the sorting, the tokens at the top are matched in turn.

[0033] In a third aspect, the present invention proposes a control method based on the above co-processor, including:

[0034] Receiving a data sending instruction given by the vector processor, and reading and sending data from the data storage module according to the data sending instruction.

[0035] The present invention has the following advantages and beneficial effects:

[0036] The co-processor provided by the present invention can work in cooperation with the vector processor, and can be used for storing, receiving, and transmitting core data in a multi-core operation platform, improving the working efficiency of the operation platform.

[0037] The co-processor provided by the present invention has strong versatility and scalability, can be widely applied to fields such as communication signal processing, image recognition, data encryption, etc., can help various vector processors efficiently complete operation tasks, and can simultaneously quickly expand and form a multi-core operation platform based on the processing core. Description of the Drawings

[0038] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not constitute a limitation to the embodiments of the present invention. In the drawings:

[0039] Figure 1 It is a schematic structural diagram of the co-processor according to an embodiment of the present invention.

[0040] Figure 2 It is a schematic structural diagram of the data storage module according to an embodiment of the present invention.

[0041] Figure 3 It is a schematic diagram of the working process of the result storage area rotation queue according to an embodiment of the present invention.

[0042] Figure 4 It is the working state machine of the control unit according to an embodiment of the present invention.

[0043] Figure 5 It is an example of token matching between the co-processor and the vector processor according to an embodiment of the present invention.

[0044] Figure 6 It is a schematic diagram of the working process of token sequence matching of the co-processor according to an embodiment of the present invention.

[0045] Figure 7 It is the instruction information received by the co-processor according to an embodiment of the present invention. Detailed Embodiments

[0046] Hereinafter, the term "comprising" or "may comprise" used in various embodiments of the present invention indicates the presence of the functions, operations, or elements of the present invention, and does not limit the addition of one or more functions, operations, or elements. Further, as used in various embodiments of the present invention, the terms "comprising", "having", and their cognates are only intended to indicate a specific feature, number, step, operation, element, component, or combination of the foregoing items, and should not be construed as precluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing items or the possibility of adding one or more features, numbers, steps, operations, elements, components, or combinations of the foregoing items.

[0047] In various embodiments of the present invention, the expression "or" or "at least one of A or / and B" includes any combination or all combinations of the recited words. For example, the expression "A or B" or "at least one of A or / and B" may include A, may include B, or may include both A and B.

[0048] Expressions (such as "first", "second", etc.) used in various embodiments of the present invention may modify various constituent elements in the various embodiments, but do not limit the corresponding constituent elements. For example, the above expressions do not limit the order and / or importance of the elements. The above expressions are only for the purpose of distinguishing one element from other elements. For example, the first user device and the second user device indicate different user devices, although both are user devices. For example, without departing from the scope of the various embodiments of the present invention, the first element may be referred to as the second element, and similarly, the second element may also be referred to as the first element.

[0049] It should be noted that: if it is described that one constituent element is "connected" to another constituent element, the first constituent element may be directly connected to the second constituent element, and a third constituent element may be "connected" between the first constituent element and the second constituent element. Conversely, when one constituent element is "directly connected" to another constituent element, it can be understood that there is no third constituent element between the first constituent element and the second constituent element.

[0050] The terms used in various embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the various embodiments of the present invention. As used herein, the singular form is also intended to include the plural form unless the context clearly indicates otherwise. Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of the present invention pertain. The terms (such as those defined in a general use dictionary) will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present invention.

[0051] To make the objectives, technical solutions, and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and are not intended to limit the present invention.

[0052] Embodiment 1

[0053] In view of the problems existing in the existing coprocessors, such as complex logic design, high resource occupancy, insufficient configurability and scalability, poor flexibility in data operations, or waste of storage resources, this embodiment provides a coprocessor based on data-driven and software-hardware collaborative work and its control method, specifically as Figure 1 shown, the coprocessor of this embodiment mainly consists of a control unit, a data receiving module, a data sending module, and a data storage module.

[0054] Among them, the control unit is responsible for the overall operation of the coprocessor. The data receiving module is responsible for receiving data and storing it in the data storage module. The data sending module is responsible for reading data from the data storage module and sending it. The data storage module is responsible for data storage.

[0055] The data storage module of this embodiment is used to store the data required for operations. The design of the data storage module needs to be combined with the working requirements of the coprocessor to meet the operation needs of the processing core. Therefore, in the data storage module of this embodiment, each address space can store 1 vector data, and a single vector data is composed of 16 scalar data spliced together. The data storage module consists of 16 sub-RAMs. One address space of each sub-RAM stores one scalar data, thus allowing the granularity of data processing to be refined to each scalar number, meeting the requirements of the processing core for data flexibility. As Figure 2 shown in the structure of the data storage module, where STACK is the data stack area, which is used for temporary data generated during the operation of the vector processor. TASK is the storage area for task data, which is used to store various variables during the operation.

[0056] In order to improve the utilization rate of the vector processor and the working efficiency of the coprocessor as much as possible, Output Data is designed as the storage area for operation results. The physical implementation of the result storage area adopts a form of multiple rotations. Multiple identical result storage areas form a rotation queue, as Figure 3As shown. After the result of the current operation is generated, another result storage area will be rotated up to store the result of the next operation. The result storage area of the current operation will be called by the data sending module. After the data sending module completes the data sending operation, the result storage area will enter the rotation queue and wait for the new operation result to be written. The design of the rotated result storage area adopts the method of parallel storage area space, which enables the vector processor to quickly start the operation of new task data after completing the operation of the current task data, avoiding the time wasted by the vector processor waiting for the coprocessor to send data. By balancing the hardware area and speed, the utilization rate of the vector processor is improved.

[0057] The data receiving module of this embodiment is used to receive data from outside the coprocessor and store it in the data storage module; the data sending module of this embodiment reads data from the data storage module and sends it outside the coprocessor.

[0058] The control unit of this embodiment controls the working process of the coprocessor through a state machine. The state machine is as Figure 4 shown. The IDLE state is the initial state, and the coprocessor is in the standby state. INIT is the initialization of the coprocessor, loading the working parameter settings. After the initialization is completed, the control unit starts to load data, and the state machine is in the LOAD state. In this state, the coprocessor will download the original data to be processed from the host computer at the same time, and will also upload the operation result back to the host computer. After the data download is completed, the state machine enters the vector processor operation state, that is, WORK. The coprocessor sends a control signal to the vector processor, and the vector processor starts the task operation. While the vector processor is performing the operation, the coprocessor also needs to receive the operation data from other vector processors and store it in the storage module, and complete the task of data sending according to the instructions given by the vector processor.

[0059] In order to improve the working efficiency of the coprocessor as much as possible. The control unit of this embodiment adopts the design idea of overlapping the secondary working processes, so that the working processes of downloading the original data and uploading the calculation results overlap, and the working processes of vector processor operation and coprocessor receiving and sending data overlap, improving the hardware utilization rate of the coprocessor data storage and receiving module, greatly compressing the delay of internal data processing of the coprocessor, so that the operation speed of the algorithm can be improved, the task operation can be accelerated, and the time overhead can be reduced.

[0060] Embodiment 2

[0061] This embodiment realizes data-driven control based on the coprocessor proposed in the above Embodiment 1.

[0062] The coprocessor realizes data-driven through data completeness check and establishing corresponding token tags. Specifically:

[0063] According to the pre-established token tags, the received external data packets are matched with the matching components. When the matching target value is consistent with the algorithm requirement target value of the vector processor, an instruction to start the operation of the vector processor is issued, thereby driving the vector processor into the working state to start the operation.

[0064] The process of matching based on token tags proposed in this embodiment is as Figure 5 shown:

[0065] When the coprocessor receives data, it adjusts the value in the corresponding token according to the task information carried by the data itself. When the token is consistent with the token value of the vector processor, the data correlation matching is completed, indicating that the data required for the operation received is complete and the operation can start.

[0066] In this embodiment, a token sequence is designed in the coprocessor, where the tokens are sorted according to the task order of the required operations. The matching process of each token in the sequence is as Figure 5 shown. During the process of receiving data, the coprocessor will modify the value in the corresponding token according to the task information corresponding to the received data. After the token at the top of the sequence completes the matching, the next token will be matched, and at the same time, the token at the top of the sequence exits the sequence. This process is as Figure 6 shown.

[0067] The superiority of using data-driven in this embodiment lies in avoiding complex and difficult-to-implement operation start control logic. Due to the diversity of the vector processors working together and the diversity of the programs running on the processors, it is difficult for the coprocessor to judge whether the received data is correct and the corresponding processing tasks through the way of hardware-form instruction control. However, using the data method can specify the correspondence between data and operation tasks at the software level, and complete the judgment of data correctness and completeness according to the information of the received data itself, avoiding redundant judgment logic.

[0068] Embodiment 3

[0069] This embodiment realizes software and hardware collaborative control based on the coprocessor proposed in the above Embodiment 1.

[0070] The coprocessor in this embodiment adopts a software and hardware collaborative working mode. The coprocessor performs corresponding processing according to the data received from the outside and the instructions given by the vector processor, and gives control signals to assist the vector processor to complete the algorithm operation.

[0071] Specifically, after the vector processor completes the operation of the current task data, the vector processor sends a data sending instruction to the coprocessor. The coprocessor reads the data from the storage module through the instruction and performs the sending work. The data sending instruction information from the vector processor is given by the algorithm personnel during software programming to indicate the data sending process of the coprocessor. The instruction information sent to the coprocessor includes data information and target information, as Figure 7 shown. The data information gives the storage address and data volume of the data to be sent, and the target information gives the sending target, write address, and data number of the data.

[0072] After the vector processor completes the current operation, it sends instruction information to the coprocessor. The coprocessor decomposes and decodes the instruction. The coprocessor first finds the storage address of the data to be sent, reads out the data according to the data volume. Then, in combination with the target information in the instruction, the data is packaged and sent outside the coprocessor to complete the data sending process. If multiple groups of data need to be sent to multiple different targets outside, the vector processor gives multiple instruction information, and the coprocessor will establish a sending queue to complete the data sending in sequence.

[0073] The signal control method combining software and hardware in this embodiment greatly improves the control ability of the vector processor over the data sending of the coprocessor. Specifying the relevant information of data sending at the software level avoids the fixed hardware circuit and improves the flexibility of data sending.

[0074] The specific implementation manners described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific implementation manners of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A coprocessor based on data-driven and collaborative work of software and hardware, characterized in that, It includes a control unit, a data receiving module, a data sending module, and a data storage module; Among them, the data receiving module is used to receive data from outside the coprocessor and store it in the data storage module; The data sending module is used to read out data from the data storage module and send it outside the coprocessor; The control unit is used to control the working process of the coprocessor; The coprocessor realizes data driving by establishing corresponding token tags through data integrity check. The data driving specifically includes: The data packets received from the outside by the coprocessor are matched with the matching component for token tags. When the matching target value is consistent with the algorithm requirement target value of the vector processor, the coprocessor gives an instruction to start the operation of the vector processor, driving the vector processor to enter the working state and start the operation.

2. The coprocessor based on data-driven and software-hardware collaborative work according to claim 1, wherein The data storage module consists of 16 sub-RAMs, and each address space of each sub-RAM stores a scalar data.

3. The coprocessor based on data-driven and software-hardware collaborative work according to claim 2, wherein Each sub-RAM includes a data stack area, a task data storage area, and an operation result storage area; Among them, the data stack area is used to store the temporary data generated by the vector processor during the operation; The task data storage area is used to store each variable during the operation; The operation result storage area adopts a form of multiple rotations.

4. The co-processor based on data-driven and software-hardware collaborative work according to claim 3, characterized in that, The form of multiple rotations is specifically: Multiple identical result storage areas form a rotation queue; After the result of the current operation is generated, another result storage area is rotated up to store the result of the next operation; The result storage area of the current operation is called by the data sending module. After the data sending module completes the data sending operation, the result storage area of the current operation will enter the rotation queue and wait for the new operation result to be written.

5. The coprocessor based on data-driven and software-hardware collaborative work according to claim 1, wherein The control unit controls the working process of the coprocessor through a state machine; The state machine includes an IDLE state, an INIT state, a LOAD state, and a WORK state; Among them, the IDLE state is the initial state, and the coprocessor is in the standby state; INIT is the initialization state of the coprocessor, and the coprocessor loads the working parameter settings; After the initialization is completed, the state machine is in the LOAD state. In this state, the coprocessor downloads the original data to be processed from the host computer, and at the same time uploads the operation results back to the host computer; After the data download is completed, the state machine enters the WORK state. In this state, the coprocessor sends a control signal to the vector processor to drive the vector processor to start the task operation. While the vector processor is performing the operation, the coprocessor also needs to receive the operation data from other vector processors, store it in the data storage module, and complete the task of data sending according to the instructions given by the vector processor.

6. The coprocessor based on data driving and software and hardware collaborative work according to any one of claims 1-5, characterized in that The token tag matching specifically includes: When the coprocessor receives data, it adjusts the value in the corresponding token according to the task information carried by the data itself. When the token is consistent with the token value of the vector processor, the data correlation matching is completed, indicating that the data required for the operation is complete, and the operation begins.

7. The coprocessor based on data-driven and software-hardware collaborative work according to claim 6, wherein The tokens are sorted according to the task order of the required operations; and according to the sorting, the tokens at the top are matched in sequence.

8. The control method of the coprocessor according to any one of claims 1-7, characterized in that Including: Receive the data sending instruction given by the vector processor, and read and send the data from the data storage module according to the data sending instruction.

Citation Information

Patent Citations

  • Image compression coprocessor with data flow control and multiple processing units

    US5699460A