Data processing method and processing system based on heterogeneous fusion multi-core
The method and system for heterogeneous multi-core processors improve inter-core communication efficiency by using a shared cache and cache adaptation unit to manage system states, reducing overheads and enhancing parallel processing capabilities.
Patent Information
- Application Number
- CN202211472080.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-11-23
AI Technical Summary
Existing multi-core processor systems face inefficiencies in inter-core communication, leading to system chaos and performance degradation due to high time and space overheads from interrupt-based or polling mechanisms, which are common in current heterogeneous multi-core processor inter-communication mechanisms.
A data processing method and system for heterogeneous fused multi-core processors that utilize a shared cache and a cache adaptation unit to manage system states through a unified state code, reducing overheads by eliminating the need for interrupt-based communication and enabling efficient, stable inter-core communication.
The proposed method and system enhance inter-core communication efficiency, reducing system overheads and complexity while allowing for parallel processing and improved performance in heterogeneous multi-core systems.
Smart Images

Figure CN115795392B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of heterogeneous multi-core processors, in particular to a data processing method and a processing system based on heterogeneous fusion multi-core. Background Art
[0002] In recent years, technologies such as artificial intelligence, cloud computing, and machine learning have emerged and developed rapidly like bamboo shoots after a spring rain. The applications of various intelligent algorithms exhibit complex and diverse characteristics, such as video image processing, speech recognition, data analysis and transmission, etc. Facing the diverse requirements such as high parallelism and high throughput computing of related programs, a single-structured processor core can no longer meet the needs. Currently, using data-intensive intelligent co-processors for collaborative acceleration to build heterogeneous multi-core processors has become one of the mainstream hardware accelerator design methods for deep learning. By integrating the high-efficiency performance of intelligent co-processors for data-intensive operations and the significant advantages of the main control processor in control-intensive tasks, it can effectively accelerate related algorithms in the field of artificial intelligence.
[0003] Compared with single-core processors and homogeneous multi-core processors, heterogeneous multi-core processors have attracted much attention in improving processor performance. Therefore, the inter-core interaction mechanism of heterogeneous multi-core processors has also become a hot topic of concern and research in the field of processor architectures at home and abroad. In a multi-core processor system, each core shares resources such as caches, external storage spaces, and network controllers. If effective communication cannot be carried out between cores, it will surely lead to chaos in the overall system, thereby affecting the overall performance of the system.
[0004] In the related technical solutions, the inter-core interaction mechanism of multi-core heterogeneous processors is usually implemented based on interrupts or polling. For the inter-core interaction mechanism based on interrupts, the main control processor needs to respond to the corresponding interrupt program, which will bring additional time and space overhead to the system; while the polling mechanism mainly judges whether there is a communication request to be processed by continuously checking the flag signal. If the flag signal is set, the corresponding data flow or process execution is triggered. Summary of the Invention
[0005] In view of this, in order to at least partially solve one of the above technical problems or defects, the purpose of the embodiments of the present invention is to provide a data processing method based on heterogeneous fusion multi-core, which can effectively reduce additional time and space overhead, and can ensure efficient and stable communication between cores. Embodiments of the technical solutions of this application also provide a corresponding data processing system based on heterogeneous fusion multi-core.
[0006] On the one hand, the technical solution of this application provides a data processing method based on heterogeneous fusion multi-core, including the following steps:
[0007] Obtain the data to be operated and the data weight, generate activation data according to the data to be operated and the data weight, and generate instruction data for the activation data;
[0008] Write the activation data and the instruction data into the shared cache, and change the status code of the shared cache after the data writing is completed;
[0009] Obtain the changed status code by polling, and arbitrate the status of the shared cache according to the status code;
[0010] According to the result of the status arbitration, output the activation data to the first intelligent coprocessor, and perform arithmetic processing according to the instruction data; or, determine that the first intelligent coprocessor is in the arithmetic state, and according to the result of the status judgment, output the activation data to the second intelligent coprocessor, and perform arithmetic processing according to the instruction data;
[0011] Write the target data obtained by the arithmetic processing into the shared cache, determine that the arithmetic is completed according to the status code, and output the target data.
[0012] In a feasible embodiment of the solution of the present application, the status code includes a read availability flag bit, a read-write status flag bit, a processor status flag bit, and a cache adaptation status flag bit;
[0013] The step of outputting the activation data to the first intelligent coprocessor according to the result of the status judgment, and performing arithmetic processing according to the instruction data; or, determining that the first intelligent coprocessor is in the arithmetic state, and outputting the activation data to the second intelligent coprocessor according to the result of the status judgment, and performing arithmetic processing according to the instruction data, includes the following steps:
[0014] Determine that the read availability flag bit, the read-write status flag bit, and the processor status flag bit are all 0, and trigger a data transfer signal;
[0015] Determine that the read availability flag bit and the read-write status flag bit are both 0 and the cache adaptation status flag bit is 1. According to the data transfer signal, output the activation data to the first intelligent coprocessor or the second intelligent coprocessor for arithmetic processing, and set the read availability flag bit to 0.
[0016] In a feasible embodiment of the solution of the present application, the step of writing the target data obtained by the arithmetic processing into the shared cache, determining that the arithmetic is completed according to the status code, and outputting the target data, includes:
[0017] Determine that the read availability flag bit and the read-write status flag bit are both 1, write the target data obtained by the arithmetic processing into the shared cache, and set both the read availability flag bit and the cache adaptation status flag bit to 0;
[0018] Determine that the read / write status flag bit is 1, the read availability flag bit and the cache adaptation status flag bit are both 0, output the target data, and set both the read / write status flag bit and the processor status flag bit to 0.
[0019] In a feasible embodiment of the solution of the present application, the obtaining the changed status code by polling and arbitrating the status of the shared cache according to the status code includes:
[0020] Read the status code according to the polling start signal;
[0021] Arbitrate the obtained status code, and obtain the source address and the number of bytes of the activation data from the shared cache according to the arbitration result;
[0022] Read the activation data into the first intelligent coprocessor or the second intelligent coprocessor according to the source address and the number of bytes.
[0023] In a feasible embodiment of the solution of the present application, the writing the target data obtained by arithmetic processing into the shared cache, determining that the arithmetic operation is completed according to the status code, and outputting the target data includes:
[0024] Determine the write-back address of the target data and temporarily store the write-back address;
[0025] Make the temporarily stored write-back address flow into the destination address register, and write the target data back to the shared cache through the destination address register.
[0026] On the other hand, a data processing system based on heterogeneous fusion multi-core is provided in the technical solution of the present application. The system includes:
[0027] A main control processor; configured to obtain data to be operated and data weights, generate activation data according to the data to be operated and the data weights, and generate instruction data for the activation data;
[0028] A cache adaptation unit; configured to write the activation data and the instruction data into a shared storage unit, modify the status code of the shared storage unit; and obtain the status code of the shared storage unit by polling and arbitrate the status of the shared storage unit; and is also configured to write the target data obtained by arithmetic processing into the shared storage unit, and determine that the arithmetic operation is completed according to the status code and output the target data;
[0029] A shared storage unit; configured to store the activation data, the instruction data, and the target data;
[0030] A number of intelligent coprocessors; configured to read the activation data and the instruction data, and obtain the target processing through parallel computing.
[0031] In a feasible embodiment of the solution of the present application, the intelligent coprocessors and the cache adaptation units are in one-to-one correspondence; the intelligent coprocessor includes an activation data storage area and a control instruction storage area;
[0032] The activation data storage area is used to cache the activation data, and the control instruction storage area is used to cache the instruction data.
[0033] In a feasible embodiment of the solution of the present application, the cache adaptation unit includes a first-in-first-out memory, a status code register, a status arbiter, and a control unit;
[0034] The control unit is connected to the main control processor through an on-chip bus; the status code register is connected to the shared storage unit in a cache-direct connection manner; the status code register is also connected to the status arbiter, the status arbiter is connected to the control unit, and the data path of the first-in-first-out memory is connected to the intelligent coprocessor;
[0035] The status code register is configured to obtain and temporarily store the status code of the shared storage unit;
[0036] The status arbiter is configured to arbitrate the status of the shared storage unit;
[0037] The first-in-first-out memory is configured to write the activation data and the instruction data into the shared storage unit; and is also configured to write the target data obtained through arithmetic processing into the shared storage unit; and output the target data;
[0038] The control unit is configured to obtain the source address and the number of bytes from the main control processor and perform data transfer.
[0039] In a feasible embodiment of the solution of the present application, the control unit further includes:
[0040] A source address temporary storage unit, configured to obtain the source address and the number of bytes in the shared storage unit before data transfer;
[0041] A logic controller, configured to control the byte counter and the source address register to perform logical operations;
[0042] A byte counter, configured to perform arithmetic operations on the number of bytes;
[0043] A source address register, configured to perform arithmetic operations on the source address.
[0044] In a feasible embodiment of the solution of the present application, the control unit further includes a destination address register;
[0045] The destination address register is used to determine the write-back address of the target data, temporarily store the write-back address, and write the target data back to the cache adaptation unit according to the write-back address.
[0046] The advantages and beneficial effects of the present invention will be partially given in the following description, and the other parts can be understood through the specific implementation manners of the present invention:
[0047] In the technical solution of the present application, by introducing the recognition and arbitration of status codes at the upper level of the intelligent coprocessor, the migration of data from the shared cache to the intelligent coprocessor is realized, and at the same time, the parallel operation of the intelligent coprocessor is realized according to the status codes. Different from the traditional inter-core communication methods based on polling or interrupts, the inter-core interaction mechanism constructed by the technical solution of the present application is uniformly managed by the system status code, avoiding the additional space-time overhead of the system and reducing the complexity of system management. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0049] Figure 1 It is a schematic structural diagram of a data processing system based on heterogeneous fusion multi-cores provided by the technical solution of the present application;
[0050] Figure 2 It is a schematic structural diagram of the cache adaptation unit in the technical solution of the present application;
[0051] Figure 3 It is a schematic execution flow diagram of the cache adaptation unit in the technical solution of the present application;
[0052] Figure 4 It is a schematic diagram of the address allocation of the shared storage space in the technical solution of the present application;
[0053] Figure 5 It is a schematic diagram of the status code mechanism in the technical solution of the present application;
[0054] Figure 6 It is a schematic calculation flow diagram of a multi-core heterogeneous fusion multi-core processor in the technical solution of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0056] It should be noted that the terms "first", "second", etc. used in this application may be used in this document to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. In addition, the terms "at least one", "a plurality of", "each", "any one", etc. used, at least one includes one, two or more than two, a plurality includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.
[0057] As pointed out in the background art section above, in a multi-core processor system, each core shares resources such as caches, external storage spaces, network controllers, etc. If effective communication cannot be carried out between the cores, it will surely lead to chaos in the overall system, thus affecting the overall performance of the system. And the existing inter-core interaction mechanisms of multi-core heterogeneous processors are usually implemented based on interrupts or polling. For the inter-core interaction mechanism based on interrupts, the main control processor needs to respond to the corresponding interrupt program, which will bring additional temporal and spatial overheads to the system; while the polling mechanism mainly judges whether there is a communication request to be processed by continuously checking the flag signal. If the flag signal is set, the corresponding data flow or process execution is triggered. The polling mechanism brings less additional overhead to the system than interrupt processing; but how to design the mapping relationship between the flag signal and the system state to further reduce the additional temporal and spatial overheads of the system and enable it to ensure efficient and stable inter-core communication has become the key to implementing such an inter-core interaction mechanism. Therefore, the technical solution of this application is based on a heterogeneous fusion multi-core processor architecture based on a data-intensive coprocessor, and proposes a set of main control processors and an inter-core data processing method that uses a cache adaptation unit to poll and arbitrate status codes, and further proposes the system status codes of the data-intensive coprocessor heterogeneous multi-core processor.
[0058] In the first aspect, as Figure 1As shown in the figure, the technical solution of this application provides a data processing system based on heterogeneous integrated multi-core. This system can take into account the advantages of both the main control processor and the data-intensive coprocessor at the same time. It utilizes the configurable ability of the main control processor, uniformly manages the storage space allocation through a shared cache, and efficiently schedules the parallel operation of multiple coprocessors through a cache adaptation unit. The system has a significant acceleration ability for neural network algorithms and can also play the advantage of high configurability.
[0059] Referring to Figure 1 , in the embodiment, the data processing system based on heterogeneous integrated multi-core mainly includes: a main control processor, several cache adaptation units, a shared storage unit, and several intelligent coprocessors. Among them, the main control processor is mainly used to obtain the data to be operated and the data weights, generate activation data according to the data to be operated and the data weights, and generate the instruction data of the activation data. The cache adaptation unit is mainly used to write the activation data and the instruction data into the shared storage unit, modify the status code of the shared storage unit; and obtain the status code of the shared storage unit through polling and arbitrate the status of the shared storage unit; it is also used to write the target data obtained by the arithmetic processing into the shared storage unit, and determine that the arithmetic is completed according to the status code and output the target data. The shared storage unit is mainly used to store the activation data, the instruction data, and the target data. The intelligent coprocessor is mainly used to read the activation data and the instruction data and obtain the target processing through parallel arithmetic processing.
[0060] Specifically, during implementation, a main control processor, such as a CPU, etc., its functions include but are not limited to preprocessing weight data, generating activation data for the intelligent co-processor to perform operations, generating intelligent co-processor operation instructions for data-intensive operations, managing the shared cache address space, and sending the address where the data is located to the adaptation unit. The intelligent co-processor for data-intensive operations, such as GPU, TPU, NPU, etc., its functions include but are not limited to processing data-intensive operations, such as weight data loading, activation data inflow, performing operations, partial sum accumulation, data pooling and activation, data write-back, etc. Multiple intelligent co-processors for data-intensive operations can work in parallel. The cache adaptation unit polls the system status code and receives the data address, schedules different activation data and weight values to be written into different intelligent co-processors, and controls the data from the shared cache unit to be shared by each intelligent co-processor for data-intensive operations. The functions of the cache adaptation unit include but are not limited to: reading the system status code in the inter-core communication partition of the shared cache and performing identification arbitration on it, and then controlling the subordinate intelligent co-processor; reading data from the shared cache and writing it into the internal secondary cache of the intelligent co-processor; writing the operation result back to the shared cache space; modifying the corresponding status code in the shared cache space. Each data-intensive intelligent co-processor corresponds to an adaptation unit. The shared storage unit (space) in the embodiment can adopt a shared DRAM storage structure; the shared DRAM structure is used to achieve the consistency of the main control processor and the intelligent co-processor for data-intensive operations at the data storage level. Among them, the shared DRAM storage is divided into a neural activation data partition, a convolution kernel weight data partition, an operation result partition, each inter-core communication partition, and an operation instruction partition.
[0061] In the embodiment, the main control processor manages the cache space, and the cache adaptation unit is used to isolate the shared cache data address from the internal address of the intelligent co-processor it manages. This not only ensures the unity of the shared cache space, has the advantages of good generality, portability, and easy programming, but also enables the intelligent co-processor for data-intensive operations to perform operation tasks more specifically, reduces redundant operations, increases the proportion of operations during its execution, and enables it to operate efficiently.
[0062] Furthermore, the introduction of the cache adaptation unit also brings the potential for instruction reuse of multiple intelligent co-processors. In the embodiment, the main control processor generates instructions for multiple intelligent co-processors to control the execution process of the intelligent co-processors. And due to the address conversion mechanism of the adaptation unit, the intelligent co-processor has the independence of the internal storage system, making the instructions of the intelligent co-processor highly reusable. This approach greatly saves the storage resources of the co-processor operation instructions, reduces the total instruction length by multiples, effectively reduces the data scale processed by the heterogeneous integrated multi-core processor, thereby reducing the on-chip storage overhead and making the management of the shared cache space more efficient and simple.
[0063] In some feasible embodiments, the co-processors in the system are in one-to-one correspondence with the cache adaptation units. The intelligent co-processor includes an activation data storage area and a control instruction storage area; wherein, the activation data storage area is used to cache the activation data, and the control instruction storage area is used to cache the instruction data.
[0064] Specifically, in the implementation process, each cache adaptation unit manages an intelligent co-processor for data-intensive applications. The internal secondary cache system of the intelligent co-processor for data-intensive applications and the shared cache of the entire system use different address spaces, and data communication and interaction between the two are realized through the cache adaptation unit. Since different address spaces are used inside and outside the intelligent co-processor for data-intensive applications, the intelligent co-processor operation instructions generated by the main control processor can be reused, thereby reducing the on-chip storage overhead. Further, since each cache adaptation unit individually manages an intelligent co-processor for data-intensive applications and different address spaces are used inside and outside the intelligent co-processor for data-intensive applications, multiple intelligent co-processors for data-intensive applications can operate in parallel.
[0065] In some feasible embodiments, the cache adaptation units in the system include a first-in-first-out memory, a status code register, a status arbiter, and a control unit. Among them, the status code register is mainly used to obtain and temporarily store the status code of the shared storage unit. The status arbiter is mainly used to arbitrate the status of the shared storage unit. The first-in-first-out memory is mainly used to write the activation data and the instruction data into the shared storage unit; and is also used to write the target data obtained by arithmetic processing into the shared storage unit; output the target data; the control unit is used to obtain the source address and the number of bytes from the main control processor and perform data transfer.
[0066] Specifically, in the implementation process, as Figure 2 shown, the cache adaptation unit is composed of a first-in-first-out memory, a status code register, a status code arbiter, and a control unit, and is connected to the main control processor through an on-chip bus. A main control processor can have control paths for multiple cache adaptation units, and each cache adaptation unit controls an intelligent co-processor for data-intensive applications. Its output data path is connected to the activation data storage area, weight data storage area, and control instruction storage area of the intelligent co-processor for data-intensive applications.
[0067] Further, during the operation of the cache adaptation unit, after obtaining the start signal input from the bus, the status code is read from the cache, and this status code is sent to the status arbiter for arbitration. When the status code meets the conditions, the status arbiter drives the logic controller, and the control system obtains the source address and the number of bytes from the main control processor through the on-chip bus to perform data transfer. The multi-level first-in-first-out memory of the cache adaptation unit functions as a buffer between the data-intensive intelligent co-processor and the shared cache. More specifically, there is a group of first-in-first-out memories between the source and destination ends of the transmission. When resources are tense and data transfer cannot be completed, the first-in-first-out memory can provide a data staging area, thereby increasing data throughput and improving performance.
[0068] The execution process of the cache adaptation unit is as Figure 3 shown, and its functions include but are not limited to reading and arbitrating the system status code, reading and loading the intelligent co-processor operation instructions, reading and loading the weight data, reading and loading the activation data, triggering the data-intensive intelligent co-processor, writing back data, and reading and rewriting the system status code.
[0069] In some feasible embodiments, the control unit may include a source address staging unit, a logic controller, a byte counter, and a source address register.
[0070] Specifically in the embodiment, during the data transfer process, for each byte of data transferred, the source address register is incremented by 1 and the value of the byte counter is decremented by 1. When the number of bytes is 0, a new source address and the number of bytes are obtained from the on-chip bus to prepare for the next data transfer. Since each cache adaptation unit is only connected to one intelligent co-processor, there is no need to obtain the destination address through the main control processor when writing data. However, since a write-back address is required when writing data back to the shared cache, this address is temporarily stored in the write-back address staging unit. When writing back data, this address flows into the destination address register, and the write-back is controlled by the logic controller.
[0071] This mechanism can be directly initialized and assigned by the control unit, thereby achieving address isolation from the system shared cache, reducing the complexity of system address management. At the same time, since only one data-intensive intelligent co-processor needs to be managed, that is, through the byte counter, the address management becomes simple and efficient, and the degree of reusability is improved.
[0072] In some feasible embodiments, the address allocation of the shared storage space is as Figure 4As shown in the figure, where Region 1 is used to store the preprocessed and rearranged tensor data, Region 2 is used to store the convolutional kernel weight data, Region 3 is used to store the returned feature map data, Region 4 is used for inter-core communication data, and Region 5 is used to store operation instructions. Among them, the system status code is stored in the inter-core communication partition 4 and can be read by the main control processor and all cache adaptation units. After receiving the start signal sent by the main control processor, the cache adaptation unit polls the status code in Region 4 and arbitrates it. If the detected status code is the corresponding value, the cache adaptation unit starts to work. When the status arbiter recognizes the corresponding system status, the cache adaptation unit reads the corresponding operation instructions from Region 5 of the shared cache and transfers them to the control unit of the data-intensive intelligent co-processor; reads the corresponding weight data from Region 2 of the shared cache and stores it in the weight storage area of the data-intensive intelligent co-processor; reads the corresponding activation data from Region 1 of the shared cache and stores it in the activation data storage area of the data-intensive intelligent co-processor. When the data-intensive intelligent co-processor finishes the operation, it sends a signal to the cache adaptation unit. The cache adaptation unit writes the operation result from the central cache area of the data-intensive intelligent co-processor into the shared cache and changes the value of the corresponding bit of the system status code in its Region 4. When the system status code indicates that all data-intensive intelligent co-processors have completed the operation, the main control processor reads the operation result from the shared cache, organizes it into output data, outputs the operation result externally, and changes the status code.
[0073] On the other hand, for the data processing system based on heterogeneous fusion multi-cores provided in the technical solution of the present application, a data processing method based on the system and each module unit in the first aspect is further provided. The method includes steps S01 - S05:
[0074] S01. Obtain the data to be operated and the data weight, generate activation data according to the data to be operated and the data weight, and generate instruction data for the activation data.
[0075] In the embodiment, the main control processor reads the data to be operated and the weight from the external memory, generates activation data that can be operated by the data-intensive intelligent co-processor, such as neural tensor data; and simultaneously generates instruction data for the data-intensive intelligent co-processor.
[0076] S02. Write the activation data and the instruction data into the shared cache, and change the status code of the shared cache after the data writing is completed.
[0077] In the embodiment, the activation data and instruction data generated in step S01 are written into the specified position in the shared cache space, and the status code is changed to trigger the corresponding intelligent co-processor.
[0078] More specifically, in the embodiment, the main control processor first preprocesses the weight data so that it can meet the data format of the data-intensive intelligent co-processor. After the preprocessing of the weight data is completed, the main control processor generates operation instructions for the intelligent co-processor. Since the internal cache of each intelligent co-processor is isolated by the cache adaptation unit, different data can be mapped to the same address of different intelligent co-processors. Usually, in a given algorithm, different data may perform the same operations, such as convolution, activation, etc. Especially when multiple co-processors cooperate to process the same-sized data segments, the operations they adopt are exactly the same. Therefore, when they independently use their own storage systems, they can be controlled by the same set of instructions. Since multiple data-intensive intelligent co-processors can use the same operation instructions, this approach greatly saves the storage resources of the operation instructions of the intelligent co-processor, reduces the total instruction length by multiples, effectively reduces the data scale processed by the multi-core processor, thereby reducing the on-chip storage overhead, and making the management of the shared cache space more efficient and simple. After the relevant data preprocessing and the generation of the operation instructions of the intelligent co-processor are completed, the main control processor writes the data into the shared cache, including the status code of the system.
[0079] S03. Obtain the changed status code by polling, and arbitrate the status of the shared cache according to the status code.
[0080] Specifically in the embodiment, the cache adaptation unit polls and checks the corresponding status code. If the condition is met, the adapter of the corresponding data-intensive intelligent co-processor starts to work.
[0081] More specifically, after the cache adapter in the embodiment obtains the start signal input from the bus, it reads the status code from the cache, and the status code is sent to the status arbiter for arbitration. When the status code meets the condition, the status arbiter drives the logic controller, and the control system obtains the source address and the number of bytes from the main control processor through the on-chip bus for data transfer.
[0082] S04. Output the activation data to the first intelligent co-processor according to the result of the status arbitration, and perform arithmetic processing according to the instruction data; or, determine that the first intelligent co-processor is in the arithmetic state, and output the activation data to the second intelligent co-processor according to the result of the status judgment, and perform arithmetic processing according to the instruction data.
[0083] Specifically in the embodiment, the cache adaptation unit polls and checks the corresponding status code. If the condition is met, the adapter of the corresponding data-intensive intelligent co-processor starts to work, rewrites the corresponding status code, and reads data from the address space accessible to them in the shared cache into the data-intensive intelligent co-processor. After the data reading is completed, the corresponding status code is rewritten. When a certain intelligent co-processor performs an operation, since the shared cache is not read or written in this state, a certain cache adaptation unit other than this can perform this step, thus realizing multi-core parallel operation.
[0084] S05. Write the target data obtained by the arithmetic processing into the shared cache, determine that the operation is completed according to the status code, and output the target data.
[0085] Specifically in the embodiment, when a certain intelligent co-processor completes an operation, the cache adaptation unit queries whether the shared cache is readable, then writes the data to the specified position in the shared cache, and rewrites the corresponding position of the status code. The main control processor checks the corresponding status code. If the condition that all intelligent co-processors have completed the operation is met, the main control processor rearranges and outputs the final data.
[0086] In some feasible implementation manners, the embodiment proposes a status code based on the heterogeneous fusion multi-core processor architecture of the main control processor and the data-intensive co-processor; for a heterogeneous multi-core processor composed of one main control processor and n data-intensive intelligent co-processors, an n + 3-bit binary number management system state is adopted.
[0087] Specifically in the embodiment, each position of the status code is defined as follows: bit 0 is the MA (memory_available) signal, that is, the read availability flag bit. When it is 1, it indicates that the shared cache is in the read state, and when it is 0, it indicates that the shared cache is in the readable state; bit 1 is RW (read / write), that is, the read / write status flag bit. When this bit is 0, it indicates the read state, and when this bit is 1, it indicates the write state; bit 2 is CW (CPU_working), that is, the processor status flag bit. When this bit is 1, it indicates that the main control processor module is preparing data and instructions in the read / write shared storage, and when this bit is 0, it indicates that the main control processor module has completed its work; and bit n + 2 is NW (NPU_working), that is, the cache adaptation status flag bit. When this bit is 1, it indicates that the cache adaptation unit n indicates that the module is reading and writing data and instructions in the shared cache and driving the lower-level tensor processor to work, and when it is 0, it indicates that the data-intensive intelligent co-processor has completed the operation and written back the result. The status code in the embodiment can be integrated with the shared cache architecture. It can be achieved by designing the status code space at the corresponding position in the inter-core communication partition of the shared cache, without the need to design specific hardware modules, which is convenient for configuration on various heterogeneous multi-core processor systems. Further, this status code is a global status code, which can be used for external query of the system status, facilitating management and system expansion or maintenance. In addition, the status code system in the embodiment has clear logic and strong scalability. When the number of intelligent co-processors is expanded, the status code logic remains unchanged, and only the corresponding NW bits need to be added to be compatible, which is easy to implement. Further, each intelligent co-processor judges its behavior through the system status code. The mode of polling and arbitrating the status code by the cache adaptation unit will not send an interrupt program to the main control processor, greatly liberating the main control processor and reducing the time and space overhead brought by inter-core interaction to the main control processor.
[0088] In some feasible embodiments, based on the foregoing definition of the status code, the process of data interaction in each unit module of the system in the data processing method of the heterogeneous fusion multi-core in the embodiment may include the following steps:
[0089] Step 1) When the main control processor is working, set MA, RW, and CW to 1. After its work is completed, set MA, RW, and CW to 0, and send a start signal to each cache adaptation unit through the on-chip bus.
[0090] Step 2) When a certain cache adaptation unit needs to move the shared cache data to the intelligent co-processor, set the corresponding NW bit to 1, RW to 0, and MA to 1. After this data transfer is completed, set MA to 0.
[0091] Step 3) After a certain intelligent coprocessor finishes its operation, check whether MA is 0. If it is 0, set the RW signal to 1, set MA to 1, write the data back to the corresponding address, then set MA to 0, and set the corresponding NW bit to 0.
[0092] Step 4) When the main control processor checks that all NW bits of the status code are 0, RW is 1, and MA is 0, it indicates that all intelligent coprocessors have finished their operations and written back the data. At this time, the main control processor rewrites CW to 1, reads the data, organizes it, and outputs it. After finishing the work, the main control processor rewrites RW and CW to 0.
[0093] Specifically, in the embodiment, as Figure 5 shown, taking the seven - bit status code to control four cache adaption units as an example, the upper four bits (bits 6, 5, 4, 3) represent the working status of the four cache adaption units, the 2nd bit represents the working status of the main control processor, and the lowest two bits represent the working status of the shared memory.
[0094] When the main control processor is working, the status code is 2’b0000111. After it finishes working, it sets the status code to 2b’0000000 and sends a start signal to the cache adaption units on the on - chip bus. Then when a certain cache adaption unit starts to work, it sets the corresponding NW bit to 1, RW to 0, and MA to 1. For example, the 1st cache adaption unit sets the status code to 2’b000101 and starts to work. After finishing data transfer, it becomes 2’b000100. Similarly, then the 2nd cache adaption unit sets the status code to 2’b001101 and starts to work. After finishing data transfer, it becomes 2’b001100... After a certain cache adaption unit finishes its operation, check whether MA is 0. If it is 0, set the RW signal to 1, set MA to 1, write the data back to the corresponding address, then set MA to 0, and set the corresponding NW bit to 0. It should be noted that among the multiple NW signals, each cache adaption unit is only sensitive to its own corresponding NW bit and has nothing to do with the status of other bits.
[0095] When the main control processor checks that the status code is 2b’0000010, it indicates that all data - intensive intelligent coprocessors have finished their operations and written back the data. At this time, the main control processor rewrites the status code to 2’b0000101, reads the data, organizes it, and outputs it. After finishing the work, the main control processor rewrites it to 2’b0000000.
[0096] Furthermore, in an embodiment of the present invention, in the case of multiple data - intensive intelligent coprocessors running in parallel, the operation flow of the heterogeneous fusion multi - core processor is as Figure 6As shown below: First, the main control processor preprocesses the data, generates intelligent co-processor control instructions, and then transfers the data to the corresponding position in the unified buffer. Then, the main control processor sends a start signal to the cache adaptation unit through the on-chip bus. Each cache adaptation unit works, polls the system status code and performs arbitration, transfers the corresponding data to the corresponding data-intensive intelligent co-processor, and rewrites the corresponding status code. At the same time, the data-intensive intelligent co-processor is started by the corresponding upper-level cache adaptation unit to perform the corresponding operation. Further, when a certain intelligent co-processor finishes the operation, the cache adaptation unit writes the operation result back to the corresponding position in the shared cache and changes the corresponding status code. When all bit status codes are in the state of completing the operation, the main control processor reads the data from the shared cache, organizes it into output data, outputs the operation result externally, and initializes the status code.
[0097] In summary, the data processing method and processing system based on heterogeneous fusion multi-core proposed by the technical solution of this application have at least the following advantages compared with the related technical solutions:
[0098] (1) Compared with the traditional heterogeneous multi-core processor interaction mechanism, the multi-core processor interaction mechanism of the present invention can greatly liberate the main control processor. The main control processor only needs to send the corresponding start signal and address through the on-chip bus. Each intelligent co-processor judges the behavior through the system status code, and the polling status code mode does not send an interrupt program to the main control processor, greatly liberating the main control processor and reducing the temporal and spatial overhead brought by inter-core interaction to the main control processor.
[0099] (2) Compared with the traditional heterogeneous multi-core processor interaction mechanism, the present invention has high portability. This interaction mechanism can be integrated with the shared cache architecture, which can be achieved by designing the system status code space at the corresponding position in the common area of the shared cache, without designing specific hardware modules. It is convenient to configure on various heterogeneous multi-core processors.
[0100] (3) Compared with the traditional heterogeneous multi-core processor interaction mechanism, this interaction mechanism relies on the global status code, and the external can query the system status through the system status code. It is convenient for the external to supervise and manage the task process of this processor system, so as to carry out expansion or maintenance.
[0101] (4) Compared with the traditional heterogeneous multi-core processor interaction mechanism, the multi-core processor interaction mechanism of the present invention has high scalability. The status code system of the present invention has clear logic. In practical applications, considering the case of expanding the number of intelligent co-processors, the logic of the inter-core interaction mechanism based on the system status code is not affected by the number of intelligent co-processors. Only by adding the corresponding NW bits can it be compatible with the original status code system. In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More precisely, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0102] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with such instruction execution systems, apparatus, or devices.
[0103] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0104] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention. The scope of the present invention is defined by the claims and their equivalents.
[0105] The above is a specific description of the preferred embodiment of the present invention. However, the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A data processing method based on heterogeneous fusion multi - cores, characterized in that, It includes the following steps: Obtain the data to be operated and the data weight, generate activation data according to the data to be operated and the data weight, and generate instruction data for the activation data; Write the activation data and the instruction data into a shared cache, and change the status code of the shared cache after the data writing is completed; Obtain the changed status code by polling, and arbitrate the status of the shared cache according to the status code; According to the result of status arbitration, output the activation data to the first intelligent co-processor, and perform arithmetic processing according to the instruction data; or, determine that the first intelligent co-processor is in an arithmetic state, and according to the result of the status judgment, output the activation data to the second intelligent co-processor, and perform arithmetic processing according to the instruction data; Write the target data obtained by the arithmetic processing into the shared cache, determine that the arithmetic operation is completed according to the status code, and output the target data; The status code includes a read availability flag bit, a read / write status flag bit, a processor status flag bit, and a cache adaptation status flag bit; The step of outputting the activation data to the first intelligent co-processor according to the result of the status judgment, and performing arithmetic processing according to the instruction data; or, determining that the first intelligent co-processor is in an arithmetic state, and outputting the activation data to the second intelligent co-processor according to the result of the status judgment, and performing arithmetic processing according to the instruction data includes the following steps: Determine that the read availability flag bit, the read / write status flag bit, and the processor status flag bit are all 0, and trigger a data transfer signal; Determine that the read availability flag bit and the read / write status flag bit are both 0 and the cache adaptation status flag bit is 1. According to the data transfer signal, output the activation data to the first intelligent co-processor or the second intelligent co-processor for arithmetic processing, and set the read availability flag bit to 0; The step of writing the target data obtained by the arithmetic processing into the shared cache, determining that the arithmetic operation is completed according to the status code, and outputting the target data includes: Determine that the read availability flag bit and the read / write status flag bit are both 1, write the target data obtained by the arithmetic processing into the shared cache, and set both the read availability flag bit and the cache adaptation status flag bit to 0; Determine that the read / write status flag bit is 1, the read availability flag bit and the cache adaptation status flag bit are both 0, output the target data, and set both the read / write status flag bit and the processor status flag bit to 0.
2. The data processing method based on heterogeneous fusion multi - cores according to claim 1, wherein The step of obtaining the changed status code by polling, and arbitrating the status of the shared cache according to the status code includes: Read the status code according to the polling start signal; Arbitrate the obtained status code, and obtain the source address and byte count of the activation data from the shared cache according to the arbitration result; Read the activation data into the first intelligent co-processor or the second intelligent co-processor according to the source address and the byte count.
3. The data processing method based on heterogeneous fusion multi - cores according to claim 2, wherein Writing the target data obtained through arithmetic processing to the shared cache, determining that the arithmetic operation is completed according to the status code, and outputting the target data includes: Determining the write-back address of the target data and temporarily storing the write-back address; Causing the temporarily stored write-back address to flow into the destination address register, and writing the target data back to the shared cache through the destination address register.
4. A data processing system based on heterogeneous fusion multi - cores for implementing the data processing method based on heterogeneous fusion multi - cores as described in any one of claims 1 - 3, characterized in that, Including: A main control processor; For obtaining data to be arithmetically processed and data weights, generating activation data according to the data to be arithmetically processed and the data weights, and generating instruction data for the activation data; A cache adaptation unit; For writing the activation data and the instruction data to a shared storage unit and modifying the status code of the shared storage unit; And polling to obtain the status code of the shared storage unit and arbitrating the status of the shared storage unit; also for writing the target data obtained through arithmetic processing to the shared storage unit, and determining that the arithmetic operation is completed according to the status code and outputting the target data; A shared storage unit; For storing the activation data, the instruction data, and the target data; A plurality of intelligent coprocessors; for reading the activation data and the instruction data and obtaining target processing through parallel arithmetic processing.
5. The data processing system based on heterogeneous fusion multi-core according to claim 4, characterized in that The intelligent coprocessors are in one-to-one correspondence with the cache adaptation unit; the intelligent coprocessor includes an activation data storage area and a control instruction storage area; The activation data storage area is used for caching the activation data, and the control instruction storage area is used for caching the instruction data.
6. The data processing system based on heterogeneous fusion multi - cores according to claim 4, wherein, The cache adaptation unit includes a first-in-first-out memory, a status code register, a status arbiter, and a control unit; The control unit is connected to the main control processor through an on-chip bus; the status code register is connected to the shared storage unit in a cache-direct connection manner; the status code register is also connected to the status arbiter, the status arbiter is connected to the control unit, and the data path of the first-in-first-out memory is connected to the intelligent coprocessor; The status code register is used for obtaining and temporarily storing the status code of the shared storage unit; The status arbiter is used for arbitrating the status of the shared storage unit; The first-in-first-out memory is used for writing the activation data and the instruction data to the shared storage unit; Also for writing the target data obtained through arithmetic processing to the shared storage unit; outputting the target data; The control unit is used for obtaining a source address and a number of bytes from the main control processor and performing data transfer.
7. The data processing system based on heterogeneous fusion multi-core according to claim 6, wherein The control unit includes: A source address temporary storage unit for obtaining the source address and the number of bytes in the shared storage unit before data transfer; A logic controller for controlling the byte counter and the source address register to perform logical operations; A byte counter for performing arithmetic operations on the number of bytes; A source address register for performing arithmetic operations on the source address.
8. The data processing system based on heterogeneous fusion multi - cores according to claim 6, wherein, The control unit further includes a destination address register; The destination address register is used for determining the write-back address of the target data, temporarily storing the write-back address, and writing the target data back to the cache adaptation unit according to the write-back address.
Citation Information
Patent Citations
Content delivery network
CN104011701A
Method and apparatus for multi-key total memory encryption based on dynamic key derivation
CN113051626A