Logic design system for controller hardware acceleration

By designing a logic design system with a hardware acceleration of controller, using customized circuits and dedicated logic circuits for parallel computing, the performance bottleneck problem of traditional database systems when processing large-scale data sets is solved, and efficient database computing operations are achieved.

CN119806644BActive Publication Date: 2025-05-16INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510307776.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-05-16
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

Traditional CPU-based database systems face performance bottlenecks when processing large-scale data sets. Although multi-core CPU technology has certain mitigation effects, since the CPU design is oriented towards general computing, its architectural complexity leads to a limited number of CPU cores that can be created, and parallel computing improvement results are not ideal.

Method used

Design a logic design system for hardware acceleration of controllers, using customized special circuits and a large number of special logic circuits to form an acceleration core, and perform parallel calculations through target column extractor, length and width parallel comparator, matrix multiplier and truth table finder to improve the throughput rate of database calculation and reduce latency.

Benefits of technology

It effectively improves the throughput rate of database computing and reduces the latency rate, improves the speed of database operations through parallel computing, and solves the performance bottleneck problem of traditional database systems when processing large-scale data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806644B_ABST
    Figure CN119806644B_ABST
Patent Text Reader

Abstract

The present application discloses a logic design system for controller hardware acceleration, which relates to the field of controller technology, including a target column extractor, a long bit width parallel comparator, a matrix multiplier and a truth table finder; the target column extractor is used to extract target column data information from a preset page, and send the target column data information to the long bit width parallel comparator; since the target column extractor does not need to read the column data information one by one, but directly reads the target column data information, the reading speed can be accelerated. Further, the long bit width parallel comparator is used to copy the target column data information to the corresponding long bit width register according to the data type, and obtain the comparison result according to the comparison operation instruction code and the target column data information, and send the comparison result to the matrix multiplier; since the long bit width parallel comparator can compare the target column data information in parallel, the throughput of database calculation can be effectively improved and the delay rate can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of controller technology, and in particular to a logic design system for controller hardware acceleration. Background Art

[0002] Database operations, such as page parsing, query filtering, data sorting, aggregation and merging, often require a lot of computing resources and time. Traditional CPU-based database systems can only use serial computing to execute computing tasks in the database one by one.

[0003] However, as the amount of data continues to grow, performance bottlenecks are faced when processing large-scale data sets. Although multi-core CPU technology has alleviated the performance bottleneck of database computing to a certain extent, the design of the CPU is oriented towards general computing, and the complexity of its architectural design requires a large amount of logical resources, resulting in a limited number of CPU cores that can be created, and the effect of parallel computing improvement is still not ideal.

[0004] Therefore, there is an urgent need for a controller hardware accelerated logic design system that can increase the speed of database computing operations, use customized dedicated circuits for computationally intensive operations in the data, and perform parallel computing by forming acceleration cores through a large number of dedicated logic circuits, which can effectively improve the throughput of database computing and reduce the latency rate. Summary of the invention

[0005] The present application provides a logic design system for controller hardware acceleration to increase the speed of database computing operations. Customized dedicated circuits are used for computationally intensive operations in the data. A large number of dedicated logic circuits are used to form acceleration cores for parallel computing, which can effectively improve the throughput of database computing and reduce latency.

[0006] The present application provides a controller hardware accelerated logic design system, the controller hardware accelerated logic design system communicates with a host end; the controller hardware accelerated logic design system comprises: a target column extractor, a long bit width parallel comparator, a matrix multiplier and a truth table finder;

[0007] The target column extractor is used to extract target column data information from a preset page, and send the target column data information to the long bit width parallel comparator; wherein the storage structure of the preset page is that the first preset position stores page meta information, the second preset position stores row meta information, the third preset position stores variable length column length information, and the fourth preset position stores row data;

[0008] The long bit width parallel comparator is used to copy the target column data information to the corresponding long bit width register according to the data type, and obtain a comparison result according to the comparison operation instruction code and the target column data information, and send the comparison result to the matrix multiplier; wherein the comparison operation instruction code is sent from the host end to the logic design system of the controller hardware acceleration;

[0009] The matrix multiplier is used to determine the search code according to the comparison result and the permutation matrix; wherein the permutation matrix is ​​sent by the host end to the logic design system of the controller hardware acceleration;

[0010] The truth table finder is used to output a matching result according to the lookup code and the logic operation truth table; wherein the matching result is used to be sent to the host end.

[0011] The present application also provides a controller hardware acceleration method, which is applied to a logic design system of the controller hardware acceleration, wherein the logic design system of the controller hardware acceleration communicates with a host end; the logic design system of the controller hardware acceleration comprises: a target column extractor, a long bit width parallel comparator, a matrix multiplier and a truth table finder;

[0012] The target column extractor extracts target column data information from a preset page, and sends the target column data information to the long bit width parallel comparator; wherein the storage structure of the preset page is that the first preset position stores page meta information, the second preset position stores row meta information, the third preset position stores variable length column length information, and the fourth preset position stores row data;

[0013] The target column data information is copied to the corresponding long bit width register according to the data type through the long bit width parallel comparator, and a comparison result is obtained according to the comparison operation instruction code and the target column data information, and the comparison result is sent to the matrix multiplier; wherein the comparison operation instruction code is sent from the host end to the logic design system of the controller hardware acceleration;

[0014] Determine the search code by the matrix multiplier according to the comparison result and the permutation matrix; wherein the permutation matrix is ​​sent by the host end to the logic design system of the controller hardware acceleration;

[0015] The truth table finder outputs a matching result according to the lookup code and the logic operation truth table; wherein the matching result is used to be sent to the host end.

[0016] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the controller hardware acceleration method when executing the computer program.

[0017] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the controller hardware acceleration method are implemented.

[0018] The present application also provides a computer program product, including a computer program, which implements the steps of the controller hardware acceleration method when executed by a processor.

[0019] Through this application, since the logic design system of the controller hardware acceleration communicates with the host side; the logic design system of the controller hardware acceleration includes: a target column extractor, a long bit width parallel comparator, a matrix multiplier and a truth table finder; wherein the target column extractor is used to extract the target column data information of the preset page, and send the target column data information to the long bit width parallel comparator; since the target column extractor does not need to read the column data information one by one, but directly reads the target column data information, the reading speed can be accelerated. Further, the long bit width parallel comparator is used to copy the target column data information to the corresponding long bit width register according to the data type, and obtain the comparison result according to the comparison operation instruction code and the target column data information, and send the comparison result to the matrix multiplier; since the long bit width parallel comparator can compare the target column data information in parallel, the execution speed can also be improved. The matrix multiplier is used to determine the search code based on the comparison result and the permutation matrix; the truth table finder is used to output the matching result based on the search code and the logic operation truth table. The final result can be obtained directly through the logic operation truth table without reading the intermediate comparison result and then calculating the final value, thereby effectively improving the throughput of database calculations and reducing the delay rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 A schematic diagram of a PostgreSQL storage format provided in an embodiment of the present application;

[0022] Figure 2 A schematic diagram of an improved storage engine for FPGA provided in an embodiment of the present application;

[0023] Figure 3A schematic diagram of the structure of a logic design system for controller hardware acceleration provided in an embodiment of the present application;

[0024] Figure 4 A schematic diagram of a storage structure of a preset page provided in an embodiment of the present application;

[0025] Figure 5 A schematic diagram of the structure of a target column extractor provided in an embodiment of the present application;

[0026] Figure 6 A schematic diagram of a storage format of a preset page provided in an embodiment of the present application;

[0027] Figure 7 A schematic diagram of an offset value calculation process provided in an embodiment of the present application;

[0028] Figure 8 A schematic diagram of register numbering provided in an embodiment of the present application;

[0029] Fig. 9 A schematic diagram of a comparison operation process provided in an embodiment of the present application;

[0030] Fig.10 A schematic diagram of the structure of a long bit width parallel comparator provided in an embodiment of the present application;

[0031] Fig.11 A schematic diagram of a matrix parallel computing implementation process provided in an embodiment of the present application;

[0032] Fig.12 A schematic diagram of a calculation process of a matrix multiplier provided in an embodiment of the present application;

[0033] Fig.13 A schematic diagram of a calculation process of a truth table finder provided in an embodiment of the present application;

[0034] Fig.14 A schematic diagram of a process of arithmetic operation of a logic design system for controller hardware acceleration provided in an embodiment of the present application;

[0035] Fig.15 A flowchart of a controller hardware acceleration method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0037] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0038] Mainstream relational databases such as MySQL, PostgreSQL, and Oracle use fixed-size pages to store data in rows. Figure 1 A schematic diagram of a PostgreSQL storage format is shown. PostgreSQL uses a fixed 8KB size as a page to store row data entered by the user. It has the following features: 1) The number of row data stored in the page is marked at a fixed position in the page. 2) Starting from the fixed position of the page and moving toward the bottom of the page, the metadata of the row data is stored, and the metadata includes the offset and length of the row. The metadata of the row data usually refers to the descriptive information of the row data, which is used to understand the structure, characteristics and context of the row data. In the database, the metadata of the row data may include the following: Table structure information: the name of the table to which the row data belongs, the storage engine of the table, the character set, etc. Column information: the name, data type, whether it is allowed to be empty, the default value, the character length, etc. of each field in the row data. Index information: the index information of the table to which the row data belongs, including the index name, index type (such as BTREE, HASH, etc.), index column, etc. Constraint information: the constraint information of the table to which the row data belongs, such as primary key constraint, unique constraint, foreign key constraint, etc. Data source and purpose: the source and purpose of the row data and its relationship with other data. Data quality information: quality indicators of row data, such as data accuracy, completeness, consistency, etc. Data update information: creation time and last update time of row data, etc. 3) Starting from the bottom of the page and moving toward the top of the page, row data is stored, and each row data contains several column fields. The length of the column field is stored in the column field attribute table of the database. For fixed-length column fields, the value in the field attribute table is greater than 0; for variable-length column fields, the value in the field attribute table is less than 0, and the actual value is stored in the header of the field, using 1B or 4B to store the field length information.

[0039] Compared with the serial computing of CPU, the core advantage of using FPGA for heterogeneous acceleration of database is that FPGA can flexibly customize the bit width of access data, use customized hardware circuits for specific computing tasks, and realize hardware acceleration through multi-parallel methods. For example, the common 64-bit CPU currently has a data access bit width of 64 bits (8B) at a time, and the number of cores under a node usually does not exceed 64. Using FPGA, you can use a 512-bit bit width to access 64B of data at a time, and at the same time, you can instantiate more computing units than the number of CPU cores for parallel computing. The basic process of FPGA for heterogeneous database acceleration is: first, load the data pages from the disk to the DDR memory of the FPGA in batches, and the computing unit in the FPGA loads several 8KB pages from the DDR at a time to the memory inside the FPGA for parsing and filtering calculations. Among them, the memory inside the FPGA can be on-chip RAM.

[0040] Before filtering calculation, it is necessary to parse the 8KB page and extract the column fields required in the SQL statement. Parallel acceleration is achieved by creating a large number of parsing and filtering calculation units.

[0041] In the related art, the extraction of column fields in the page is mostly done by parsing byte by byte. First, the page row meta information is read from the 8KB page to obtain the offset and length of each row in the page, and each byte in the row is read byte by byte according to the column field attribute table. When the target column field in the SQL statement is encountered, the column field is copied to the comparator until all the target column fields are copied to the comparator, and then the next column is processed.

[0042] In the related art, there is also a storage engine improvement method for FPGA, which places the length information of the variable-length field at the end of the row meta information, and puts the fixed-length fields of the same type together, so that the FPGA can quickly calculate the position of the target column in the page, avoiding the byte-by-byte analysis of the column field, so as to improve the processing efficiency of the FPGA. For details, please refer to Figure 2 A schematic diagram of an improved storage engine for FPGA is shown.

[0043] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0044] Figure 3 1 is a schematic diagram of a logic design system for controller hardware acceleration provided by an embodiment of the present disclosure. The controller may be an FPGA. Figure 3As shown, the FPGA hardware accelerated logic design system 10 communicates with the host end 20; in this embodiment, taking the controller as an FPGA as an example, the FPGA hardware accelerated logic design system 10 provided in this embodiment includes: a target column extractor 301, a long bit width parallel comparator 302, a matrix multiplier 303 and a truth table finder 304;

[0045] The target column extractor 301 is used to extract the target column data information of the preset page and send the target column data information to the long bit width parallel comparator 302; wherein the storage structure of the preset page is that the first preset position stores page meta information, the second preset position stores row meta information, the third preset position stores variable length column length information and the fourth preset position stores row data.

[0046] In one example, see Figure 4 A schematic diagram of a storage structure of a preset page is shown. The storage structure of the preset page has the following technical features:

[0047] Feature 601: Set the maximum supported data bit width in the FPGA ( ), starting from the top of the page, a fixed position stores the page meta information of this page, including: number of rows and variable length column length. length.

[0048] Feature 602: Use The length stores the row meta information. Assume that the row meta information is of fixed length ( ), the number of row metadata that can be stored at one time is: .

[0049] Feature 603: Storage The variable length column length information of each row, each variable length column takes up 2B, and the total occupied space is calculated by Align it and fill the rest with 0.

[0050] Feature 604: Storage The row data contains several column fields, which store fixed-length fields first and variable-length fields later. The column fields are stored according to the register bit width length preset in the FPGA ( ) are aligned and stored, and the total length of the row data is based on Align and store, and fill the rest with 0.

[0051] Feature 605: Repeat features 602-604 until there is no free space in the page to store the next row of data.

[0052] The advantage of such a setting is that the improved storage format of this embodiment is used to facilitate the FPGA to use continuous burst memory access, read multiple continuous row data at a time, and improve the memory access efficiency of the DDR memory. Compared with the prior art, there is a gap between the meta information and the row data, resulting in the problem of inability to access the memory continuously. In the present invention, there is no gap between the meta information and the row data, and continuous memory access can be performed, thereby improving the memory access efficiency.

[0053] In one example, see Figure 5 A schematic diagram of the structure of a target column extractor is shown.

[0054] In one example, the target column extractor includes: a page read module and a target column copy module; the page read module includes a memory access controller, a routing state machine and a meta-information register group; the target column copy module includes: a column offset accumulator, a row cache register group and a register copier.

[0055] In one example, the memory access controller includes an on-chip RAM for temporarily storing data, and the memory access controller is used to read data of a preset page from a preset disk and store the data of the preset page in the on-chip RAM for temporarily storing data.

[0056] In one example, the longest bus data width supported by the FPGA is used to read data externally. The internal RAM contains an on-chip RAM for temporary data storage, the width of which is the longest bus width of the FPGA, and the depth is an integer multiple of the amount of data read in a bus burst. Based on the AXI bus, a 512-bit memory access interface is used.

[0057] In one example, the routing state machine is used to distribute data of a preset page stored in an on-chip RAM for temporarily storing data.

[0058] In one example, the routing state machine is used to distribute page meta information, row meta information and variable-length column length information to the meta information register group; the routing state machine is used to distribute row data to the row cache register group.

[0059] It is worth noting that, in the present invention, if the filter expression in the SQL statement does not involve a variable-length column field, the variable-length column length information is forwarded and skipped to save the required clock cycle.

[0060] In one example, the meta information register group is used to store the number of rows, row offset, and variable-length column length in the data of a preset page.

[0061] In one example, the column offset accumulator is used to parallely calculate the offset of each column relative to the starting position according to a preconfigured fixed-length column length list and a variable-length column length, and send the offset value of the target column field in the row to the register copier.

[0062] In one example, the row cache register group is used to split row data and store them into multiple registers that can be accessed in parallel.

[0063] In one example, a RAM cache line is split and stored in multiple registers that can be accessed in parallel. In the example: Register width 4B, the number of registers is the maximum row length supported in the page ;in, Used to indicate the maximum supported length of a data table row.

[0064] The advantage of such a setting is that the improved storage format of this embodiment is used to reduce the FPGA's cache capacity requirements and parsing delay for the page. In the related art, at least 8KB of storage space (PostgreSQL) is required, and the parsing calculation can be started only after all 8KB of data is cached. In this embodiment, the row meta information and row data are stored continuously according to the preset rules. In theory, there is no need to use RAM for data caching. In order to take into account the burst transmission efficiency, a RAM cache of 1-2KB can be used.

[0065] In one example, the register copier is used to copy corresponding register contents from the row cache register group to the long bit width parallel comparator in parallel according to the target column required in the comparison instruction.

[0066] In an example, for a clearer explanation, an example is given as follows:

[0067] Assume that the hardware and software configuration is as follows:

[0068] 1) Use the PostgreSQL database, the maximum supported length of the data table row length .

[0069] 2) Using PostgreSQL database, the length of its row metadata .

[0070] 2) Using PostgreSQL database, the length of common data types is: integer = 4B, floating point = 8B, date = 4B, and each Chinese character occupies 2B.

[0071] 3) The sample query table has 8 fields, including 6 fixed-length fields and 2 variable-length fields.

[0072] 4) FPGA hardware platform, using AXI bus for internal interconnection, with a maximum supported bus width =512bit.

[0073] 5) FPGA hardware platform, register width in row cache register group .

[0074] Assume that the user creates the following table 1:

[0075] Table 1: Table information created by the user

[0076]

[0077] In this embodiment, see Figure 6 A structural schematic diagram of a storage format of a preset page is shown.

[0078] In one example, the first row stores page meta information, including the number of variable-length columns and the number of rows in the page, and occupies a fixed 64B.

[0079] In one example, the meta information of each row is stored (starting from 64B), and the offset of each row in the page is marked in the meta information. Each row of meta information occupies 4B, and 64B / 4B=16 rows of data can be stored at a time.

[0080] In one example, the length information of the variable-length column is stored starting from 128B. In this example, there are 2 variable-length columns, each of which occupies 2B, so each row of variable-length columns occupies 4B, and a 64B bit width can store 16 rows of variable-length column meta information. If there are many variable-length columns, use multiple 64B bit width rows to store the meta information of the variable-length columns.

[0081] In an example, fixed-length columns are stored first in the row data, followed by variable-length columns. The starting position of each column field is guaranteed to be 4B aligned, and the rest is filled with 0.

[0082] The benefit of this setting is that it is convenient for the FPGA to use fewer clock cycles to extract the data of the target column. In the related art, the FPGA needs to parse the target column field position and extract the data byte by byte. Especially when variable-length fields exist, multiple clock cycles are required to perform data parsing and splicing operations. In this embodiment, after the row metadata, row data, and column field data are stored according to preset rules, the target field can be mapped to an aligned register operation, thereby greatly reducing the combinational logic operations required for column field extraction.

[0083] For the structural diagram of the storage format of the above-mentioned preset page, the working process of the column field extraction module of the FPGA in this embodiment is as follows:

[0084] 1) Use 512-bit bit width and burst transmission mode to batch obtain multiple 8KB memory pages and put them into the internal temporary RAM. The bit width of the temporary RAM is also 512 bits, and 512 bits of data can be output in one clock cycle.

[0085] 2) In the routing state machine, data is continuously read from the temporary RAM, 512 bits (64B) of data are obtained in each clock cycle, and the number of rows, row offset, and length of the variable-length column are stored in the meta-information register. If the filter expression of this SQL query statement does not involve variable-length columns, the reading of the variable-length column length part is skipped.

[0086] 3) After the metadata is read, the number of rows is 3; the row offsets are 192, 256, and 320 respectively; the lengths of the variable-length columns in each row are: (10|12), (16|18), and (14|14).

[0087] 4) In the column offset accumulator, the variable-length column length is obtained according to the fixed-length column length configured by the user, and the offset value of each column in the row is calculated in parallel using a parallel adder. For details, see Figure 7 A schematic diagram of an offset value calculation process is shown.

[0088] This embodiment supports pre-screening of variable-length column comparison operations. Compared with fixed-length column comparison operations, variable-length column comparison operations consume more clock cycles. When the length of the variable-length column is different from the constant length input by the user, there is no need to compare the data content of the variable-length column. If there is a variable-length column in this filtering expression, and the variable-length column has the following characteristics in the truth table:

[0089] When the comparison result of the variable-length column is 0, the comparison result of the filter expression must be 0, which is called the top-level AND operation column in the present invention.

[0090] For the top-level AND operation column, when the length of the variable-length column is different from the constant length in the comparison operation, the processing of the row of data is skipped, which saves the time-consuming comparison of variable-length columns.

[0091] For example: Figure 6 In the query, the user query SQL is where (order date < 2020-01-01 or product price > 100) and delivery information = 'AA city BB district'.

[0092] For the filter condition: Delivery information = 'AA city BB district', the logical operation is and. When this condition is not met, there is no need to compare the entire row of data. When the length of the delivery information field in a row is different from the length of 'AA city BB district', there is no need to compare the data in this row.

[0093] Optionally, when the column field is not a top-level AND operation column, the variable-length column field comparator input result can be set to 0, and no comparison calculation is required.

[0094] 5) After reading the meta information, the routing state machine sends the row data information to the row cache register group, splits the data, and assigns it to the register group that can be accessed in parallel. =4B for alignment, so the content of the field can be represented by multiple register numbers. For details, see Figure 8 A schematic diagram of register numbering is shown.

[0095] 6) The register copier will convert the offset of the target column into a register number according to the target column used in the comparison instruction, and copy the target column data from the row register group to the comparator register. Since the column field is always aligned according to the register bit width, there will not be a situation where two column field data are contained in one register. At the same time, since the row data is stored in the register that can be accessed in parallel in the FPGA, the target column field data required in multiple comparison instructions can be copied in parallel.

[0096] The benefit of this arrangement is that it supports full parallel computation of all comparison operations in a filter expression. In this embodiment, the target column fields are all stored in aligned registers. Since registers in FPGA can be accessed in parallel, full parallel computation of all comparison operations in a filter comparison expression can be supported.

[0097] This embodiment supports parallel comparison operations on multiple rows of data. In this embodiment, when the number of comparators in the FPGA is greater than the number of comparison operations required by the current SQL filter expression, parallel comparison operations on multiple rows of data can be supported. For example: FPGA creates parallel comparators for 4 integers, 4 floating points, 4 dates, and 4 strings. When the filter expression in SQL only requires 2 integers, 1 floating point, and 1 date, the target column extraction module in the present invention will assign the target columns in the two rows to the parallel comparators, and use a shift mask to distinguish different rows. In one comparison operation, the comparison operation of 2 rows of data is completed simultaneously. For details, please refer to Fig. 9 A schematic diagram of a comparison operation process is shown.

[0098] The long bit width parallel comparator 302 is used to copy the target column data information to the corresponding long bit width register according to the data type, and obtain the comparison result according to the comparison operation instruction code and the target column data information, and send the comparison result to the matrix multiplier 303; wherein the comparison operation instruction code is sent by the host end to the FPGA hardware accelerated logic design system 10.

[0099] In the design of related technologies, data often needs to be read from the external DDR to the on-chip RAM first, and then read from the on-chip RAM to the register for logical operation. DDR requires multiple clock cycles to complete a data transmission, RAM requires at least one clock cycle to obtain data, and registers can be read by multiple logic circuits as input for parallel calculation within one clock cycle.

[0100] As a general-purpose processor, the CPU has an internal register design of 64 bits (8 bytes). An 8-byte data can be loaded for operation in one clock cycle. At the same time, the number of registers is limited and needs to complete different functions (for example, X86 has 16 general-purpose registers). For a large number of parallel computing requirements, it often causes a large computing delay.

[0101] As a customizable hardware circuit, FPGA can freely set the internal register width, number of registers, and hardware circuit functions using registers. For processing logic with fixed operation modes, using FPGA to design dedicated circuits for calculation can greatly improve computing performance.

[0102] In the related art, comparators of multiple data types are created in each comparison calculation unit. In this embodiment, a parallel comparator composed of multiple long bit width registers is created to improve the comparison performance while reducing the consumption of logic resources on the FPGA chip.

[0103] In one example, the long bit width parallel comparator includes: a plurality of comparison operation instruction code registers, a plurality of long bit width input registers, a plurality of comparators, a plurality of comparison result registers and a shift controller.

[0104] In one example, the comparison operation instruction code register is used to store the comparison operation instruction code; the long bit width input register is used to calculate multiple comparison operation instruction codes of the same data type in parallel.

[0105] The advantage of this arrangement is that it simplifies the comparison operation logic design, and changes the serial comparison operation in the related technology to a parallel comparison operation, thereby improving the operation efficiency.

[0106] In one example, the comparator is used to configure the comparison constant value and the multiplexer in the comparator according to the operator in the comparison operation instruction code configured by the user, and select the comparison operation output result; the comparison result register is used to store the comparison operation output result.

[0107] The benefit of this setting is that in scenarios where the filter expression in the SQL statement uses fewer comparators, it can support parallel comparison of multiple rows to improve calculation performance.

[0108] In one example, the shift controller is used to shift the comparison operation output result according to the shift register value and output it as the comparison result.

[0109] In one example, the number of comparators of the same data type is determined by the maximum number of values ​​of the same data type in multiple tables in the database.

[0110] The benefit of such a setting is that it reduces the consumption of logic resources. In the related art, it is necessary to create comparators of all data types for the comparison operation in each filter expression, resulting in a large number of redundant comparators. This embodiment greatly reduces the number of redundant comparators.

[0111] In an example, for details, see Fig.10 A schematic diagram of the structure of a long bit width parallel comparator is shown.

[0112] In the figure, 802 indicates that the comparison operation of the same data is placed in a group of long bit width registers, and multiple comparison operations of the same data type can be calculated in parallel. If necessary, multiple groups of parallel comparators can be created for a certain data type.

[0113] In the figure, 803 indicates that the comparison constant value and the multiplexer in the comparator are configured according to the operator in the comparison instruction code configured by the user, and the comparison operation output result is selected.

[0114] In the figure, 804 indicates that the result of the comparison operation is shifted according to the shift register value in the shift controller and then output as the result.

[0115] The number of different types of comparators is determined by the maximum number of comparators required for different data types in the table to be queried. Optionally, the user can limit the maximum number of supported parallel comparison filter conditions. Optionally, the user can set multiple comparators for frequently queried table fields.

[0116] To explain it more clearly, let's take an example:

[0117] For example: create three tables in the database: tableA, tableB, tableC. The columns in each table are:

[0118] tableA: 2 integers, 2 dates.

[0119] tableB: 2 floating point values, 1 integer value, and 2 strings.

[0120] tableC: 1 floating point, 1 date, 3 integers, 1 string.

[0121] Assume that the FPGA is designed to support all fields in all tables. The required comparators are: 3 integers, 2 dates, 2 floating points, and 2 strings.

[0122] Optionally, multiple comparators may be created in the FPGA, or comparators of corresponding data types may be created for fields that need to be frequently queried.

[0123] The SQL statement entered by the user will be filtered and calculated table by table, which means that the filter expression in a single table will not use all comparators.

[0124] During initialization, all constants in comparison instructions are copied to corresponding registers according to data types, and the results of different comparison operations are selected according to the operators.

[0125] When the column field extraction module copies the target column field in the table to the long bit width register, all comparison instructions can perform parallel comparison operations within one clock cycle, thereby greatly improving the performance of comparison operations.

[0126] When the number of comparisons used in this SQL query is relatively small and there are multiple comparators in the FPGA, comparison calculations on multiple rows of data can be performed simultaneously, further improving the parallel comparison performance.

[0127] The matrix multiplier 303 is used to determine the search code according to the comparison result and the permutation matrix; wherein the permutation matrix is ​​sent from the host end to the logic design system 10 for FPGA hardware acceleration.

[0128] In one example, the matrix multiplier is specifically used to: store the permutation matrix into a permutation matrix register group; the permutation matrix register group includes a plurality of permutation matrix row registers.

[0129] In one example, the matrix multiplier is specifically used to: perform a bitwise AND operation on the permutation matrix row register and the comparison result register, and output a search code.

[0130] In FPGA, registers can be read and accessed by multiple computing logic blocks at the same time. This feature can realize parallel calculation of multiple rows and columns in matrix operations, which can greatly improve the performance of matrix calculations compared to CPU calculations.

[0131] Since the filter expression in the SQL statement input by the user is variable, and the comparator connection designed in the FPGA is fixed, the results of a parallel comparison operation are often invalid, and the valid comparison results need to be merged together. In the technical solution of the present invention, a key technical problem needs to be solved, namely: how to efficiently merge the calculation results required by the filter expression in this SQL together.

[0132] In the prior art, a relatively easy method is to configure the comparison registers used in this SQL by the user, and then the FPGA performs shift operations one by one to splice the comparison results used in the filter expression in this SQL. However, this method consumes a large number of clock cycles to perform shift and splicing operations. In this embodiment, the problem is solved by using the feature of FPGA that can realize matrix parallel computing. For details, please refer to Fig.11 A schematic diagram of a matrix parallel computing implementation process is shown.

[0133] Feature 901: The result calculated by the comparator is connected to a long bit width register, called a comparison result register. The calculation result of each comparator corresponds to a bit in the register.

[0134] Feature 902: The data in the permutation matrix is ​​stored in long bit width registers by row, forming a plurality of permutation matrix row registers.

[0135] Feature 903: Perform a bitwise AND operation on each permutation matrix row register and the comparison result register, and if the result is 0, output 0, otherwise output 1. (It is worth noting that since there is at most one 1 in each row of the permutation matrix, this calculation operation is equivalent to a matrix multiplication operation).

[0136] Feature 904: The operation results of all permutation matrix row registers and comparison result registers constitute the search code as output.

[0137] For a clearer explanation, see Fig.12 The figure shows a schematic diagram of the calculation process of a matrix multiplier. Four integer, four floating point, four date, and four string comparators are instantiated in the FPGA. Each comparator is connected to 1 bit in the comparator result register, and the comparison result register is 16 bits wide. The permutation matrix register group consists of multiple 16-bit wide registers (RO-R15), and each register stores a row of the permutation matrix. During calculation, each register in the permutation matrix is ​​calculated in parallel with the comparison result register.

[0138] The calculation method is: first perform a bitwise logical AND operation, and then output whether the entire 16-bit operation result is 1. When the SQL filter expression entered by the user requires 2 integers, 1 floating point, 2 dates, and 1 string. In the comparison results of the parallel comparators in the FPGA, the first 2 bits of the 4 integer comparators are valid, the first bit of the floating point comparator is valid, the first 2 bits of the date comparator are valid, and the first bit of the string comparator is valid. If the shift + splicing method is used to splice the valid comparison results. The following steps are required: right shift the floating point comparison result by 2 bits, and perform an AND operation with the integer result to obtain result A. Right shift the date comparison result by 3 bits, and perform an AND operation with result A to obtain result B. Right shift the string comparison result by 65 bits, and perform an AND operation with result B to obtain result C, and result C is used as the search code. In this calculation process, the latter step depends on the calculation result of the previous step, and multiple clock cycles are required to complete, resulting in a large delay. Using the scheme of the present invention, the calculation process of the matrix does not have data dependence, and can be calculated in parallel, thereby reducing the delay of the calculation process.

[0139] The advantage of this setting is that it utilizes the parallel access characteristics of registers in FPGA and the parallel calculation of matrix rows and columns. In one clock cycle, the calculation results required for the filter expression in this SQL statement can be merged together to form a search code, thereby improving calculation efficiency.

[0140] The truth table finder 304 is used to output a matching result according to the lookup code and the logic operation truth table; wherein the matching result is used to be sent to the host end.

[0141] In one example, the truth table finder is specifically used to: use the value generated by the comparison operation result of the logic operation truth table as an index, and write the corresponding result value into an on-chip RAM of a preset bit width.

[0142] In one example, the truth table finder is specifically used to: respond to a query calculation request, use the search code calculated by the permutation matrix as an index, and obtain the corresponding value from the on-chip RAM of a preset bit width as an output result; the output result is used to determine whether the row meets the filtering condition.

[0143] In digital circuits, complex logic operations can be converted into truth table lookups through logic algebra. FPGAs often use this method to accelerate operations.

[0144] For the newly input SQL statement, the SQL statement recoding module generates the logic operation truth table (the second truth table in the previous article), and uses the value generated by the comparison operation result as the index to write the corresponding result value into a 1-bit wide and 1-bit deep The FPGA on-chip RAM (N is the number of comparison calculations used by the current SQL statement).

[0145] When performing query calculations, the search code calculated by the permutation matrix is ​​used as the index, and the corresponding value is obtained from the RAM as the output result. This is used to determine whether the row meets the filter condition.

[0146] For a clearer explanation, see Fig.13 A schematic diagram of the calculation process of a truth table finder is shown.

[0147] For the table and SQL statement in the figure, after the SQL statement recoding module, the second truth table generated uses the comparison operation result as the index and fills the corresponding result value into the truth table RAM. When a query is needed, only one clock cycle is needed to obtain the result of the complex logic operation.

[0148] The benefit of this setting is that it reduces the on-chip RAM space required for the truth table. Since in real scenarios, SQL statements do not use the results of all comparator operations, in this embodiment, the permutation matrix merges the valid comparison results, and the FPGA only needs to create a truth table corresponding to the valid bits, so the required on-chip RAM space is greatly reduced. Only one clock cycle is required to obtain the calculation results of complex logical operations. The FPGA calculation results can be selected in the form of a bitmap to mark which rows in the page meet the filtering conditions. The processing matching row module receives the calculation results of the FPGA, finds the index number of the row that meets the filtering conditions in the page from the bitmap, and then retrieves the row data according to the page row meta information to complete the filtering query process.

[0149] In one example, the FPGA hardware accelerated logic design system is also used for arithmetic operations.

[0150] For example, see Fig.14 A schematic diagram of the arithmetic operation process of a controller hardware accelerated logic design system is shown.

[0151] In one example, the FPGA hardware accelerated logic design system is also used for a column-based storage database.

[0152] The embodiment of the present application provides a controller hardware acceleration method, which can be seen in Fig.15 A flowchart of a controller hardware acceleration method is shown. Specifically, it can be applied to a logic design system of FPGA hardware acceleration, the logic design system of FPGA hardware acceleration communicates with a host end; the logic design system of FPGA hardware acceleration includes: a target column extractor, a long bit width parallel comparator, a matrix multiplier and a truth table finder;

[0153] S1501. Extract target column data information from a preset page through a target column extractor, and send the target column data information to a long bit width parallel comparator; wherein the storage structure of the preset page is that the first preset position stores page meta information, the second preset position stores row meta information, the third preset position stores variable length column length information, and the fourth preset position stores row data.

[0154] S1502. Copy the target column data information to the corresponding long bit width register according to the data type through the long bit width parallel comparator, obtain the comparison result according to the comparison operation instruction code and the target column data information, and send the comparison result to the matrix multiplier; wherein the comparison operation instruction code is sent by the host end to the FPGA hardware accelerated logic design system.

[0155] S1503. Determine the search code according to the comparison result and the permutation matrix through a matrix multiplier; wherein the permutation matrix is ​​sent from the host end to the logic design system accelerated by the FPGA hardware.

[0156] S1504. Output a matching result through a truth table finder according to the lookup code and the logic operation truth table; wherein the matching result is used to be sent to the host end.

[0157] In one example, the target column extractor includes: a page read module and a target column copy module; the page read module includes a memory access controller, a routing state machine and a meta-information register group; the target column copy module includes: a column offset accumulator, a row cache register group and a register copier.

[0158] In one example, the memory access controller includes an on-chip RAM for temporarily storing data, and the data of a preset page is read from a preset disk through the memory access controller, and the data of the preset page is stored in the on-chip RAM for temporarily storing data.

[0159] In one example, data of a preset page stored in an on-chip RAM for temporarily storing data is distributed through a routing state machine.

[0160] In one example, the page meta information, row meta information and variable-length column length information are distributed to the meta information register group through a routing state machine; the routing state machine is used to distribute the row data to the row cache register group.

[0161] In one example, the number of rows, row offset, and variable-length column length in the data of the preset page are stored in the meta-information register group.

[0162] In one example, the column offset accumulator calculates the offset of each column relative to the starting position in parallel according to a preconfigured fixed-length column length list and a variable-length column length, and sends the offset value of the target column field in the row to the register copier.

[0163] In one example, the row data is split by a row cache register group and stored in a plurality of registers that can be accessed in parallel.

[0164] In one example, the register copier copies the corresponding register contents from the row cache register group to the long bit width parallel comparator in parallel according to the target column required in the comparison instruction.

[0165] In one example, the long bit width parallel comparator includes: a plurality of comparison operation instruction code registers, a plurality of long bit width input registers, a plurality of comparators, a plurality of comparison result registers and a shift controller.

[0166] In one example, the comparison operation instruction code is stored in a comparison operation instruction code register; and the input register with a long bit width is used to calculate multiple comparison operation instruction codes of the same data type in parallel.

[0167] In one example, the comparison constant value and the multiplexer in the comparator are configured according to the operator in the comparison operation instruction code configured by the user to select the comparison operation output result; the comparison result register is used to store the comparison operation output result.

[0168] In one example, the comparison operation output result is shifted by a shift controller according to a shift register value and then output as a comparison result.

[0169] In one example, the number of comparators of the same data type is determined by the maximum number of values ​​of the same data type in multiple tables in the database.

[0170] In one example, the permutation matrix is ​​stored in a permutation matrix register group through a matrix multiplier; the permutation matrix register group includes a plurality of permutation matrix row registers.

[0171] In one example, a matrix multiplier performs a bitwise AND operation on the permutation matrix row register and the comparison result register to output a lookup code.

[0172] In one example, a truth table finder uses a value generated by a comparison operation result of a logic operation truth table as an index, and writes a corresponding result value into an on-chip RAM of a preset bit width.

[0173] In one example, a truth table finder responds to a query calculation request, uses a search code calculated by a permutation matrix as an index, and obtains a corresponding value from an on-chip RAM of a preset bit width as an output result; the output result is used to determine whether the row meets the filtering condition.

[0174] In one example, the FPGA hardware accelerated logic design system is also used for arithmetic operations.

[0175] In one example, the FPGA hardware accelerated logic design system is also used for a column-based storage database.

[0176] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.

[0177] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above controller hardware acceleration method embodiments.

[0178] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to extract target column data information from a preset page through a target column extractor when running, and send the target column data information to a long bit width parallel comparator; wherein the storage structure of the preset page is that the first preset position stores page meta information, the second preset position stores row meta information, the third preset position stores variable length column length information, and the fourth preset position stores row data;

[0179] The target column data information is copied to the corresponding long bit width register according to the data type through the long bit width parallel comparator, and the comparison result is obtained according to the comparison operation instruction code and the target column data information, and the comparison result is sent to the matrix multiplier; wherein the comparison operation instruction code is sent from the host end to the logic design system of the controller hardware acceleration;

[0180] Determine the search code by a matrix multiplier according to the comparison result and the permutation matrix; wherein the permutation matrix is ​​sent from the host end to the logic design system of the controller hardware acceleration;

[0181] The matching result is outputted by the truth table finder according to the search code and the logic operation truth table; wherein the matching result is used for sending to the host end.

[0182] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0183] The embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, extracting target column data information from a preset page through a target column extractor, and sending the target column data information to a long bit width parallel comparator; wherein the storage structure of the preset page is that the first preset position stores page meta information, the second preset position stores row meta information, the third preset position stores variable length column length information, and the fourth preset position stores row data;

[0184] The target column data information is copied to the corresponding long bit width register according to the data type through the long bit width parallel comparator, and the comparison result is obtained according to the comparison operation instruction code and the target column data information, and the comparison result is sent to the matrix multiplier; wherein the comparison operation instruction code is sent from the host end to the logic design system of the controller hardware acceleration;

[0185] Determine the search code by a matrix multiplier according to the comparison result and the permutation matrix; wherein the permutation matrix is ​​sent from the host end to the logic design system of the controller hardware acceleration;

[0186] The matching result is outputted by the truth table finder according to the search code and the logic operation truth table; wherein the matching result is used for sending to the host end.

[0187] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0188] The above is a detailed introduction to a controller hardware acceleration method provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A logic design system for controller hardware acceleration, characterized in that: The logic design system accelerated by the controller hardware communicates with the host end; The logic design system of the controller hardware acceleration includes: a target column extractor, a long bit width parallel comparator, a matrix multiplier and a truth table finder; The target column extractor is used to extract target column data information from a preset page, and send the target column data information to the long bit width parallel comparator; wherein the storage structure of the preset page is that the first preset position stores page meta information, the second preset position stores row meta information, the third preset position stores variable length column length information, and the fourth preset position stores row data; The long bit width parallel comparator is used to copy the target column data information to the corresponding long bit width register according to the data type, and obtain a comparison result according to the comparison operation instruction code and the target column data information, and send the comparison result to the matrix multiplier; wherein the comparison operation instruction code is sent from the host end to the logic design system of the controller hardware acceleration; The matrix multiplier is used to determine the search code according to the comparison result and the permutation matrix; wherein the permutation matrix is ​​sent by the host end to the logic design system of the controller hardware acceleration; The truth table finder is used to output a matching result according to the lookup code and the logic operation truth table; wherein the matching result is used to be sent to the host end.

2. The system according to claim 1, characterized in that The target column extractor includes: a page reading module and a target column copying module; the page reading module includes a memory access controller, a routing state machine and a meta information register group; the target column copying module includes: a column offset accumulator, a row cache register group and a register copier.

3. The system according to claim 2, characterized in that The memory access controller comprises a memory for temporarily storing data, and the memory access controller is used to read the data of the preset page from a preset disk, and store the data of the preset page in the memory for temporarily storing data.

4. The system according to claim 2, characterized in that The routing state machine is used to distribute the data of the preset page stored in the memory for temporarily storing data.

5. The system according to claim 4, characterized in that The routing state machine is used to distribute the page meta information, the row meta information and the variable-length column length information to the meta information register group; the routing state machine is used to distribute the row data to the row cache register group.

6. The system according to claim 2, characterized in that The meta information register group is used to store the number of rows, row offset, and variable length column length in the data of the preset page.

7. The system according to claim 2, characterized in that The column offset accumulator is used to calculate the offset of each column relative to the starting position in parallel according to the pre-configured fixed-length column length list and variable-length column length, and send the offset value of the target column field in the row to the register copier.

8. The system according to claim 2, characterized in that The row cache register group is used to split the row data and store them into multiple registers that are accessed in parallel.

9. The system according to claim 2, characterized in that The register copier is used to copy the corresponding register content from the row cache register group to the long bit width parallel comparator in parallel according to the target column required in the comparison instruction.

10. The system according to claim 1, characterized in that The long bit width parallel comparator comprises: a plurality of comparison operation instruction code registers, a plurality of long bit width input registers, a plurality of comparators, a plurality of comparison result registers and a shift controller.

11. The system according to claim 10, characterized in that The comparison operation instruction code register is used to store the comparison operation instruction code; the long bit width input register is used to parallelly calculate a plurality of comparison operation instruction codes of the same data type.

12. The system according to claim 10, characterized in that The comparator is used to configure the comparison constant value and the multiplexer in the comparator according to the operator in the comparison operation instruction code configured by the user, and select the comparison operation output result; the comparison result register is used to store the comparison operation output result.

13. The system according to claim 10, characterized in that The shift controller is used to shift the comparison operation output result according to the shift register value and then output it as the comparison result.

14. The system according to claim 10, characterized in that The number of comparators of the same data type is determined by the maximum number of comparators of the same data type across multiple tables in the database.

15. The system according to claim 1, characterized in that The matrix multiplier is specifically used for: storing the permutation matrix into a permutation matrix register group; the permutation matrix register group includes a plurality of permutation matrix row registers.

16. The system according to claim 15, characterized in that The matrix multiplier is specifically used for: performing a bitwise AND operation on the permutation matrix row register and the comparison result register, and outputting the search code.

17. The system according to claim 1, characterized in that The truth table finder is specifically used to: use the value generated by the comparison operation result of the logic operation truth table as an index, and write the corresponding result value into a memory with a preset bit width.

18. The system according to claim 17, characterized in that The truth table finder is specifically used to: respond to a query calculation request, use the search code calculated by the permutation matrix as an index, and obtain the corresponding value from a memory of a preset bit width as an output result; the output result is used to determine whether the row in the preset page meets the filtering condition.

19. The system according to claim 1, characterized in that The controller hardware accelerated logic design system is also used for arithmetic operations.

20. The system according to claim 1, characterized in that The controller hardware accelerated logic design system is also used for a column-type storage database.

21. A controller hardware acceleration method, characterized in that: A logic design system applied to the controller hardware acceleration, wherein the logic design system for the controller hardware acceleration communicates with a host end; The logic design system of the controller hardware acceleration includes: a target column extractor, a long bit width parallel comparator, a matrix multiplier and a truth table finder; The target column extractor extracts target column data information from a preset page, and sends the target column data information to the long bit width parallel comparator; wherein the storage structure of the preset page is that the first preset position stores page meta information, the second preset position stores row meta information, the third preset position stores variable length column length information, and the fourth preset position stores row data; The target column data information is copied to the corresponding long bit width register according to the data type through the long bit width parallel comparator, and a comparison result is obtained according to the comparison operation instruction code and the target column data information, and the comparison result is sent to the matrix multiplier; wherein the comparison operation instruction code is sent from the host end to the logic design system of the controller hardware acceleration; Determine the search code by the matrix multiplier according to the comparison result and the permutation matrix; wherein the permutation matrix is ​​sent by the host end to the logic design system of the controller hardware acceleration; The truth table finder outputs a matching result according to the lookup code and the logic operation truth table; wherein the matching result is used to be sent to the host end.

22. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the controller hardware acceleration method as claimed in claim 21 when executing the computer program.

23. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the controller hardware acceleration method as claimed in claim 21.

24. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the controller hardware acceleration method as claimed in claim 21 are implemented.

Citation Information

Patent Citations

  • Systems and methods for performing instructions specifying ternary tile logic operations

    CN110909883A

  • Hardware acceleration card, acceleration system, method, equipment, medium and program product

    CN118503204A