Database acceleration system and database acceleration method

The database acceleration system composed of a central processing unit and a vector processor uses operation plans and preset execution units to process database data, solving the problem of insufficient instruction domain specificity and improving database operation efficiency.

CN120316096BActive Publication Date: 2025-09-05INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510780609.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-05
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In the prior art, when database data is processed using fixed data processing instructions, the instructions are prone to lack of domain specificity, resulting in low database operation efficiency.

Method used

The database acceleration system consists of a central processing unit, an internal bus and a vector processor. It generates an operation plan by analyzing operation requests and uses the preset execution units in the vector processor to process the target data, thereby achieving instruction expansion and improving database operation efficiency.

Benefits of technology

By expanding the instructions of each preset execution unit, the lack of domain specificity of instructions is avoided and the operational efficiency of the database is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316096B_ABST
    Figure CN120316096B_ABST
Patent Text Reader

Abstract

The present application discloses a database acceleration system and a database acceleration method, which relate to the field of data processing technology, and include: a central processing unit (CPU), an internal bus, a vector processor, and a plurality of preset execution units; the CPU receives an operation request sent by a user terminal; parses the operation request into an operation plan; retrieves corresponding target data from a database according to the operation plan; sends the operation plan and the target data to the internal bus; the internal bus sends the operation plan and the target data to the vector processor; the vector processor controls one or more preset execution units to process the target data according to the operation plan to obtain a processing result corresponding to the operation plan; the processing result is returned to the CPU; and the CPU outputs the processing result to the user terminal. By extending the instructions of each preset execution unit, the situation of insufficient domain specificity of the instructions is avoided, and the operation efficiency of the database is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a database acceleration system and a database acceleration method. Background Art

[0002] A database query is a query that retrieves data from a database using a specific query statement. The database query process primarily uses structured query language to set query conditions and filter target data, and then returns the corresponding query results to the user in a table or other format. With the surge in demand for database queries, improving database response speed has become increasingly important.

[0003] In the related art, current database queries primarily utilize structured query languages ​​(SQLs) to process the corresponding data in the database using fixed data processing instructions to obtain processing results. However, this approach of processing database data using fixed data processing instructions in the related art is prone to insufficient domain specificity of the instructions, resulting in low database operation efficiency. Summary of the Invention

[0004] The present application provides a database acceleration system and a database acceleration method to at least solve the problem in the related art that database data is processed through fixed data processing instructions, which is prone to insufficient instruction domain specificity, resulting in low database operation efficiency.

[0005] The present application provides a database acceleration system, comprising: a central processing unit (10) and a coprocessor (20), wherein the coprocessor (20) comprises at least an internal bus (201) and a vector processor (202); the central processing unit (10) is communicatively connected to the internal bus (201); the internal bus (201) is communicatively connected to the vector processor (202); the vector processor (202) comprises at least a plurality of preset execution units (2021), and each preset execution unit (2021) comprises one or more instructions;

[0006] The central processing unit (10) receives an operation request sent by a user terminal (30); parses the operation request into an operation plan, wherein the operation plan is an operation step corresponding to the operation request; retrieves corresponding target data from a database according to the operation plan; and sends the operation plan and the target data to an internal bus (201);

[0007] Internal bus (201), sends operation plan and target data to vector processor (202);

[0008] The vector processor (202) controls one or more preset execution units (2021) to process the target data according to the operation plan to obtain a processing result corresponding to the operation plan; and returns the processing result to the central processing unit (10);

[0009] The central processing unit (10) outputs the processing result to the user terminal (30).

[0010] This application also provides a database acceleration method, including:

[0011] The CPU receives the operation request from the user end; parses the operation request into an operation plan; loads the corresponding operation microcode from the host memory according to the operation plan; and sends the operation microcode to the internal bus;

[0012] Internal bus, which sends the operation microcode to the vector processor;

[0013] The vector processor pre-fetches corresponding target data from the database data of the memory controller through the operation microcode; controls one or more preset execution units through the operation microcode to process the target data to obtain the processing result corresponding to the operation plan; and returns the processing result to the central processing unit;

[0014] The central processing unit outputs the processing results to the user end.

[0015] The database acceleration system and database acceleration method provided in the embodiments of the present application are composed of a central processing unit, an internal bus, a vector processor and multiple preset execution units to form a database acceleration system; the central processing unit receives an operation request sent by a user terminal; parses the operation request into an operation plan; retrieves corresponding target data from the database according to the operation plan; sends the operation plan and database data to the internal bus; the internal bus sends the operation plan and target data to the vector processor; the vector processor controls one or more preset execution units to process the target data according to the operation plan to obtain a processing result corresponding to the operation plan; the processing result is returned to the central processing unit; and the central processing unit outputs the processing result to the user terminal. By expanding the instructions of each preset execution unit, the situation of insufficient domain specificity of the instructions is avoided, and the operating efficiency of the database is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1Schematic diagram of the structure of the database acceleration system provided in the embodiment of the present application Figure 1 ;

[0018] Figure 2 Schematic diagram of the structure of the database acceleration system provided in the embodiment of the present application Figure 2 ;

[0019] Figure 3 A flowchart of the database acceleration method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0021] It should be noted that the terms "center," "longitudinal," "transverse," "length," "width," "thickness," "upper," "lower," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," "circumferential," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended solely for ease of description and simplification of the present application. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limiting the present application. The terms "mounted," "connected," and "connected" should be interpreted broadly, and may include, for example, fixed, removable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediary; and internal communication between two components. The terms "parallel," "perpendicular," and "equal" encompass the described conditions and conditions similar to the described conditions, provided that the range of the similar conditions is within an acceptable range of deviation, as determined by a person of ordinary skill in the art taking into account the measurement in question and the errors associated with the measurement of the particular quantity (i.e., the limitations of the measurement system). For example, "parallel" includes both absolute parallelism and approximate parallelism, where the acceptable deviation range for approximate parallelism may be, for example, within 5°; "perpendicular" includes both absolute perpendicularity and approximate perpendicularity, where the acceptable deviation range for approximate perpendicularity may also be, for example, within 5°. "Equal" includes both absolute equality and approximate equality, where the acceptable deviation range for approximate equality may be, for example, that the difference between the two is less than or equal to 5% of either. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.

[0022] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0023] A database query is a query that retrieves data from a database through a user's specific query statement. The database query process mainly uses structured query language to set query conditions and filter target data, and returns the corresponding query results to the user in the form of a table, etc. It is often used in scenarios such as data analysis and system development. With the surge in demand for database queries, it is particularly important to improve the response speed of the database. In the related art, current database queries mainly use structured query language to process the corresponding data in the database with fixed data processing instructions to obtain processing results. However, in the related art, the method of processing database data through fixed data processing instructions is prone to insufficient domain specificity of the instructions, which leads to low database operation efficiency.

[0024] In order to solve the above technical problems, the embodiments of the present application propose the following technical concepts: the inventor considers the operation requests sent by the user end, builds a database acceleration system based on the central processing unit and the coprocessor, uses the central processing unit to parse the operation request into an operation plan, and retrieves the target data from the database according to the operation plan, uses the internal bus to send the operation plan and the target data to the vector processor, uses one or more preset execution units in the vector processor to process the target data according to the operation plan to obtain the processing result, and outputs the processing result from the central processing unit to the user end. By expanding the instructions of each preset execution unit, the situation of insufficient domain specificity of the instruction is avoided, and the operation efficiency of the database is improved.

[0025] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0026] In conjunction with the specific application environment architecture or specific hardware architecture that the execution of the database acceleration system depends on, the specific application environment architecture or specific hardware architecture is described here. Figure 1 , Figure 1 Schematic diagram of the structure of the database acceleration system provided in the embodiment of the present application Figure 1 .

[0027] like Figure 1The database acceleration system includes: a central processing unit 10 and a coprocessor 20, the coprocessor 20 includes at least an internal bus 201 and a vector processor 202; the central processing unit 10 is communicatively connected to the internal bus 201; the internal bus 201 is communicatively connected to the vector processor 202; the vector processor 202 includes at least a plurality of preset execution units 2021, each preset execution unit 2021 includes one or more instructions.

[0028] The coprocessor may be a coprocessor based on an open source instruction set or other processors.

[0029] Among them, the open source instruction set can be RISC-V or other instruction sets.

[0030] Among them, the RISC-V coprocessor is a coprocessor with an open standard instruction set architecture designed based on the reduced instruction set computing principle.

[0031] In this embodiment, the structure data of the preset execution unit includes one or more of the following: an operation code, a function code, a sub-function code, a target register, a first source register, a second source register, and a mask pattern.

[0032] In this embodiment, the opcode field is opcode, which is used for custom instructions; the function code field is funct7, which is used to distinguish different instructions; the sub-function code field is funct3, which is used to further subdivide instructions; the target register field is rd; the first source register field is rs1; the second source register field is rs2; the mask mode field is vm, where the number 0 indicates no mask and the number 1 indicates mask control.

[0033] The structure data of the operation code includes one or more of the following: arithmetic operation code, load code, store code, branch code, system instruction and custom instruction.

[0034] Among them, the arithmetic operation code is 0110011, and example instructions are add, sub, and xor; the load code is 0000011, and example instructions are lw, lb, and lh; the storage code is 0100011, and example instructions are sw, sb, and sh; the branch code is 1100011, and example instructions are beq, bne, and blt; the system instruction is 1110011, and example instructions are ecall and csrrw; the custom instruction is 1101011, and example instructions are vfilter and vhash.

[0035] For example, 1101011 is an opcode reserved for non-standard extensions in the open source instruction set, used to implement custom instructions. After the CPU recognizes opcode = 1101011, it routes the instruction to the vector processor and selects the corresponding hardware circuit based on funct7 / funct3.

[0036] For example, when vm=1, mask control is enabled, and operations are performed only on vector elements whose corresponding bits of vm are 1; when vm=0, mask control is disabled, and all elements participate in the calculation.

[0037] Where vm is a 1-bit mask register or consists of specific bits of the vector mask register v0. Each mask bit corresponds to a vector element, for example, bit i of vm controls the comparison of v2[i] and v3[i].

[0038] In this embodiment, the multiple preset execution units 2021 include one or more of the following: a data filtering execution unit, a hash connection execution unit, an aggregation calculation execution unit, a regular matching execution unit, a sorting execution unit, a full table scan execution unit, and an index tree operation execution unit.

[0039] Among them, the data filtering execution unit includes comparison operation instructions; the hash connection execution unit includes hash table construction instructions, hash detection instructions and nested loop connection instructions; the aggregation calculation execution unit includes sum calculation instructions, count calculation instructions, maximum and minimum value instructions and average value instructions; the regular matching execution unit includes pattern compilation instructions, vector matching instructions and result extraction instructions; the sorting execution unit includes radix sorting instructions and merge sorting instructions; the full table scan execution unit includes initialization scan instructions, batch data acquisition instructions and column projection instructions; the index tree operation execution unit includes initialization index instructions, precise operation instructions and range operation instructions.

[0040] The index tree operation execution unit is an index B+ tree operation execution unit.

[0041] The instruction format of the data filtering execution unit is:

[0042]

[0043] The last three digits of the function code funct7 indicate the data type, and 001 indicates fixed point (default).

[0044] In addition, the last three digits of the function code funct7, 010, represent floating point (single precision); 011 represents floating point (double precision); 100 represents 16-bit fixed point; 101 represents character type; 110 represents Boolean type; and 111 represents special mode (such as NULL value processing).

[0045] The last five digits of the function code funct7 indicate the instruction type, and 00001 indicates the data filtering instruction vfilter.

[0046] Sub-function code funct3=000 is equal to mode; funct3=001 is greater than mode; funct3=010 is less than mode; funct3=011 is greater than or equal to mode; funct3=100 is less than or equal to mode.

[0047] Among them, the equal mode, greater than mode, less than mode, greater than or equal to mode and less than or equal to mode are collectively referred to as comparison operation instructions.

[0048] The operation is vd[i]=(vs1[i] op vs2[i])?1:0. For example, vfilter v1, v2, v3, vm, funct3=101, funct7=1100001#v1=(v2>v3)?1:0; if v2=[10, 20, 30], v3=[5, 25, 30], then v1=[1, 0, 0].

[0049] The instruction format of the hash join execution unit is:

[0050]

[0051] The last five digits of the function code funct7 represent the instruction type, and 00010 represents the hash connection instruction vhash.

[0052] Sub-function code funct3=010 is a hash table construction instruction.

[0053] In addition, funct3=011 is a hash detection instruction; funct3=100 is a nested loop connection instruction.

[0054] For example, vhash.buildv1, v2, funct7=0000001 # Build a hash table with the key of v2 and store it in v1; vhash.probe v3, v1, v4, funct7=0000001 # Use the key of v4 to search in the hash table of v1.

[0055] The instruction format of the aggregate computing execution unit is:

[0056]

[0057] The last five digits of the function code funct7 represent the instruction type, and 00011 represents the aggregate calculation instruction vagg.

[0058] Sub-function code funct3=010 is the sum calculation instruction.

[0059] In addition, funct3=011 is the maximum value instruction; funct3=100 is the minimum value instruction; funct3=101 is the average value instruction; funct3=110 is the count calculation instruction.

[0060] Example: vagg.SUM v1,v2,v3#v1=SUM(v2+v3).

[0061] In addition, the aggregate function accelerator supports sum calculation instructions, count calculation instructions, maximum and minimum value instructions, and average value instructions, a multi-level accumulator array vectorized group key matching engine, and a distributed result cache.

[0062] Among them, the instruction format of the regular matching execution unit is:

[0063]

[0064] The last five digits of the function code funct7 represent the instruction type, and 00100 represents the regular matching instruction vregex.

[0065] Sub-function code funct3=001 is the mode compilation instruction.

[0066] In addition, funct3=010 is a vector matching instruction; funct3=011 is a result extraction instruction.

[0067] For example: vregex.match v1, v2, v3#v1= (whether v3 matches the regular expression of v2).

[0068] Among them, the instruction format of the sorting execution unit is:

[0069]

[0070] Among them, the last five bits of the function code funct7 represent the instruction type, and 00101 represents the sort instruction vsort.

[0071] Sub-function code funct3 = 000 is a radix sort instruction. In addition, funct3 = 001 is a merge sort instruction.

[0072] Example 1, vsort.radix v1, v2 # Sort the data in v2 by radix and store the result in v1.

[0073] Input: v2 = [3, 1, 4, 2]

[0074] Output: v1 = [1, 2, 3, 4]

[0075] Example 2, vsort.merge v1, v2, v3 #Merge the sorted v2 and v3, and store the result in v1.

[0076] Input: v2=[1, 3], v3=[2, 4]

[0077] Output: v1 = [1, 2, 3, 4]

[0078] Among them, the instruction format of the full table scan execution unit is:

[0079]

[0080] The last five digits of the function code funct7 represent the instruction type, and 00110 represents the full table scan instruction vscan.

[0081] Sub-function code funct3 = 000 is the initialization scan instruction. In addition, funct3 = 001 is the batch acquisition data instruction; funct3 = 010 is the column projection instruction.

[0082] Among them, the initialization scan instruction is: vscan.init v1, 0x8000#Load the table descriptor from memory address 0x8000 to v1.

[0083] A table description includes: the starting address of the table data in memory, the byte size of each row of data in the table, and the offset of each column in the table relative to the starting position of the row (up to 8 columns are supported).

[0084] The batch data acquisition instruction is: vscan.next v2#load the next batch of data into v2.

[0085] The data layout is: column storage, v2 is organized by column, for example: v2[0:63] = column 1, v2[64:127] = column 2; row storage: v2 is organized by row.

[0086] The row and column adaptive processing is as follows: Column memory processing optimization mode: Dynamically identify hot columns in the row and prioritize them; Use vsan.next instruction to batch obtain column data; Column data prefetching and vector register alignment

[0087] Hybrid processing mode: Use column-based processing for frequently accessed columns; use row-based processing for operations that require access to the entire row; and select the optimal access mode.

[0088] Among them, the column projection instruction is: vscan.project v3,v2,0x0F #Select columns 1-4 from v2 (mask 0x0F=00001111).

[0089] Input: v2=[A1,B1,C1,D1,A2,B2,C2,D2,...]

[0090] Output: v3=[A1,A2,...,B1,B2,...,C1,C2,...]

[0091] The instruction format of the index tree operation execution unit is:

[0092]

[0093] The last five digits of the function code funct7 represent the instruction type, and 00111 represents the index query instruction vindex.

[0094] Sub-function code funct3=000 is the initialization index instruction.

[0095] In addition, funct3=001 is a precise operation instruction; funct3=010 is a range operation instruction.

[0096] Exemplarily, the index initialization instruction is: vindex.init v1, 0x9000 #Load the B+ tree index descriptor (address 0x9000) to v1.

[0097] The index descriptor is: root node address, key type and key byte number.

[0098] For example, the precise query instruction is: vindex.search v2, v1, v3#Search for key v3 in index v1 and store the result in v2.

[0099] Input: v3=

[42] (query key value)

[0100] Output: v2 = [0x1234] (matching row address)

[0101] Exemplarily, the range operation instruction is: vindex.range v2, v1, v3#query the entries in v1 that satisfy v3[0]≤key≤v3[1].

[0102] Input: v3 = [10, 20] (range boundary)

[0103] Output: v2=[0x1234,0x5678,...] (matching row address list)

[0104] The central processing unit 10 receives an operation request sent by the user terminal 30; parses the operation request into an operation plan, wherein the operation plan is the operation steps corresponding to the operation request; retrieves corresponding target data from the database according to the operation plan; and sends the operation plan and target data to the internal bus 201.

[0105] In this embodiment, the original data in the database comes from the hard disk. Specifically, the central processing unit retrieves the original data from the hard disk into the database and runs the database.

[0106] Specifically, the central processing unit 10 receives an operation request sent by the user terminal 30; performs syntax parsing and semantic checking on the operation request to obtain a preliminary operation plan; intervenes in the initial operation plan according to a preset plug-in, rewrites the operation plan based on rules and cost models to identify accelerable operation segments (such as large-scale data filtering, complex connections and aggregate calculations, etc.), and obtains an operation plan containing acceleration instructions; retrieves corresponding target data from the database according to the operation plan; and sends the operation plan and target data to the internal bus 201.

[0107] The operation request may be an SQL operation request or other operation request.

[0108] The preset plug-in may be a Hook plug-in or other plug-ins.

[0109] In addition, it is necessary to make a feasibility judgment on the operation plan containing the acceleration instruction to check whether the operation plan can trigger the preset execution unit.

[0110] In addition, for accelerable operations, an optimized vector instruction sequence is generated to insert necessary control instructions and data handling instructions.

[0111] For example: the traditional WHERE condition corresponds to the vfilter.cond instruction; the JOIN operation corresponds to the vhash.build+vhash.probe instruction sequence; the GROUP BY aggregation corresponds to the vagg.group+vagg.reduce instruction sequence.

[0112] The internal bus 201 sends the operation plan and target data to the vector processor 202.

[0113] The vector processor 202 controls one or more preset execution units 2021 to process the target data according to the operation plan to obtain the processing results corresponding to the operation plan; and returns the processing results to the central processing unit 10;

[0114] Exemplarily, the one or more preset execution units are a data filtering execution unit, a hash connection execution unit, and an aggregation calculation execution unit.

[0115] Exemplarily, the vector processor executes database-specific instructions in parallel: a vector filter execution unit, a parallel comparator array processes predicate judgments on target data; a hash join execution unit, a dedicated hash calculation unit accelerates the detection phase; an aggregation calculation execution unit, a multi-level accumulator array implements parallel aggregation.

[0116] In addition, if only part of the operation plan needs to be accelerated, the data corresponding to the part needs to be accelerated, and the data corresponding to the remaining part of the operation plan needs to be processed normally to obtain a processing result that integrates accelerated processing and normal processing.

[0117] In addition, memory consistency management is also required: maintaining cache consistency through extended protocols; using atomic operations on critical data structures, such as index B+ tree operations.

[0118] Furthermore, if data processing is abnormal, when abnormal data is detected during vectored execution, the exception context is saved and an interrupt is triggered, and the CPU takes over the exception handling.

[0119] In addition, when each preset execution unit competes for accelerator resources, it is called in a round-robin manner based on priority and time slice.

[0120] The central processing unit 10 outputs the processing result to the user terminal 30.

[0121] In this embodiment, the user terminal may be a display terminal, a mobile device, or other terminals.

[0122] The database acceleration system provided in this embodiment is composed of a central processing unit, an internal bus, a vector processor and multiple preset execution units. The central processing unit receives an operation request sent by a user terminal, parses the operation request into an operation plan, retrieves corresponding target data from the database according to the operation plan, sends the operation plan and database data to the internal bus, and the internal bus sends the operation plan and target data to the vector processor. The vector processor controls one or more preset execution units to process the target data according to the operation plan to obtain a processing result corresponding to the operation plan, and returns the processing result to the central processing unit. The central processing unit outputs the processing result to the user terminal. By expanding the instructions of each preset execution unit, the lack of domain-specificity of the instructions is avoided, thereby improving the operating efficiency of the database.

[0123] Figure 2 Schematic diagram of the structure of the database acceleration system provided in the embodiment of the present application Figure 2 .

[0124] like Figure 2 ,exist Figure 1 Based on the embodiment, the coprocessor 20 also includes: a memory controller 203; the memory controller 203 is communicatively connected to the internal bus 201; wherein the central processing unit 10 has pre-retrieved the database data in the database to the memory controller 203, and the memory controller 203 includes multiple storage channels.

[0125] The memory controller may be an HBM memory controller or other controllers.

[0126] Among them, the HBM memory controller is a high-bandwidth memory controller.

[0127] The plurality of storage channels may be any number of 4, 6 or 8, or any other number.

[0128] The central processing unit 10 receives an operation request sent by the user terminal 30 ; parses the operation request into an operation plan; loads the corresponding operation microcode from the host memory according to the operation plan; and sends the operation microcode to the internal bus 201 .

[0129] In this embodiment, the operation microcode is a layer of abstraction of the hardware structure by a computer or processor, or a data structure for implementing complex machine instructions.

[0130] The format of the microcode includes: operation code, address attributes, format string of output address, batch size and binary mask.

[0131] Among them, opcode, type: uint8_t, description: The opcode field follows the open source instruction set standard and is used to identify the specific micro-operation type (such as load, store, or data movement).

[0132] Address attribute, type: uint64_t, description: points to the address of the predicate condition, used to implement conditional execution. The condition judgment result can be read through this address to decide whether to execute the current microinstruction.

[0133] Output address format string, type: uint64_t, description: the target address of the result buffer. The data will be written to this location after the storage operation is completed. It may be used in memory-mapped I / O or direct memory access scenarios.

[0134] Batch size, type: uint32_t, description: Specifies the batch size for single processing, supports batch operation optimization, for example, processing multiple data units at a time to reduce overhead.

[0135] Binary mask, type: uint8_t, description: Column selection mask, used to select specific columns when operating on multiple columns of data. Each bit may correspond to a column, and a mask value of 1 indicates that the column is selected.

[0136] In addition, if the microcode is missing, it will fall back to the central processing unit for execution and record the missing microcode characteristics, triggering the dynamic microcode generation process.

[0137] The internal bus 201 sends the operation microcode to the vector processor 202.

[0138] The vector processor 202 pre-fetches the corresponding target data from the database data of the memory controller 203 through the operation microcode; controls one or more preset execution units 2021 through the operation microcode to process the target data to obtain the processing result corresponding to the operation plan; and returns the processing result to the central processing unit 10.

[0139] In this embodiment, the discussion on the preset execution unit has been Figure 1 The corresponding embodiments are described in detail and will not be repeated here.

[0140] The central processing unit 10 outputs the processing result to the user terminal 30.

[0141] In this embodiment, the discussion about the user terminal has been Figure 1 The corresponding embodiments are described in detail and will not be repeated here.

[0142] Continue to refer Figure 2 , the vector processor 202 also includes: a vector register 2022.

[0143] The vector processor 202 pre-fetches corresponding target data from the database data of the memory controller 203 by operating the microcode.

[0144] In this embodiment, the database data is original records or information sets stored in a structured manner.

[0145] One or more preset execution units 2021 process the target data according to the operation microcode to obtain a processing result corresponding to the operation plan; and output the processing result to the vector register 2022.

[0146] The vector register 2022 returns the processing result to the central processing unit 10.

[0147] Continue to refer Figure 2 , the coprocessor 20 also includes: a scalar processor 204.

[0148] The coprocessor 20 verifies the operation microcode to obtain a verification result.

[0149] After determining that the verification result is passed, the scalar processor 204 performs initialization processing on the operation microcode to complete the initialization of the operation microcode.

[0150] Continue to refer Figure 2 , the scalar processor 204 also includes: a control register 2041.

[0151] The control register 2041 activates the operation microcode and establishes a mapping relationship between the instruction operation code and the operation microcode to complete the initialization of the operation microcode.

[0152] In addition, when adding new processing microcode, there is no need to restart the system, and the new microcode can be directly loaded to overwrite the old microcode.

[0153] The database acceleration system provided in this embodiment also includes a memory controller, and the central processing unit has pre-retrieved the database data in the database to the memory controller; receives the operation request sent by the user end through the central processing unit; parses the operation request into an operation plan; loads the corresponding operation microcode from the host memory according to the operation plan; sends the operation microcode to the internal bus; the internal bus sends the operation microcode to the vector processor; the vector processor pre-fetches the corresponding target data from the database data of the memory controller through the operation microcode; controls one or more preset execution units to process the target data through the operation microcode to obtain the processing result corresponding to the operation plan; returns the processing result to the central processing unit; the central processing unit outputs the processing result to the user end, retrieves the database data in the database to the memory controller through the central processing unit, and the operation microcode pre-fetches the corresponding target data from the database data of the memory controller, thereby realizing the near storage of the database acceleration system, breaking through the storage wall bottleneck, accelerating the database, and further improving the operation efficiency of the database.

[0154] In addition, the database acceleration system provided in this embodiment integrates a multi-channel memory controller and supports concurrent access, thereby improving bandwidth utilization.

[0155] Figure 3 This is a flow chart of the database acceleration method provided in the embodiment of the present application. The execution subject of this embodiment can be Figure 2 The database acceleration system in the embodiment shown is not particularly limited in this embodiment. Figure 3 As shown, the method includes:

[0156] S301: The central processing unit receives an operation request sent by the user terminal; parses the operation request into an operation plan; loads the corresponding operation microcode from the host memory according to the operation plan; and sends the operation microcode to the internal bus.

[0157] Specifically, the corresponding operation microcode is loaded from the host memory according to the operation plan through the microcode loading instruction.

[0158] Among them, the microcode loading instruction is xLOAD_MICROCODE.

[0159] Among them, the format of the microcode load instruction is: microcode address, source register, sub-function code, target register and operation code, and the function is: to load microcode from the host memory to the coprocessor cache.

[0160] For example, LOAD_MICROCODE 0x8000# loads microcode from 0x8000.

[0161] In addition, it is necessary to identify the special microcode corresponding to the operation plan and check whether the corresponding version already exists in the coprocessor microcode cache.

[0162] S302: Internal bus, sending the operation microcode to the vector processor.

[0163] S303: The vector processor pre-fetches the corresponding target data from the database data of the memory controller through the operation microcode; controls one or more preset execution units through the operation microcode to process the target data to obtain the processing result corresponding to the operation plan; and returns the processing result to the central processing unit.

[0164] In this embodiment, the microcode is decomposed into storage microcode and computing microcode. Accordingly, in step S303, the corresponding target data is pre-fetched from the database data of the memory controller through the operation microcode, specifically: the corresponding target data is pre-fetched from the database data of the memory controller through the storage microcode.

[0165] In this embodiment, the storage microcode is responsible for data handling, storage format conversion (row-column conversion) and prefetch strategy control.

[0166] For example, store_microcode_trans defines the data transfer path from the memory controller to the vector processor, supports dynamic adjustment of prefetch size, sequential access prefetches 64KB, and random access prefetches 8KB.

[0167] In this embodiment, computing microcode is used to implement hardware acceleration logic for specific complex database operations.

[0168] Illustratively, compute_microcode_regex contains the logic for regular expression matching operations.

[0169] The prefetch process is as follows: Query start: The access pattern analysis unit is initialized, and the memory controller loads data according to the default row format; Pattern recognition: The access pattern analysis unit detects sequential scan columns, marks them as hot data, and triggers column-based conversion; Prefetch trigger: The intelligent prefetch engine prefetches 256KB of data for subsequent columns; Instruction execution: vscan.project directly operates on column-based data, avoiding row parsing overhead, for example: vscan.project v3, v2, 0x03 # selects columns 1-2; Dynamic adjustment: If subsequent queries change to random access columns, the prefetch strategy is switched to 8KB, and the columns may be converted to hot data.

[0170] In this embodiment, the target data includes hot data and cold data, and the database data includes multiple columns of data. Accordingly, the above steps specifically include:

[0171] S3031: Obtain the number of accesses to each column of data in a predetermined dominant access mode according to a preset counter array by storing microcode.

[0172] Specifically, the process of determining the dominant access mode in step S3031 includes:

[0173] S30311: Obtain the number of times a memory controller is accessed by multiple access modes within a preset time.

[0174] In this embodiment, the structure of the access mode includes one or more of the following types: prefetch data size, prefetch trigger condition, and prefetch address cache.

[0175] In this embodiment, the multiple access modes include at least a sequential access mode and a random access mode.

[0176] Among them, the sequential access mode is: prefetch data size: prefetch data from the memory controller to utilize high-bandwidth concurrency; prefetch trigger condition: trigger prefetch after detecting three consecutive sequential accesses; prefetch address calculation: predict the next batch of data addresses based on the current address step size.

[0177] Among them, the random access mode is: prefetch data size: conservatively prefetch 8KB to reduce invalid data movement; prefetch trigger condition: triggered based on access locality (such as cache line alignment); prefetch address cache: record the address range of recent random accesses, and give priority to prefetching high-frequency areas.

[0178] The relevant bandwidth allocation is as follows: in sequential mode, the memory controller channel bandwidth is allocated to 70% prefetch and 30% real-time request; in random mode, it is adjusted to 50% prefetch and 50% real-time request.

[0179] S30312: Determine a dominant access pattern from multiple access patterns based on the number of accesses.

[0180] Specifically, a dominant access pattern is determined from multiple access patterns according to the number of accesses and a preset threshold.

[0181] The preset threshold may be any number among 100 times, 200 times or 300 times, or other numbers.

[0182] S3032: Sort the access times to obtain an access sequence.

[0183] S3033: Determine the frequently accessed column data in each column data according to the access sequence.

[0184] In this embodiment, a preset top number of columns of data in the access sequence are determined as high-frequency access column data.

[0185] The preset top quantity may be top 10%, top 20%, or other quantities.

[0186] S3034: Determine the frequently accessed column data as hot data, and determine the remaining column data in each column data as cold data.

[0187] In addition, after step S3034, the method further includes steps a and b:

[0188] Step a: Store the hot data in the memory controller in column format and record the starting address, length, and data type of each column of data.

[0189] Specifically, the memory controller stores the data continuously in columns, aligns each column of data to a 512-bit vector register, and records the starting address, length, and data type of each column of data for parsing by the vscan.project instruction.

[0190] Step b: Store the cold data in the memory controller in rows.

[0191] Specifically, cold data is stored in the memory controller in rows to reduce conversion overhead.

[0192] In addition, lazy transformations are used to trigger row-column transformations only when a column is promoted to hot data.

[0193] Specifically, in step S303, the target data is processed by controlling one or more preset execution units through the operation microcode to obtain the processing result corresponding to the operation plan. Specifically, the target data is processed by controlling one or more preset execution units through the calculation microcode to obtain the processing result corresponding to the operation plan.

[0194] In this embodiment, the processing result includes one or more sub-processing results; accordingly, the above steps specifically include:

[0195] S3035: By calculating the microcode, operating the corresponding one or more preset execution units, processing the target data to obtain one or more sub-processing results.

[0196] S3036: Sort the sub-processing results to obtain a processing result sequence.

[0197] S3037: Aggregate the processing result sequence to obtain an aggregated processing result.

[0198] S3038: Merge the processing result sequence to obtain a merged processing result.

[0199] S3039: Convert the merged processing results to obtain a processing result.

[0200] Specifically, the merged processing results are converted into a row storage format to obtain the processing results.

[0201] S304: The central processing unit outputs the processing result to the user end.

[0202] The database acceleration method provided in this embodiment receives an operation request sent by a user terminal through a central processing unit; parses the operation request into an operation plan; loads a corresponding operation microcode from a host memory according to the operation plan; sends the operation microcode to an internal bus; the internal bus sends the operation microcode to a vector processor; the vector processor pre-fetches corresponding target data from the database data of a memory controller through the operation microcode; controls one or more preset execution units to process the target data through the operation microcode to obtain a processing result corresponding to the operation plan; returns the processing result to the central processing unit; the central processing unit outputs the processing result to the user terminal, retrieves the database data in the database to the memory controller through the central processing unit, and the operation microcode pre-fetches the corresponding target data from the database data of the memory controller, thereby realizing near storage of the database acceleration system and improving the operation efficiency of the database.

[0203] In addition, the database acceleration method provided in this embodiment processes database data by decomposing microcode into storage microcode and computing microcode to enhance the dynamic adaptability of hardware.

[0204] It should be noted that the database acceleration method provided in this embodiment also includes a dynamic conversion mechanism, specifically:

[0205] Mask control: 8-bit mask to select the target column, for example: 0x0F selects the first 4 columns.

[0206] Column offset configuration: Specify the offset of the corresponding column in the row through the control register.

[0207] Row-column conversion unit: A dedicated circuit realizes real-time conversion of rows to columns or columns to rows, supporting parallel processing of multiple rows.

[0208] Cache optimization: Conversion results are cached in vector registers to avoid repeated conversions.

[0209] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0210] The above is a detailed introduction to a database acceleration system and a database acceleration method provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A database acceleration system, characterized in that: include: A central processing unit (10) and a coprocessor (20), wherein the coprocessor (20) comprises at least an internal bus (201) and a vector processor (202); the central processing unit (10) is communicatively connected to the internal bus (201); the internal bus (201) is communicatively connected to the vector processor (202); the vector processor (202) comprises at least a plurality of preset execution units (2021), each preset execution unit (2021) comprising one or more instructions; the plurality of preset execution units (2021) comprising one or more of the following: a data filtering execution unit, a hash connection execution unit, an aggregate calculation execution unit, a regular matching execution unit, a sorting execution unit, a full table scan execution unit, and a plurality of other execution units. Unit and index tree operation execution unit; the data filtering execution unit includes comparison operation instructions; the hash connection execution unit includes hash table construction instructions, hash detection instructions and nested loop connection instructions; the aggregation calculation execution unit includes sum calculation instructions, count calculation instructions, maximum and minimum value instructions and average value instructions; the regular matching execution unit includes pattern compilation instructions, vector matching instructions and result extraction instructions; the sorting execution unit includes radix sort instructions and merge sort instructions; the full table scan execution unit includes initialization scan instructions, batch data acquisition instructions and column projection instructions; the index tree operation execution unit includes initialization index instructions, precise operation instructions and range operation instructions; The central processing unit (10) receives an operation request sent by a user terminal (30), wherein the operation request is an SQL operation request; performs syntax analysis and semantic checking on the operation request to generate a preliminary operation plan; Intervening in the preliminary operation plan through a preset plug-in, rewriting the preliminary operation plan based on rules and cost models, identifying accelerable operation segments and generating an operation plan containing acceleration instructions; wherein the operation plan is the operation steps corresponding to the operation request; corresponding target data is retrieved from the database according to the operation plan; the operation plan and the target data are sent to the internal bus (201); the target data includes hot data and cold data, and the database data includes a plurality of column data; The internal bus (201) sends the operation plan and the target data to the vector processor (202); The vector processor (202) controls one or more preset execution units (2021) to process the target data according to the operation plan to obtain a processing result corresponding to the operation plan; and returns the processing result to the central processing unit (10); The central processing unit (10) outputs the processing result to the user terminal (30); The coprocessor (20) further includes: a memory controller (203); the memory controller (203) is communicatively connected to the internal bus (201); The central processing unit (10) has previously retrieved the database data in the database to the memory controller (203), and the memory controller (203) includes a plurality of storage channels; The central processing unit (10) receives an operation request sent by a user terminal (30); parses the operation request into an operation plan; loads a corresponding operation microcode from a host memory according to the operation plan; and sends the operation microcode to the internal bus (201); wherein the microcode is decomposed into a storage microcode and a calculation microcode; The internal bus (201) sends the operation microcode to the vector processor (202); The vector processor (202) obtains the number of accesses to each column of data in a predetermined dominant access mode according to a preset counter array through the storage microcode; sorts the access numbers to obtain an access sequence; determines the high-frequency access column data in each column of data according to the access sequence; determines the high-frequency access column data as the hot data, and determines the remaining column data in each column of data as the cold data; controls one or more preset execution units (2021) to process the target data through the computing microcode to obtain a processing result corresponding to the operation plan; and returns the processing result to the central processing unit (10).

2. The database acceleration system according to claim 1, characterized in that: The vector processor (202) further includes: a vector register (2022); The vector processor (202) pre-fetches corresponding target data from the database data of the memory controller (203) through the operation microcode; The one or more preset execution units (2021) process the target data according to the operation microcode to obtain a processing result corresponding to the operation plan; and output the processing result to the vector register (2022); The vector register (2022) returns the processing result to the central processing unit (10).

3. The database acceleration system according to claim 1, characterized in that: The coprocessor (20) further includes: a scalar processor (204); The coprocessor (20) verifies the operation microcode to obtain a verification result; The scalar processor (204), after determining that the verification result is verification passed, performs initialization processing on the operation microcode to complete the initialization of the operation microcode.

4. The database acceleration system according to claim 3, characterized in that: The scalar processor (204) further includes: a control register (2041); The control register (2041) activates the operation microcode and establishes a mapping relationship between the instruction operation code and the operation microcode to complete the initialization of the operation microcode.

5. The database acceleration system according to claim 1, characterized in that: The structural data of the preset execution unit (2021) includes one or more of the following: Operation code, function code, sub-function code, destination register, first source register, second source register and mask pattern.

6. The database acceleration system according to claim 5, characterized in that: The structure data of the operation code includes one or more of the following: Arithmetic operation codes, load codes, store codes, branch codes, system instructions and custom instructions.

7. A database acceleration method, characterized in that: The database acceleration system according to claim 1 comprises: The central processing unit receives an operation request sent by a user terminal; parses the operation request into an operation plan; loads corresponding operation microcode from a host memory according to the operation plan; and sends the operation microcode to the internal bus; wherein the microcode is decomposed into a storage microcode and a calculation microcode; The internal bus sends the operation microcode to the vector processor; The vector processor pre-fetches corresponding target data from the database data of the memory controller through the storage microcode; controls one or more preset execution units through the computing microcode to process the target data to obtain a processing result corresponding to the operation plan; and returns the processing result to the central processing unit; The central processing unit outputs the processing result to the user terminal; The target data includes hot data and cold data, and the database data includes a plurality of column data; Accordingly, the pre-fetching corresponding target data from the database data of the memory controller by the storage microcode includes: Obtaining, by means of the storage microcode, the number of accesses to each column of data by a predetermined dominant access mode according to a preset counter array; Sort the number of visits to obtain an access sequence; determining, according to the access sequence, high-frequency access column data among the columns of data; The frequently accessed column data is determined as the hot data, and the remaining column data in the columns of data is determined as the cold data.

8. The database acceleration method according to claim 7, characterized in that: The process of determining the dominant access mode includes: Obtaining the number of times the memory controller is accessed by multiple access modes within a preset time; A dominant access pattern is determined from the plurality of access patterns according to the respective access counts.

9. The database acceleration method according to claim 8, characterized in that: The structure of the access mode includes one or more of the following types: prefetch data size, prefetch trigger condition, and prefetch address cache.

10. The database acceleration method according to claim 7, characterized in that: After determining the frequently accessed column data as the hot data and determining the remaining column data in each column data as the cold data, the method further includes: The hot data is stored in a memory controller in column format, and the starting address, length, and data type of each column of data are recorded; The cold data is stored in a memory controller in row format.

11. The database acceleration method according to claim 7, characterized in that: The processing result includes one or more sub-processing results; Accordingly, controlling the one or more preset execution units to process the target data through the computing microcode to obtain a processing result corresponding to the operation plan includes: The computing microcode is used to operate one or more corresponding preset execution units to process the target data to obtain one or more sub-processing results; Sort the sub-processing results to obtain a processing result sequence; Aggregating the processing result sequence to obtain an aggregated processing result; Merging the processing result sequences to obtain merged processing results; The merged processing result is converted to obtain the processing result.

Citation Information

Patent Citations

  • Accelerator, acceleration method and electronic equipment

    CN114579078A