Chip

By introducing arithmetic logic operation data table module into the chip, replacing or simplifying the function of ALU, the problem of CPU and GPU generating a large amount of heat is solved, and the power consumption is reduced and the computing speed is improved.

CN120144086APending Publication Date: 2025-06-13RUGAO ENQI DAILY NECESSITIES DEPARTMENT STORE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510216676.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In existing computer systems, the large amount of heat generated by CPUs and GPUs lead to high power consumption.

Method used

By introducing arithmetic logical operation data table module into the chip, array tables and indexes are used to store and find operation results, thereby replacing or simplifying the function of the ALU and reducing the power consumption of the chip.

Benefits of technology

The effect of reducing chip power consumption is achieved, while improving calculation speed and reducing heat generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144086A_ABST
    Figure CN120144086A_ABST
Patent Text Reader

Abstract

The invention discloses a chip which comprises an arithmetic logic operation data table module which comprises an array table and indexes, operation results are stored in the array table, and the indexes are in one-to-one correspondence with the operation results; the central processing unit comprises a register and a controller, the controller reads a first operation instruction from the register, and the controller decodes the first operation instruction to obtain a value to be operated; the arithmetic logic operation data table module obtains a to-be-operated value from the controller and takes the to-be-operated value as an index to search an operation result, and the controller writes the obtained operation result back to the register. Parts or all functions of the ALU are replaced by the arithmetic logic operation data table module, the functions of the ALU are simplified or the ALU is removed, and the power consumption of a chip is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of microelectronics technology, and more particularly to a chip. Background Art

[0002] Reference Figure 1 A central processing unit (CPU) generally includes structures such as a control unit (CU), an arithmetic logic unit (ALU), registers, and a cache. The ALU is one of the core components of the CPU and is responsible for performing arithmetic and logical operations in most instructions.

[0003] In the current computer system, storage space is no longer a problem. The problem is the large amount of heat generated by the CPU and GPU. Summary of the Invention

[0004] In view of the above defects or deficiencies in the prior art, it is desirable to provide a chip with reduced power consumption.

[0005] In a first aspect, the chip of the present invention includes:

[0006] An arithmetic logic operation data table module, which includes an array table and an index. The array table stores operation results, and the index corresponds to the operation results one by one;

[0007] A central processing unit, which includes registers and a controller. The controller reads a first operation instruction from the registers, decodes the first operation instruction to obtain a value to be operated, the arithmetic logic operation data table module obtains the value to be operated from the controller and uses the value to be operated as an index to find the operation result, and the controller writes the obtained operation result back to the registers.

[0008] According to the technical solution provided by the embodiments of the present application, by replacing part or all of the functions of the ALU with an arithmetic logic operation data table module, the functions of the ALU are simplified or the ALU is removed, reducing the power consumption of the chip. Brief Description of the Drawings

[0009] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent:

[0010] Figure 1 Schematic diagram of an existing CPU structure;

[0011] Figure 2 Schematic diagram of the structure of the chip according to the embodiment of the present invention;

[0012] Figure 3 Schematic diagram of the structure of the chip according to the embodiment of the present invention;

[0013] Figure 4Schematic diagram of the calculation speed test results of the ALU and the data table;

[0014] Figure 5 Schematic diagram of the calculation speed test results of the ALU and the data table;

[0015] Figure 6 Schematic diagram of the calculation speed test results of the ALU and the data table. Detailed implementation manners

[0016] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the convenience of description, only the parts related to the invention are shown in the drawings.

[0017] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and embodiments.

[0018] Please refer to Figure 2 and 3 , the chip of the present invention is characterized in that it includes: an arithmetic logic operation data table module, which includes an array table and an index, the array table stores operation results, and the index corresponds to the operation results one by one; a central processing unit, which includes a register and a controller, the controller reads a first operation instruction from the register, the controller decodes the first operation instruction to obtain a value to be operated, the arithmetic logic operation data table module obtains the value to be operated from the controller and uses the value to be operated as an index to look up the operation result, and the controller writes the obtained operation result back to the register.

[0019] In the embodiment of the present invention, the arithmetic logic operation data table module includes an array table and an index, and the index corresponds to the operation results one by one. When it is necessary to calculate data, only the value to be operated needs to be used as an index to directly look up the operation result in the data table, and there is no need to perform a calculation process, thereby improving the calculation speed. All functions of the ALU are designed into multiple data tables. Whether it is an 8-bit chip, a 16-bit chip, a 32-bit or 64-bit XOR, or more, its value range is fixed. That is to say, the number of bits of the chip represents the maximum number. Within such a limited data range, we can always establish a table to arrange and combine all the data required for ALU calculation and record the results in the table. In this way, all the requirements for obtaining results through ALU calculation can obtain the results by looking up the table. The array table can be designed in any base.

[0020] The controller reads the first arithmetic instruction from the register, decodes the first arithmetic instruction to obtain the value to be operated on, the arithmetic logic operation data table module obtains the value to be operated on from the controller and uses the value to be operated on as an index to look up the operation result in the data table, and the controller writes the obtained operation result back to the register. The controller (CU) updates the flag bits in the program status register (PSW), and the controller is ready to execute the next instruction.

[0021] The existing ALU is responsible for performing arithmetic and logic operations in most instructions. The array table includes an arithmetic operation table, a logic operation table, a comparison operation table, and a forward and backward shift operation table.

[0022] The arithmetic operation table includes an addition data table, a subtraction data table, a multiplication data table, and a division data table. Specifically, the addition data table add

[256]

[256] , the subtraction data table sub

[256]

[256] , the multiplication data table mul

[256]

[256] , etc., are respectively assigned constant values. For example, add[1][9] represents 1 + 9 and is assigned the value 10; the subtraction data table sub[3][4] represents 3 - 4 and is assigned the value -1; mul[8][9] represents 8 × 9 and is assigned the value 72. Then when the controller (CU) needs the ALU to perform an operation, it can directly obtain the result by taking the array value, thus eliminating the calculation of the ALU.

[0023] The following demonstrates the comparison of the calculation speed of ALU multiplication and the calculation speed of the data table multiplication:

[0024] In this example, the results of addition, subtraction, multiplication, and division for traversing 256 are written to a file in advance and then read into memory in the program of the data table multiplication.

[0025]

[0026]

[0027] The program for demonstrating the multiplication speed of the data table and the calculation speed of the ALU is as follows:

[0028]

[0029]

[0030]

[0031]

[0032] Figures 4 to 6 For the calculation speed test results of the ALU and the data table, from this test result, the multiplication of the data table replaces the multiplication of the ALU; the multiplication of the data table and the multiplication of the ALU are similar in speed under the current CPU architecture conditions.

[0033] The logic operation tables include AND data tables, OR data tables, NOT, XOR data tables, etc. Specifically, corresponding data tables are established, such as AND

[256]

[256] , OR

[256]

[256] , NOT

[256] , XOR

[256]

[256] , etc.

[0034] The comparison operation tables are based on comparison operations, such as >, <, ==, !=, etc. A table is established for each symbol. For example, the greater-than table (Big

[256]

[256] ), and the Boolean values 0 and 1 are written into it.

[0035] The forward and backward shift operation tables perform shift operations. The forward and backward shift operation tables can be established in the form of Fmov

[256] [8][2] and Bmov

[256] [8][2]. 256 represents an 8-bit value; 8 represents the different results of a number that can be shifted 8 times; 2 represents the data shifted out forward or backward, which is used for the subsequent integration of data.

[0036] The arithmetic logic operation data table module is used to partially replace the function of the ALU to simplify the function of the ALU or reduce the usage frequency of the ALU. The arithmetic logic operation data table module completely replaces the function of the ALU to cancel the ALU. Thereby reducing the power consumption of the CPU, that is, reducing the chip power consumption.

[0037] Furthermore, it also includes a data calculation module. If the value to be operated is outside the index range, the data calculation module calculates the value to be operated to obtain a calculation result, and the controller writes the obtained calculation result back to the register.

[0038] In the embodiment of the present invention, when processing out-of-table data, the data calculation module calculates the value to be operated to obtain a calculation result.

[0039] Here, it is necessary to consider how to split a large integer into bytes and perform the above operations.

[0040] The union in C language perfectly implements this function. A union is a special data type that allows different data types to be stored in the same memory location. By defining a union that contains an integer member and a byte array member, the splitting of an integer into bytes can be achieved. As follows:

[0041] #include<stdio.h>

[0042] / / Define a union for splitting an integer into bytes

[0043] union IntToBytes{

[0044] int integer; / / Integer type member

[0045] unsigned char bytes[sizeof(int)]; / / Character array member for storing bytes

[0046] };

[0047] int main(){

[0048] union IntToBytes converter;

[0049] / / Assign a value to the integer member

[0050] converter.integer = 0x12345678;

[0051] / / Print the value of the integer

[0052] printf("Value of the integer: 0x%08X\n", converter.integer);

[0053] / / Print each byte of the integer

[0054] for(int i = 0; i < sizeof(int); i++){

[0055] printf("Value of the %d-th byte: 0x%02X\n", i, converter.bytes[i]);

[0056] }

[0057] return 0;

[0058] }

[0059] At the same time, a union can also be used to combine bytes into an integer. As follows:

[0060] #include <stdio.h>

[0061] / / Define the union

[0062] union ByteToInt{

[0063] unsigned int integer; / / Integer type member

[0064] unsigned char bytes[4]; / / Character array member, 4 bytes corresponding to a 32-bit integer

[0065] };

[0066] int main(){

[0067] union ByteToInt converter;

[0068] / / Assume the byte data to be merged

[0069] converter.bytes[0] = 0x01;

[0070] converter.bytes[1] = 0x02;

[0071] converter.bytes[2] = 0x03;

[0072] converter.bytes[3] = 0x04;

[0073] / / Print the merged integer value

[0074] printf("Merged integer: 0x%08X\n", converter.integer);

[0075] return 0;

[0076] }

[0077] After splitting data of different number systems into 256 - number - system data, perform look - up table calculations on the split 256 - number - system data.

[0078] It can be but is not limited to, the high - precision addition calculation process is as follows:

[0079] Reverse storage: Store the string in reverse order into the array starting from the units digit (e.g., "123" is stored as [3, 2, 1]), which is convenient for aligning operations from the low - order bits;

[0080] Add digit by digit: Each calculation is c[i] = a[i] + b[i] + carry, the carry carry = c[i] / 10, and the current digit retains c[i] % 10;

[0081] Process the carry of the highest digit: If there is still a carry after adding the highest digit, the length of the result needs to be extended;

[0082] Remove leading zeros: Redundant zeros at the end of the result array need to be deleted before outputting in reverse order (e.g., [0, 3, 2, 1] is converted to [3, 2, 1]).

[0083] It can be but is not limited to, the high - precision subtraction calculation process is as follows:

[0084] Compare sizes: If the minuend is less than the subtrahend, swap the two and mark the result as negative;

[0085] Borrow handling: If the current digit a[i] < b[i], borrow 1 from the higher digit (a[i] += 10, a[i + 1] -= 1);

[0086] Result correction: If all higher digits are zero, at least one significant digit needs to be retained (e.g., if the result is 0).

[0087] It can be but is not limited to, the high-precision multiplication calculation process is as follows:

[0088] Rule of multiplying each digit: The result of a[i] * b[j] is accumulated to c[i + j - 1] (simulating vertical multiplication);

[0089] Carry handling: After the outer loop ends, calculate the carry of each digit c[i] uniformly (c[i + 1] += c[i] / 10, c[i] %= 10);

[0090] Number of digits of the result: The maximum number of digits of the product is the sum of the number of digits of the two numbers (e.g., 999 * 999 = 998001, number of digits 3 + 3 = 6).

[0091] It can be but is not limited to, the high-precision division calculation process is as follows:

[0092] Trial quotient strategy: By expanding the divisor multiple (such as multiplying by 10^n) to reduce the number of subtraction operations and improve efficiency;

[0093] Subtraction simulation: Use high-precision subtraction to subtract the expanded divisor in a loop, and record the number of times as the quotient of the current digit;

[0094] Remainder handling: The remainder after each subtraction needs to be retained and combined with the next digit for continued operation;

[0095] Complexity optimization: Through bit weight adjustment (such as divisor × 10^k), reduce the time complexity from O(N) to O(logN).

[0096] Furthermore, the array table is in base 256.

[0097] In the embodiments of the present invention, 256 is the 8th power of 2 (i.e., the range of 1 byte), each element occupies 1 byte, and the memory is naturally aligned. Modern CPUs use bytes as the smallest addressing unit. After alignment, the access speed is faster, and memory fragmentation is reduced. The data table stored in consecutive bytes can better utilize the CPU cache line (usually 64 bytes), reduce cache misses, and improve performance.

[0098] Furthermore, the array table includes an arithmetic operation table, a logical operation table, a comparison operation table, and a forward and backward shift operation table.

[0099] Furthermore, the arithmetic and logical operation data table module is stored in the hard disk or the cache inside the central processing unit.

[0100] In an embodiment of the present invention, the arithmetic logic operation data table module can be stored in a hard disk, which communicates with a central processing unit and reads data into the memory. The arithmetic logic operation data table module can be stored in the cache inside the central processing unit to improve the data processing speed.

[0101] Furthermore, the central processing unit further includes a plurality of auxiliary controllers, which are used for parallel operation of data. Adding a plurality of auxiliary controllers enables efficient parallelism in data table calculations, thereby improving the parallel processing ability and data processing speed.

[0102] The calculation method of the data table is more efficient and faster. The characteristic of the ALU is that it inputs data and obtains the result through a circuit. The characteristic of the data table is that it combines known results to calculate another unknown result. The results are the same, but the data table method may be faster. For example, for trigonometric functions, the ALU calculation method obtains data step by step through a circuit (or a coprocessor); while the data table calculation method can not only imitate the ALU calculation method to calculate data step by step, but also establish a trigonometric function data table and directly look up the table to obtain the data.

[0103] Once the ALU is designed, its operation set is fixed. All the functions of the ALU are striving to implement existing human theories. The calculation method of the data table can almost be realized as long as one dares to think and the theory is mature.

[0104] When the ALU executes complex arithmetic and logic operations, especially when dealing with complex mathematical operations or graphics processing, a large amount of heat is generated. Removing the ALU directly reduces the power consumption of the chip. On the same chip area, it is beneficial to expand the scale of the controller, register, or on-chip cache, and several auxiliary controllers can be added for efficient parallelism in data table calculations, etc.

[0105] The characteristic of the data table calculation method is that multiple data tables need to be established, occupying memory space. Depending on the use of the data table, it is at least several hundred K, which can be ignored considering today's memory measured in G and hard disk measured in T. The system that removes the ALU and replaces it with the data table calculation method is cheaper.

[0106] The above description is only the preferred embodiment of the present application and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but also should cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with (but not limited to) the technical features with similar functions disclosed in the present application.

Claims

1. A chip, characterized in that: include: An arithmetic logic operation data table module, comprising an array table and an index, wherein the array table stores operation results, and the index corresponds to the operation results one by one; A central processing unit includes a register and a controller, wherein the controller reads a first operation instruction from the register, the controller decodes the first operation instruction to obtain a value to be operated, the arithmetic logic operation data table module obtains the value to be operated from the controller and uses the value to be operated as an index to find an operation result, and the controller writes the obtained operation result back to the register.

2. The chip according to claim 1, characterized in that: It also includes a data calculation module. If the value to be calculated is outside the range of the index, the data calculation module calculates the value to be calculated to obtain a calculation result, and the controller writes the obtained calculation result back to the register.

3. The chip according to claim 1, characterized in that: The array table is in 256-base.

4. The chip according to claim 1, characterized in that: The array table includes an arithmetic operation table, a logic operation table, a comparison operation table and a forward and backward shift operation table.

5. The chip according to claim 1, characterized in that: The arithmetic logic operation data table module is stored in a hard disk or a high-speed cache inside the central processing unit.

6. The chip according to claim 1, characterized in that: The central processing unit also includes a plurality of auxiliary controllers, and the plurality of auxiliary controllers are used for parallel operation of data.