Operating system kernel mode calculation method and device, equipment, medium and product
By generating lookup table sets and performing bit string decomposition, this method replaces traditional GEMM operations, achieving efficient parallel computing in the operating system kernel mode. This solves the problem of low computational efficiency in traditional methods and improves resource utilization and computational efficiency.
Patent Information
- Application Number
- CN202510777726.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-11-07
AI Technical Summary
In traditional operating system kernel-mode computation methods, the low computational density and poor energy efficiency of floating-point operations conflict with the kernel's requirements for high real-time performance and low latency. Furthermore, the GEMM computation method does not fully utilize hardware characteristics, resulting in resource constraints and low computational efficiency.
By obtaining the input matrix and weight matrix, pre-computation processing is performed to generate a lookup table set. The weight matrix is then decomposed into a bit string. Matrix multiplication is performed using the target index value and the lookup table set, replacing the traditional element-by-element multiplication operation. Parallel table lookup is then performed using SIMD instructions.
It enables efficient inference computation with low computational resource consumption, significantly reducing instruction cycles and power consumption, and improving computational efficiency and resource utilization.
Smart Images

Figure CN120909755A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a method and device for kernel computing of an operating system, equipment, medium and product. BACKGROUND
[0002] With the increasing complexity of modern OS (Operating System), the rapid diversification of hardware, and the steady development of artificial intelligence, it is urgent to explore the potential of AI (Artificial Intelligence) in improving the decision-making of the OS kernel.
[0003] In the traditional method for kernel computing of an operating system, the deployment of an AI model in a kernel environment faces two key bottlenecks, making it difficult for the AI model to meet the deployment requirements of the kernel environment. First, the low computational density and poor energy efficiency of floating-point operations are in conflict with the strict requirements of the kernel for high real-time performance and low latency. Second, the traditional GEMM (General Matrix Multiply) computing method does not fully utilize hardware characteristics, and occupies a large amount of processor and memory resources, further exacerbating the resource shortage of the kernel environment. Therefore, the current method for kernel computing of an operating system has low computational efficiency. SUMMARY
[0004] Therefore, it is necessary to provide a method, device, equipment, medium and product for kernel computing of an operating system that can improve computational efficiency.
[0005] In a first aspect, the present application provides a method for kernel computing of an operating system, comprising:
[0006] obtaining an input matrix and a weight matrix;
[0007] performing pre-computation processing on the input matrix to obtain a set of lookup tables;
[0008] performing bit string decomposition processing on the weight matrix to determine a target index value;
[0009] using the target index value and the set of lookup tables, performing computation processing on the input matrix according to the corresponding positions of matrix multiplication to obtain a target computation result.
[0010] In one embodiment, the pre-computation processing on the input matrix to obtain the set of lookup tables comprises:
[0011] grouping the input matrix according to a preset grouping dimension to obtain an activation sequence corresponding to each lookup table, the size of the activation sequence being the grouping dimension, and the values in the activation sequence being the values of the input matrix in the group;
[0012] traversing each activation sequence to obtain a lookup table;
[0013] According to each of the lookup tables, the set of lookup tables is obtained.
[0014] In one of the embodiments, traversing each activation sequence to obtain a lookup table includes:
[0015] For each of the activation sequences, a lookup calculation formula corresponding to the index value is obtained by using a preset binary form of the index value;
[0016] According to the elements in the activation sequence, the lookup calculation formula is calculated and processed to obtain a lookup value corresponding to the index value;
[0017] According to each of the lookup values, the lookup table is composed.
[0018] In one of the embodiments, the weight matrix is subjected to bit string decomposition processing to determine a target index value, including:
[0019] The weight matrix is subjected to decomposition processing according to bit positions to obtain a plurality of query matrices, the number of the query matrices being consistent with the number of the bit positions;
[0020] Based on each of the query matrices, a target index value is determined.
[0021] In one of the embodiments, using the target index value and the set of lookup tables, an input matrix is calculated and processed according to a corresponding position of matrix multiplication to obtain a target calculation result, including:
[0022] According to the corresponding position of matrix multiplication, a target lookup table is determined from the set of lookup tables;
[0023] Based on each of the target index values, a target lookup value is determined from the target lookup table;
[0024] Using a preset calculation rule, each of the target lookup values is calculated and processed to obtain a target calculation result.
[0025] In one of the embodiments, the method further includes:
[0026] For each of the query matrices, the query matrix is divided according to bit positions to obtain a bit position matrix corresponding to each bit position;
[0027] Each of the bit position matrices is stored in an integer form.
[0028] In a second aspect, the application further provides an operating system kernel state computing device, including:
[0029] A data acquisition module is configured to acquire an input matrix and a weight matrix;
[0030] a precomputation module, configured to perform precomputation processing on the input matrix to obtain a lookup table set;
[0031] a matrix decomposition module, configured to perform bit string decomposition processing on the weight matrix to determine a target index value;
[0032] a result calculation module, configured to perform calculation processing on the input matrix according to the corresponding positions of matrix multiplication by using the target index value and the lookup table set, to obtain a target calculation result.
[0033] In a third aspect, the present application also provides a computer device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the method of the first aspect when executing the computer program.
[0034] In a fourth aspect, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect.
[0035] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the method of the first aspect.
[0036] The above-mentioned operating system kernel state computing method, device, equipment, medium and product, wherein the operating system kernel state computing method of the first aspect comprises: obtaining an input matrix and a weight matrix; performing precomputation processing on the input matrix to obtain a lookup table set; performing bit string decomposition processing on the weight matrix to determine a target index value; performing calculation processing on the input matrix according to the corresponding positions of matrix multiplication by using the target index value and the lookup table set, to obtain a target calculation result. In the present application, the lookup table set is constructed according to the input matrix, the traditional element-by-element multiplication operation of GEMM is replaced by the lookup table method, and parallel lookup table based on the SIMD instruction can be performed, so that efficient inference calculation can be completed under low computing resource occupation. In the AI model of low bit width integer quantization type, the method in the present embodiment can greatly reduce the number of instruction cycles and power consumption under the condition of lacking special hardware support, and greatly improve the calculation efficiency and resource utilization. BRIEF DESCRIPTION OF DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other related drawings can also be obtained without creative labor.
[0038] Figure 1 It is a flowchart of the operating system kernel state computing method in an embodiment.
[0039] Figure 2 Flowchart of steps for obtaining a lookup table set in one embodiment;
[0040] Figure 3 Structural diagram of a lookup table set in one embodiment;
[0041] Figure 4 Structural diagram of a plurality of lookup tables in one embodiment;
[0042] Figure 5 Flowchart of steps for determining a target index value in one embodiment;
[0043] Figure 6 Flowchart of steps for obtaining a target calculation result in one embodiment;
[0044] Figure 7 Flowchart of steps for compressively storing a weight matrix in one embodiment;
[0045] Figure 8 Flowchart of an operating system kernel mode calculation method in another embodiment;
[0046] Figure 9 Structural block diagram of an operating system kernel mode calculation device in one embodiment;
[0047] Figure 10 Internal structural diagram of a computer device in one embodiment. DETAILED DESCRIPTION
[0048] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0049] It should be noted that the terms "first", "second", etc. used in the present application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "include" and "have" and any variations thereof used in the present application are intended to cover non-exclusive inclusion. The term "a plurality of" used in the present application refers to two or more. The term "and / or" used in the present application refers to one of the options or any combination of a plurality of options.
[0050] The operating system kernel state computing method provided by the embodiments of the present application can be applied to a processor or an acceleration chip with a SIMD (Single Instruction, Multiple Data) instruction set. The processor can be an x86 architecture, and a lookup instruction in an AVX (Advanced Vector Extensions) instruction set is used for calculation. Such a processor can be a central processor in a server. The processor can also be an ARM (Advanced RISC Machine) architecture, and a NEON instruction set is used to implement efficient lookup calculation operation. In this case, the processor can be a processor in a mobile phone, an embedded system, a Raspberry Pi, or an NPU host.
[0051] In one embodiment, as shown in Figure 1 An operating system kernel state computing method is provided. The embodiments take the method applied to a server as an example. It can be understood that the method can also be applied to a terminal and can also be applied to a system including a terminal and a server and can be implemented through interaction between the terminal and the server. In the embodiments, the method includes the following steps.
[0052] In step 102, an input matrix and a weight matrix are obtained.
[0053] In the embodiments of the present application, the input matrix can be input data of one layer in an AI model, such as a neural network. For example, in an image processing task, the input matrix can be a pixel block of an image or a feature map in a convolutional neural network. In a speech recognition task, the input matrix can be a mel-frequency coefficient or other acoustic feature representation. In a natural language processing task, the input matrix can be an embedded and encoded word vector, a position vector, or a context representation. In an Internet of Things and industrial monitoring scenario, the input matrix can be multi-dimensional sensor sampling data, such as acceleration, temperature, and voltage. In a recommendation system, the input matrix can be user behavior encoding or commodity feature vectors. The input matrix X ∈ R N×K can be represented as
[0054] ,
[0055] The weight matrix in the embodiments of the present application can be obtained by training according to a specific AI model and task requirement. The weight matrix is usually in a low-bit representation form, such as a 1-bit, 2-bit, or 4-bit compression format, and is used to represent coefficient information of connection strength in each layer of the model. The weight matrix W ∈ R M×K can be represented as
[0056] .
[0057] Step 104, pre-computing the input matrix to obtain a lookup table set.
[0058] Step 106, bit string decomposition processing is performed on the weight matrix to determine the target index value.
[0059] The string decomposition processing refers to splitting the weight matrix into a 1-bit matrix according to the bit, and then determining the target index value based on the array corresponding to the 1-bit matrix.
[0060] Step 108, using the target index value and the lookup table set, according to the corresponding position of the matrix multiplication, the input matrix is calculated to obtain the target calculation result.
[0061] The target lookup value can be determined from the lookup table set based on the target index value, and the target calculation result is obtained by calculating each target lookup value using a preset calculation rule. The calculation rule corresponds to the pre-computation rule of the input matrix.
[0062] The above operating system kernel state calculation method obtains the input matrix and the weight matrix; pre-computing the input matrix to obtain a lookup table set; bit string decomposition processing is performed on the weight matrix to determine the target index value; using the target index value and the lookup table set, according to the corresponding position of the matrix multiplication, the input matrix is calculated to obtain the target calculation result. In this embodiment, the lookup table set is constructed according to the input matrix, and the traditional GEMM (General Matrix Multiply, general matrix multiplication) element-by-element multiplication operation is replaced by the lookup table method, and parallel lookup table based on SIMD instruction can be performed, and efficient inference calculation can be completed under low calculation resource occupation. In the low bit width integer quantization type AI model, the method in this embodiment can greatly reduce the number of instruction cycles and power consumption under the condition of lacking special hardware support, greatly improving the calculation efficiency and resource utilization.
[0063] In one exemplary embodiment, based on Figure 1 As shown in the embodiment shown in Figure 2 According to the input matrix, the pre-computation process to obtain the lookup table set includes:
[0064] Step 202, grouping the input matrix according to the preset grouping dimension to obtain the activation sequence corresponding to each lookup table, the size of the activation sequence is the grouping dimension, and the value in the activation sequence is the value of the input matrix in the group.
[0065] Step 204, traversing each activation sequence to obtain a lookup table.
[0066] Step 206, according to each lookup table, the lookup table set is obtained.
[0067] In some embodiments, step 204 can further include: for each of the activation sequences, obtaining a lookup calculation formula corresponding to the index value by using a preset binary form of the index value; calculating the lookup calculation formula according to the elements in the activation sequence to obtain a lookup value corresponding to the index value; and composing the lookup table according to the lookup values.
[0068] wherein the grouping dimension g represents that after grouping the input matrix X, g elements are included in each group. As shown in Figure 3 For an input matrix X of N×K and a grouping dimension g, the size of each lookup table is 2 g , and the number of lookup tables is N×K / g.
[0069] For example, for an input matrix X of 32 elements, g=4, 4 elements in each group, the size of each lookup table is 16 rows, i.e., each lookup table includes 16 index values, the value range of the index value is [0, 15], and the number of lookup tables is 8.
[0070] As shown in Figure 1 Take the first lookup table as an example for illustration, and combine the foregoing embodiments. The input matrix X is divided into groups of 4 elements, the first group is (a, b, c, d), which is the activation sequence corresponding to the first lookup table. In the 16 rows of the first lookup table, "0" in the 4-bit binary form of the index value indicates that the corresponding position in the lookup value calculation formula is "-"; "1" in the binary form indicates that the corresponding position in the lookup value calculation formula is "+". The index value of the first row is 0, the binary form is "0000", and the lookup value calculation formula is "-d-c-b-a"; the index value is 7, the binary form is "0111", and the lookup value calculation formula is "-d+c+b+a".
[0071] It can be understood that the second group of the input matrix X can be (e, f, g, h), which is the activation sequence corresponding to the second lookup table. In the 16 rows of the second lookup table, the index value of the first row is 0, the binary form is "0000", and the lookup value calculation formula is "-h-g-f-e"; the index value is 8, the binary form is "1000", and the lookup value calculation formula is "+h-d-f-e".
[0072] In a possible implementation, each lookup table includes a calculation lookup value and a mirror lookup value; in the process of generating the lookup table, further including: taking, among the index values, index values with the same highest bit in binary form as calculation index values, and taking the rest of the index values as mirror index values; performing calculation processing on the lookup values corresponding to the calculation index values by using the activation sequence to obtain the calculation lookup values; performing inversion mirror processing on each of the calculation index values and the activation sequence to obtain mirror lookup values corresponding to each of the mirror index values; and sorting the calculation lookup values and the mirror lookup values according to the sizes of the index values to obtain the lookup table.
[0073] For 4-bit binary data form index values, the calculation index values can be 8 smaller ones with the highest bit being 0, i.e., 0-7, in binary form 0000-0111, or 8 larger ones with the highest bit being 1, i.e., 8-15, in binary form 1000-1111. It can be understood that the lookup value calculation formula "-d-c-b-a" corresponding to the index value 0 and the lookup value calculation formula "+d+c+b+a" corresponding to the index value 15 are opposite numbers, i.e., the index value 0 and the index value 15 are mirror images, and the corresponding lookup values are also mirror images. After obtaining the lookup value corresponding to one of the index values 0 or 15, the lookup value corresponding to the other index value can be obtained through mirror processing. The mirror processing can be performed through the following instructions to obtain the mirror lookup value corresponding to the mirror index value:
[0074] vec_lut[t]=-vec_lut[15-t],
[0075] wherein t represents the mirror index value, 15-t represents the calculation index value corresponding to the mirror index value, the instruction represents accessing the lookup value of the calculation index value 15-t, taking the negative of the lookup value, and storing the negative value in the position of the mirror index value t as the lookup value of the mirror index value t.
[0076] The lookup tables in the final lookup table set can be arranged in order of the index values from small to large.
[0077] For example, the lookup value calculation formula "+d-c-b-a" corresponding to the calculation index value 8 can be obtained by fixing the highest bit of the index value in binary form to 1, i.e., the first bit of the activation sequence in the lookup value calculation formula is +. Based on the above instructions, the acquisition process of the mirror index value 7 corresponding to the calculation index value 8 can be "-(+d-c-b-a)".
[0078] In this embodiment, each lookup table is generated through mirror processing, only half of the lookup values need to be calculated, which can effectively reduce the memory access and calculation amount, and can control the index values in the generated lookup table to be continuously accessible and ensure the uniformity of the access logic.
[0079] In a possible implementation, the input matrix can be pre-processed by using AVX2 instructions such as the _mm256_add_ps and _mm256_sub_ps instructions to obtain a set of lookup tables.
[0080] For example, the input matrix can be a 1x32 32-bit floating point number, g=4, and for 8 lookup tables, the above instructions can be used to perform parallel calculation and processing on each lookup table, and 8 lookup tables are constructed to form a set of lookup tables. Specifically, as shown in Figure 4 The same color represents the same lookup table, Figure 4 The outermost part is the first row of the 8 lookup tables, and for the first row index value "0000" of the 8 lookup tables, the corresponding lookup value calculation formula is the subtraction of the 4 values in the activation sequence. The _mm256_sub_ps instruction can be used to simultaneously process the first row of the 8 lookup tables. For example, for the first lookup table, the lookup value calculation formula is "-d-c-b-a"; and for the second lookup table, the lookup value calculation formula is "-h-g-f-e", thereby quickly completing the pre-calculation and generating the set of lookup tables.
[0081] In an example embodiment, based on Figure 2 As shown in Figure 5 The process of determining the target index value includes:
[0082] In step 502, the weight matrix is decomposed by bit to obtain a plurality of query matrices, and the number of query matrices is consistent with the number of bit positions.
[0083] In step 504, the target index value is determined based on each query matrix.
[0084] In combination with the foregoing embodiments, the weight matrix W is an n-bit matrix, which is decomposed by bit into n 1-bit query matrices. The i-th query matrix W (i) ∈R M×K / g The query matrix weight parameter is {-1, 1}.
[0085] For example, one of the decomposed query matrices W (i) may be represented as:
[0086] ,
[0087] The bits of the target index value are w0, w1, w2, and w3, respectively, and the above index value in the target lookup table is taken as the target index value.
[0088] In a possible implementation, the process of performing calculation processing on the input matrix according to the corresponding positions of the matrix multiplication by using the target index values and the lookup table set can further include: determining a target lookup table from the lookup table set according to the corresponding positions of the matrix multiplication; determining target lookup values from the target lookup table based on the target index values; and performing calculation processing on the target lookup values by using a preset calculation rule to obtain the target calculation result.
[0089] The preset calculation rule can include bit shifting and result correction, and the bit shifting result may be represented as:
[0090]
[0091] For example, the calculation process is as shown in Figure 6 The weight matrix W is an INT4 matrix and can be represented as:
[0092] ,
[0093] According to the preset decomposition rule, the weight matrix W can be decomposed into the following query matrix:
[0094] ,
[0095] In the target lookup table corresponding to the query matrix, the query matrix includes the index values "0101", "0011", "1111" and "0000" in the target lookup table, and the corresponding lookup values are:
[0096] R (0) = 2; R (1) = 4; R (2) = 10; R (2) = -10.
[0097] The target calculation result R can be represented as:
[0098] R = 0.5R (0) + R (1) + 2R (2) + 4R (2) - 0.5Sum(X) = 1 + 4 + 20 - 40 - 5 = -20.
[0099] In a possible implementation, the process of performing calculation processing on the input data based on the lookup table to obtain the target calculation result can be completed during online.
[0100] In an example embodiment, the method further includes: for each query matrix, dividing the query matrix by bit to obtain a bit matrix corresponding to each bit; and compressing and storing each bit matrix in integer form.
[0101] As shown in Figure 7 , for the weight matrix W (32x16, INT2), each INT2 can be represented by 2 bits. The weight matrix decomposition forms a query matrix as shown in Figure 4 . The low bit matrix is formed by the low bits of each weight value in each query matrix, as shown in Figure 4 ; and the high bit matrix is formed by the high bits of each weight value, as shown in Figure 4 . The high bit matrix and the low bit matrix are packed and stored in int8 format.
[0102] In a possible implementation, the operation of compressed storage can be completed offline.
[0103] In this embodiment, 4 int2 are stored as 1 int8, avoiding the problem of storing a 1-bit matrix in the original bit width and occupying a large amount of memory, thereby effectively saving storage space.
[0104] In an exemplary embodiment, as shown in Figure 8 , a method for operating a system kernel state is provided, comprising:
[0105] Step 801, obtaining an input matrix and a weight matrix.
[0106] Step 802, grouping the input matrix according to a preset grouping dimension to obtain an activation sequence corresponding to each lookup table.
[0107] Wherein, the size of the activation sequence is the grouping dimension, and the value in the activation sequence is the value of the input matrix in the group.
[0108] Step 803, traversing each activation sequence to form a lookup table.
[0109] Wherein, for each activation sequence, a lookup calculation formula corresponding to the index value is obtained by using the binary form of a preset index value; according to the elements in the activation sequence, the lookup calculation formula is calculated and processed to obtain a lookup value corresponding to the index value; and according to each lookup value, the lookup table is formed.
[0110] Step 804, obtaining a lookup table set according to each lookup table.
[0111] Step 805, performing bit string decomposition processing on the weight matrix to determine a target index value.
[0112] Step 806, determining a target lookup table from the lookup table set according to the corresponding position of the matrix multiplication.
[0113] Step 807, determining a target lookup value from the target lookup table based on each target index value.
[0114] Step 808, performing calculation processing on each target lookup value by using a preset calculation rule to obtain a target calculation result.
[0115] The target index value is determined by decomposing the weight matrix by bit strings, including: decomposing the weight matrix by bit to obtain a plurality of query matrices, the number of query matrices being consistent with the number of bit positions; and determining the target index value based on each query matrix.
[0116] For each query matrix, the query matrix is divided by bit to obtain a bit matrix corresponding to each bit position; and each bit matrix is stored in integer form.
[0117] It should be understood that although each step in the flowchart involved in the above embodiments is displayed in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise stated herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in the above embodiments can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps. It can be understood that the steps in different embodiments can be freely combined as needed, and various non-contradictory schemes formed by combination are within the scope of protection of the present application.
[0118] Based on the same inventive concept, the embodiments of the present application also provide an operating system kernel state computing device for implementing the above-mentioned operating system kernel state computing method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more operating system kernel state computing device embodiments provided below can refer to the limitations of the operating system kernel state computing method in the above text, which will not be repeated here.
[0119] In an exemplary embodiment, as shown in Figure 9 An operating system kernel state computing device is provided, including: a data acquisition module 902, a pre-computation module 904, a matrix decomposition module 906, and a result calculation module 908, wherein:
[0120] The data acquisition module 902 is configured to acquire an input matrix and a weight matrix.
[0121] The precomputation module 904 is configured to perform precomputation on the input matrix to obtain a set of lookup tables.
[0122] The matrix decomposition module 906 is configured to perform bit string decomposition on the weight matrix to determine a target index value.
[0123] The result computation module 908 is configured to perform computation on the input matrix according to the corresponding positions of matrix multiplication by using the target index value and the set of lookup tables to obtain a target computation result.
[0124] In one of the embodiments, the precomputation module 904 is further configured to group the input matrix according to a preset grouping dimension to obtain an activation sequence corresponding to each of the lookup tables, the activation sequence has a size of the grouping dimension, and values in the activation sequence are values of the input matrix in a group; perform table generation by traversing each of the activation sequences to obtain the lookup tables; and obtain the set of lookup tables according to the lookup tables.
[0125] In one of the embodiments, the precomputation module 904 is further configured to, for each of the activation sequences, obtain a lookup computation formula corresponding to the index value by using a preset binary form of the index value; perform computation on the lookup computation formula according to elements in the activation sequence to obtain a lookup value corresponding to the index value; and compose the lookup table according to the lookup values.
[0126] In one of the embodiments, the matrix decomposition module 906 is further configured to perform decomposition on the weight matrix according to bit positions to obtain a plurality of query matrices, the number of the query matrices is consistent with the number of the bit positions.
[0127] The target index value is determined based on the query matrices.
[0128] In one of the embodiments, the result computation module 908 is further configured to determine a target lookup table from the set of lookup tables according to the corresponding positions of matrix multiplication; determine target lookup values from the set of lookup tables based on the target index values; and perform computation on the target lookup values according to a preset computation rule to obtain the target computation result.
[0129] In one of the embodiments, the data acquisition module 902 is further configured to, for each of the query matrices, divide the query matrix according to bit positions to obtain a bit matrix corresponding to each of the bit positions; and store the bit matrices in an integer form.
[0130] Each of the modules in the operating system kernel state computing device described above can be realized by software, hardware, or a combination thereof. Each of the modules described above can be embedded in or independent of a processor in a computer device in a hardware form, or stored in a memory in a computer device in a software form, so as to be called and executed by a processor to perform operations corresponding to each of the modules.
[0131] In an example embodiment, a computer device, which can be a server, is provided, and an internal structure diagram of the computer device can be as shown in FIG. 1. Figure 10 The computer device includes a processor, a memory, an input / output interface, and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store an input matrix and a weight matrix. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with terminals outside through a network connection. The computer program is executed by the processor to implement an operating system kernel state computing method.
[0132] Those skilled in the art can understand that Figure 10 The structure shown in FIG. 1 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0133] In an example embodiment, a computer device is also provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0134] In an example embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0135] In an example embodiment, a computer program product is provided, which includes a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0136] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the above-mentioned embodiment methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., but is not limited thereto. The processor involved in the embodiments provided in the present application can be a general processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but is not limited thereto.
[0137] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the range disclosed in the present application.
[0138] The above embodiments only express several implementation ways of the present application, and the description is specific and detailed, but it should not be understood as a limitation to the patent scope of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, several modifications and improvements can be made, which all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. An operating system kernel mode computing method, characterized in that, The operating system kernel state computing method comprises: An input matrix and a weight matrix are obtained; The input matrix is pre-computed to obtain a lookup table set; The weight matrix is bit string decomposed to determine a target index value; The input matrix is computed according to the target index value and the lookup table set and a corresponding position of matrix multiplication to obtain a target computation result.
2. The method of claim 1, wherein, The input matrix is pre-computed to obtain a lookup table set, which comprises: The input matrix is grouped according to a preset grouping dimension to obtain an activation sequence corresponding to each lookup table, the size of the activation sequence being the grouping dimension, and the values in the activation sequence being the values of the input matrix in the group; Each activation sequence is iterated to obtain a lookup table; The lookup table set is obtained according to each lookup table.
3. The method of claim 2, wherein, Each activation sequence is iterated to obtain a lookup table, which comprises: For each activation sequence, a lookup computation formula corresponding to the index value is obtained by using a preset binary form of the index value; The lookup computation formula is computed according to the elements in the activation sequence to obtain a lookup value corresponding to the index value; The lookup table is composed according to each lookup value.
4. The method of claim 1, wherein, The weight matrix is bit string decomposed to determine a target index value, which comprises: The weight matrix is decomposed according to bits to obtain a plurality of query matrices, the number of the query matrices being consistent with the number of the bits; The target index value is determined based on each query matrix.
5. The method of claim 4, wherein, The input matrix is computed according to the target index value and the lookup table set and a corresponding position of matrix multiplication to obtain a target computation result, which comprises: A target lookup table is determined from the lookup table set according to the corresponding position of matrix multiplication; A target lookup value is determined from the target lookup table based on each target index value; Each target lookup value is computed according to a preset computation rule to obtain the target computation result.
6. The method of claim 4, wherein, The method further comprises: For each query matrix, the query matrix is divided according to bits to obtain a bit matrix corresponding to each bit; Each bit matrix is stored in an integer form.
7. An operating system kernel mode computing device, comprising: The operating system kernel state computing device comprises: A data acquisition module for obtaining an input matrix and a weight matrix; A pre-computation module for pre-computing the input matrix to obtain a lookup table set; A matrix decomposition module for bit string decomposing the weight matrix to determine a target index value; A result computation module for computing the input matrix according to the target index value and the lookup table set and a corresponding position of matrix multiplication to obtain a target computation result.
8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the steps of the method of any one of claims 1 to 6.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.