Vector register device, vector register addressing method and electronic device

By designing a vector register device that supports direct addressing and indirect addressing, the problem of data arrangement in the vector processor is solved, and the interleaving arrangement of data and the optimization of matrix row-column vector simultaneous addressing is realized.

CN119847945BActive Publication Date: 2025-06-06JIXIN COMM TECH (NANJING) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510325221.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-06
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

In the prior art, the non-parallelity of data arrangement of vector processors is poor, and it is impossible to effectively solve the problem of non-parallel data arrangement, especially when matrix row vectors are addressed simultaneously.

Method used

A vector register device is designed, including a decoding unit, an address vector unit and a vector register unit, supporting two methods: direct addressing and indirect addressing. Through the setting of address vector units and vector register units, any vector units in the vector register can be flexibly read, realizing the interleaving arrangement of data.

Benefits of technology

It solves the problem of data arrangement not parallel in vector processors, especially the problem of simultaneous addressing of matrix row and column vectors, optimizes the data arrangement in vector processors, and improves the processor's computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119847945B_ABST
    Figure CN119847945B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of register technology, and provides a vector register device, a vector register addressing method and an electronic device. The vector register device comprises: a decoding unit translates register address coding to obtain a translation address, and sends it to an address vector unit or a vector register unit; the address vector unit determines a target address register according to the translation address, reads data in the target address register, and sends it to the vector register unit; the vector register unit reads data in a target vector register according to the data in the target address register, or reads data in a target vector register according to the translation address. The present invention has two modes of direct addressing and indirect addressing through the setting of an address vector unit and a vector register unit. Indirect addressing is performed through a group of address registers to realize reading any vector unit within the addressing range, and solves the problem of non-parallel data arrangement in a vector processor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of register technology, and in particular to a vector register device, a vector register addressing method and an electronic device. Background Art

[0002] The original general-purpose processors were all scalar processors, that is, one instruction performed an independent operation and obtained a data result. In the Flynn classification, such processors are also called single instruction single data (SISD). As the demand for data computing increases, an obvious method is to perform the same operation on a set of data, so that the parallel processing capability of data is increased without changing the instruction stream, that is, single instruction multiple data (SIMD). SIMD technology initially realized parallel computing of byte, half word, and word type data by splitting the data of 64-bit registers into multiple 8-bit, 16-bit, and 303-bit forms. Later, in order to further increase the parallelism of computing, SIMD technology began to meet the application's demand for computing power by increasing the register bit width. Typical SIMD structures are Intel's MMX, ARM's NEON, CEVA, TI's c6000, etc.

[0003] Another way to improve data parallelism is vector computing technology. Vector Processor System (VPS) is a parallel processing computer system oriented to vector parallel computing and based on pipeline structure. The use of parallel processing structures such as lookahead control and overlapping operation technology and operation pipeline plays an important role in improving the computing speed. Like traditional SIMD technology, it also increases the parallelism of computing by expanding the register bit width, but the difference is that vector registers are variable-length registers.

[0004] The register group of a traditional vector processor is composed of a long register. Each operation calculates the data of the entire vector register together. It is not flexible enough and cannot solve the problem of non-parallel data arrangement. Summary of the invention

[0005] The present invention provides a vector register device, a vector register addressing method and an electronic device, which are used to solve the defect of poor data arrangement in the prior art, and solve the problem when the data arrangement is not parallel.

[0006] The present invention provides a vector register device, comprising: a decoding unit, an address vector unit and a vector register unit, wherein the decoding unit is connected to the address vector unit and the vector register unit respectively, the address vector unit is connected to the vector register unit, the address vector unit comprises an address register group, the address register group comprises a plurality of address registers, the vector register unit comprises a vector register group, the vector register group comprises a plurality of vector registers, and the vector register comprises a plurality of vector units;

[0007] The decoding unit is used to translate the register address code to obtain a translation address, and send the translation address to the address vector unit or the vector register unit;

[0008] The address vector unit is used to determine the target address register according to the translation address, read the data in the target address register, and send the data in the target address register to the vector register unit;

[0009] The vector register unit is used to read the data in the target vector register according to the data in the target address register, or to read the data in the target vector register according to the translation address.

[0010] According to a vector register device provided by the present invention, the translation address includes a first translation address representing direct addressing and a second translation address representing indirect addressing;

[0011] The first translation address includes a vector register index and a unit index;

[0012] The second translated address includes an address register index.

[0013] According to a vector register device provided by the present invention, the vector register unit includes an address calculation module, a vector register module and a unit integration module;

[0014] The address calculation module is used to determine the target vector register according to the vector register index, and generate a mask factor and a shift factor based on the unit index and the unit length in the register address encoding;

[0015] The vector register module is used to read the data in the target vector register;

[0016] The unit integration module is used to perform a shift operation and a mask operation on the data in the target vector register based on the mask factor and the shift factor to obtain final vector data.

[0017] According to a vector register device provided by the present invention, the address calculation module is further used to disassemble each address unit in the target address register to obtain the target vector register address;

[0018] The vector register module is further used to read data in the corresponding target vector unit according to the target vector register address;

[0019] The unit integration module is used to perform a splicing operation on the data in the target vector unit to obtain final vector data.

[0020] According to a vector register device provided by the present invention, the vector unit is obtained by cutting the vector register based on the target bit width;

[0021] The address register includes a plurality of address units, and the number of address units in each of the address registers is the same as the number of vector units in each of the vector registers.

[0022] According to a vector register device provided by the present invention, the number of bits of the address unit satisfies the following formula:

[0023] ;

[0024] Where x represents the number of bits in the address unit, N represents the number of vector registers in the vector register group, and M represents the number of vector units in each vector register. Indicates rounding up.

[0025] The present invention also provides a vector register addressing method, comprising:

[0026] Translate the register address code to obtain the translation address;

[0027] Determining an addressing mode according to the translation address;

[0028] If the addressing mode is direct addressing, determining a target vector register based on the translation address, and reading data in the target vector register;

[0029] If the addressing mode is indirect addressing, determining a target address register based on the translation address, and reading data in the target address register;

[0030] The data in the target address register is used as the target vector register address, and based on the target vector register address, the data in the target vector register is read.

[0031] According to a vector register addressing method provided by the present invention, determining a target vector register based on the translation address and reading data in the target vector register comprises:

[0032] determining the target vector register according to the vector register index in the translation address, and generating a mask factor and a shift factor based on the unit index in the translation address and the unit length in the register address encoding;

[0033] Reading data in the target vector register;

[0034] Based on the mask factor and the shift factor, a shift operation and a mask operation are performed on the data in the target vector register to obtain final vector data.

[0035] According to a vector register addressing method provided by the present invention, the data in the target address register is used as the target vector register address, and based on the target vector register address, the data in the target vector register is read, including:

[0036] Disassembling each address unit in the target address register to obtain a target vector register address;

[0037] Read data in the corresponding target vector unit according to the target vector register address;

[0038] The data in the target vector unit is concatenated to obtain final vector data.

[0039] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-mentioned vector register addressing methods when executing the computer program.

[0040] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the vector register addressing method described in any one of the above is implemented.

[0041] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned vector register addressing methods.

[0042] The vector register device, vector register addressing method and electronic device provided by the present invention are divided according to vector units, and the non-traditional vector register can only read the vector register by row, so that the data reading from the register is flexible, and it can have greater flexibility in the reading stage, so that it has great flexibility for data processing. At the same time, each vector unit in the vector register can be addressed, and the interleaving arrangement of data is supported when addressing, and separate hardware is no longer required to arrange the data, which saves silicon cost and processing time. Through the setting of the address vector unit and the vector register unit, it has two modes of direct addressing and indirect addressing, that is, it can read the vector by row like a traditional vector processor, and it can also flexibly read any vector unit in the vector register by indirect addressing, and realize the interleaving arrangement of data in the data reading stage. Indirect addressing is performed through a group of address registers during indirect addressing, and any vector unit within the addressing range is read, which solves the problem of non-parallel data arrangement in the vector processor, especially the problem of simultaneous addressing of matrix row and column vectors, and optimizes the data arrangement in the vector processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0044] Figure 1 It is a structural schematic diagram of the vector register device provided by the present invention.

[0045] Figure 2 It is a schematic diagram of the structure of the address register group in the embodiment provided by the present invention.

[0046] Figure 3 It is a schematic diagram of the structure of the vector register group in the embodiment provided by the present invention.

[0047] Figure 4 It is a schematic diagram of the structure of the vector register unit in the embodiment provided by the present invention.

[0048] Figure 5 It is a flow chart of the vector register addressing method provided by the present invention.

[0049] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention.

[0050] Reference numerals:

[0051] 1: decoding unit; 2: address vector unit; 3: vector register unit; 301: vector register group; 302: address calculation module; 303: vector register module; 304: unit integration module. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0053] Combine the following Figure 1-Figure 6 The present invention describes a vector register device, a vector register addressing method and an electronic device.

[0054] Figure 1 FIG. 1 is a schematic diagram showing the structure of a vector register device according to an exemplary embodiment. Figure 1 As shown, in an exemplary embodiment, the vector register device includes: a decoding unit 1, an address vector unit 2 and a vector register unit 3, wherein the decoding unit 1 is connected to the address vector unit 2 and the vector register unit 3 respectively, the address vector unit 2 is connected to the vector register unit 3, the address vector unit 2 includes an address register group, the address register group includes a plurality of address registers, the vector register unit 3 includes a vector register group, the vector register group includes a plurality of vector registers, and the vector register includes a plurality of vector units;

[0055] The decoding unit 1 is used for translating the register address code to obtain a translation address, and sending the translation address to the address vector unit 2 or the vector register unit 3;

[0056] The address vector unit 2 is used to determine the target address register according to the translation address, read the data in the target address register, and send the data in the target address register to the vector register unit 3;

[0057] The vector register unit 3 is used to read the data in the target vector register according to the data in the target address register, or to read the data in the target vector register according to the translation address.

[0058] In the embodiment of the present invention, Figure 1 As shown, the decoding unit 1 is connected to the address vector unit 2 and the vector register unit 3 respectively, and the address vector unit 2 is connected to the vector register unit 3. Figure 2As shown, the address register group includes K address registers, each address register has X bits. Each address register is composed of M address units, each address unit has x bits. Each address unit is composed of two parts, the high bit represents the vector register index, and the low bit represents the unit index. The vector register index indicates a specific vector register, and the unit index indicates a specific vector unit in the vector register. The two together constitute the addressing address of a vector register, that is, the target vector register address. The addressing address is any vector unit in the vector register, so that data at any position of the vector register group can be read through this group of addresses.

[0059] The vector register unit 3 includes a vector register group, such as Figure 3 As shown, the vector register group includes N vector registers with a bit width of W. Each vector register group is cut into M parts according to the target bit width w to obtain M vector units. Each vector unit can be accessed, and the target bit width is the minimum bit width supported.

[0060] In the embodiment of the present invention, a register address code {A, L} consisting of an address code and a unit length is input, where A represents the address code and L represents the unit length. The decoding unit 1 translates the register address code to obtain a translation address. When the addressing mode used is different, the translation address obtained is also different. If direct addressing is used, the address code is translated into , which is indexed by a vector register and cell index If it is indirect addressing, the address code is translated into , which represents the address register index. When the translation address represents direct addressing, the decoding unit 1 sends the translation address to the vector register unit 3, and when the translation address represents indirect addressing, the decoding unit 1 sends the translation address to the address vector unit 2.

[0061] During indirect addressing, the address vector unit 2 determines the target address register according to the translation address, and the target address register is the address register in the address register group indicated in the translation address. After the target address register is determined according to the translation address, the data in the target address register is read, and the data in the target address register is sent to the vector register unit 3. The vector register unit 3 determines the corresponding target vector unit according to the data received in the target address register, and reads the data in the target vector unit. Since the address unit in the address register is composed of a high bit and a low bit, it is possible to arbitrarily read the data of multiple vector units from any vector register according to the address unit in the target address register, and combine them into a vector.

[0062] During direct addressing, the vector register unit 3 directly determines the target vector register according to the received translation address, the target vector register being the vector register indicated by the translation address, and directly reads the data in the target vector register.

[0063] In the embodiment of the present invention, the vector register is divided according to the vector unit, and the non-traditional vector register can only read the vector register by row, so that the data reading from the register is flexible, and it can have greater flexibility in the reading stage, so that it has great flexibility for data processing. At the same time, each vector unit in the vector register can be addressed, and the interleaving arrangement of data is supported when addressing, and separate hardware is no longer required to arrange the data, which saves silicon overhead and processing time. Through the setting of the address vector unit 2 and the vector register unit 3, there are two modes of direct addressing and indirect addressing, that is, the vector can be read by row like a traditional vector processor, and any vector unit in the vector register can be flexibly read by indirect addressing, and the interleaving arrangement of data is realized in the data reading stage. Indirect addressing is performed through a group of address registers during indirect addressing to realize reading any vector unit within the addressing range, which solves the problem of non-parallel data arrangement in the vector processor, especially the problem of simultaneous addressing of matrix row and column vectors, and optimizes the data arrangement in the vector processor.

[0064] In an exemplary embodiment of the present invention, the translation address includes a first translation address representing direct addressing and a second translation address representing indirect addressing;

[0065] The first translation address includes a vector register index and a unit index;

[0066] The second translated address includes an address register index.

[0067] In the embodiment of the present invention, if direct addressing is used, the register address code is translated into , which is indexed by a vector register and cell index If it is indirect addressing, the register address code is translated into , indicating the address register index.

[0068] In an exemplary embodiment of the present invention, the vector register unit 3 includes an address calculation module 302, a vector register module 303 and a unit integration module 304;

[0069] The address calculation module 302 is used to determine the target vector register according to the vector register index, and generate a mask factor and a shift factor based on the unit index and the unit length in the register address encoding;

[0070] The vector register module 303 is used to read the data in the target vector register;

[0071] The unit integration module 304 is used to perform a shift operation and a mask operation on the data in the target vector register based on the mask factor and the shift factor to obtain final vector data.

[0072] In the embodiment of the present invention, Figure 4 As shown, the vector register unit 3 is composed of a vector register group 301, an address calculation module 302, a vector register module 303 and a unit integration module 304. The address calculation module 302 is used to calculate the address of the read data. Specifically, when directly addressing, the reading work is divided into two parts. First, according to the translation address Register index of Find the corresponding target vector register, and then translate the address The unit index and unit length L generate a mask factor and shift factor, the target small bit width w, the unit index , the shift factor is ; The mask factor is (1<< )-1. The address calculation module 302 is used to determine the vector register index , mask factor and shift factor. This direct addressing method is faster and has lower overhead, but it has low flexibility and can only read part of the continuous units in a vector register.

[0073] The vector register module 303 uses the vector register index Read the data in the destination vector register.

[0074] The unit integration module 304 receives the read data, performs a shift operation on the data, and then performs a mask operation to obtain a final result.

[0075] In the embodiment of the present invention, direct addressing can read multiple continuous vector units at any starting point in a vector register. Its addressing mode is relatively fast and has relatively low overhead. Compared with the traditional SIMD register group, the utilization rate of the vector register is improved. True vector processing is achieved, and the data in the vector register is read according to a starting point and length, while the traditional vector processor can only read a complete line.

[0076] In an exemplary embodiment of the present invention, the address calculation module 302 is further used to disassemble each address unit in the target address register to obtain the target vector register address;

[0077] The vector register module 303 is further used to read data in the corresponding target vector unit according to the target vector register address;

[0078] The unit integration module 304 is used to perform a splicing operation on the data in the target vector unit to obtain final vector data.

[0079] In the embodiment of the present invention, during indirect addressing, the decoding unit 1 translates the register address code to determine the target address register from the address register group, disassembles each address unit in the target address memory, and forms a target vector register address. , where the i-th address Indexed by vector register and cell index constitute.

[0080] The vector register module 303 reads out M target vector units according to the target vector register address. , sent to the unit integration module 304.

[0081] The unit integration module 304 performs a concatenation operation on the lower L units of the received M target vector units, and then adds high-order bits to the output bit width to obtain a final result, where L≤M.

[0082] In the embodiment of the present invention, the indirect addressing mode can realize reading of a vector unit at any position in the vector register group, and its flexible performance enhances the ability of the vector processor to process non-parallel data.

[0083] In an exemplary embodiment of the present invention, the vector unit is obtained by cutting the vector register based on the target bit width;

[0084] The address register includes a plurality of address units, and the number of address units in each of the address registers is the same as the number of vector units in each of the vector registers.

[0085] In the embodiment of the present invention, as mentioned above, there are N vector register groups, each vector register has a bit width of W bits, and each vector register group is cut into M parts according to the target small bit width w, that is, W = w × M. Each part is called a vector unit, and each vector unit can be accessed.

[0086] The address register group has a total of K address registers, each of which has a bit width of X bits. Each address register consists of M address units, each of which has x bits, that is, X=x×M.

[0087] In an exemplary embodiment of the present invention, the number of bits of the address unit satisfies the following formula:

[0088] ;

[0089] Where x represents the number of bits in the address unit, N represents the number of vector registers in the vector register group, and M represents the number of vector units in each vector register. Indicates rounding up.

[0090] In the embodiment of the present invention, the number of bits of the address unit is related to the number N of the vector register groups and the number M of the address units in the address register. The specific relationship is referred to the above formula.

[0091] The vector register device provided in the embodiment of the present invention supports addressing by row and by column while performing matrix operations, and supports addressing by positive diagonal and anti-diagonal lines while performing matrix operations, thereby increasing the efficiency of matrix processing, eliminating the need for separate data placement processing, and improving the efficiency of matrix-related calculations. At the same time, it also supports discontinuous skip step addressing of vectors, and implements stride reading of data in a vector through addressing of a register stack without the need for additional hardware.

[0092] Figure 5 FIG. 1 is a flow chart showing a vector register addressing method according to an exemplary embodiment. Figure 5 As shown, in an exemplary embodiment, the vector register addressing method includes steps 510 to 550, which are described in detail as follows.

[0093] Step 510: translate the register address code to obtain a translated address.

[0094] In the embodiment of the present invention, the register address code {A, L} consisting of the address code and the unit length is translated to obtain a translated address.

[0095] Step 520: determine an addressing mode according to the translation address.

[0096] In the embodiment of the present invention, when the adopted addressing modes are different, the obtained translation addresses are also different. Therefore, the corresponding addressing mode can be determined according to the obtained translation address.

[0097] Step 530: If the addressing mode is direct addressing, determine the target vector register based on the translation address, and read the data in the target vector register.

[0098] In the embodiment of the present invention, during direct addressing, the target vector register is directly determined according to the translation address, the target vector register is a vector register in the vector register group indicated by the translation address, and data in the target vector register is directly read.

[0099] Step 540: If the addressing mode is indirect addressing, determine the target address register based on the translated address, and read the data in the target address register.

[0100] In the embodiment of the present invention, during indirect addressing, a target address register is determined according to the translation address, and the target address register is an address register in the address register group indicated in the translation address.

[0101] Step 550: Use the data in the target address register as the target vector register address, and read the data in the target vector register based on the target vector register address.

[0102] In the embodiment of the present invention, after the target address register is determined according to the translation address, the data in the target address register is read, and the data in the target address register is used as the target vector register address, and then based on the target vector register address, the data in the target vector register is read.

[0103] In an exemplary embodiment of the present invention, determining a target vector register based on the translation address and reading data in the target vector register include the following steps, which are described in detail as follows.

[0104] determining the target vector register according to the vector register index in the translation address, and generating a mask factor and a shift factor based on the unit index in the translation address and the unit length in the register address encoding;

[0105] Reading data in the target vector register;

[0106] Based on the mask factor and the shift factor, a shift operation and a mask operation are performed on the data in the target vector register to obtain final vector data.

[0107] In the embodiment of the present invention, when directly addressing, the reading work is divided into two parts. First, according to the translation address Register index of Find the corresponding target vector register, and then translate the address The cell index of And the unit length L generates a mask factor and shift factor. According to the register index Read the data in the target vector register. Perform shift and mask operations on the read data to get the final result.

[0108] In an exemplary embodiment of the present invention, the method of using the data in the target address register as the target vector register address and reading the data in the target vector register based on the target vector register address includes the following steps, which are described in detail as follows.

[0109] Each address unit in the target address register is disassembled to obtain a target vector register address.

[0110] The data in the corresponding target vector unit is read according to the target vector register address.

[0111] The data in the target vector unit is concatenated to obtain final vector data.

[0112] In the embodiment of the present invention, during indirect addressing, the register address code is translated to determine the target address register from the address register group, and each address unit in the target address memory is disassembled to form the target vector register address. . According to the target vector register address, read out M target vector units , perform concatenation operation on the lower L units of the M target vector units read, and then add high-order bits to the output bit width to obtain the final result.

[0113] The use of the vector register device provided by the embodiment of the present invention in a vector processor can greatly improve the ability of the vector processor to process non-parallel data. It has great application space in data switching, reordering, interleaving, column addressing, diagonal addressing, etc. With indirect addressing, the vector processor will no longer need logic such as reordering, interleaving, scaling, and expansion, reducing the silicon overhead of this part; it also does not need to perform data switching when reading and writing storage. Special computing requirements, such as column addressing, diagonal addressing, stride addressing, etc. can all be supported, improving the computing power of the vector processor.

[0114] Take a vector register device as an example, wherein the bit width of the vector register is W=128, the target bit width is w=16, the number of vector units in the vector register is M=8, the number of vector registers in the vector register group is N=8, the number of address registers is K=4, the number of bits of the address unit is x=6, and the number of bits of the address register is X=24.

[0115] In one example, for the interleaving operation, if the original sequence Store in Figure 3 In the vector register 1 shown, it is necessary to interleave it, and the result is If the indirect addressing method of the vector register device provided by the present invention is used, a single read can be completed, and the target vector register address of the address register is set to .

[0116] In common vector processor instruction architectures, such as Intel's AVX and ARM's neon instruction sets, there are no special instructions of this type. The positions need to be moved one by one through move instructions, so at least 8 instructions are required to complete this operation. However, the vector register device provided by the present invention only needs one instruction to complete it. In this example, the instruction effect is improved by 87.5%.

[0117] In one example, for the reverse operation, if the original sequence Stored in vector register 1, it needs to be reversed, the result is If the indirect addressing method of the vector register device provided by the present invention is used, a single read can be completed, and the target vector register address of the address register is set to .

[0118] The neon instruction set of Arm includes a VREV reverse instruction, which requires separate hardware support and has certain hardware overhead, while the vector register device provided by the present invention can be completed only through configuration.

[0119] In one example, for the expansion operation, if the original sequence Stored in vector register 1, the first two units need to be expanded by 4, the result is If the indirect addressing method of the vector register device provided by the present invention is used, a single read can be completed, and the address unit of the address register is set to .

[0120] The vector instruction sets of Intel and ARM both require separate hardware to support, while the vector register device provided by the present invention can be completed only through configuration without involving additional hardware overhead.

[0121] In one example, for the stride operation, if the original sequence Stored in vector register 1, it needs to be read with strides, and the result is If the indirect addressing method of the vector register device provided by the present invention is used, a single read can be completed, and the target vector register address of the address register is set to .

[0122] The vector instruction sets of Intel and ARM both require separate hardware to support, while the vector register device provided by the present invention can be completed only through configuration without involving additional hardware overhead.

[0123] In one example, for column addressing operations, if a (4x4) matrix is ​​placed in a vector register, it occupies 0-3 vector units in vector registers No. 1-4 of the vector register group. If the indirect addressing method of the vector register device provided by the present invention is used, the first column of the matrix is ​​read, and the target vector register address of the address register is set to .

[0124] In common vector processor instruction architectures, such as Intel's AVX and ARM's neon instruction sets, there are no special instructions of this type. The positions need to be moved one by one through move instructions, so at least 4 instructions are required to complete this operation. The vector register device provided by the present invention only requires one instruction, and the instruction effect in this example is improved by 75%.

[0125] In one example, for diagonal addressing operations, if a (4x4) matrix is ​​placed in a vector register, it occupies 0-3 vector units of vector registers 1-4 of the vector register group. If the indirect addressing method of the vector register device provided by the present invention is used to read the main diagonal elements of the matrix, the target vector register address of the address register is set to .

[0126] In common vector processor instruction architectures, such as Intel's AVX and ARM's neon instruction sets, there are no special instructions of this type. The positions need to be moved one by one through move instructions, so at least 4 instructions are required to complete this operation; while the vector register device provided by the present invention only requires one instruction, and the instruction effect in this example is improved by 75%.

[0127] Figure 6 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 6 As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communications interface 620 and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute the vector register addressing method, which includes: translating the register address encoding to obtain a translation address;

[0128] Determining an addressing mode according to the translation address;

[0129] If the addressing mode is direct addressing, determining a target vector register based on the translation address, and reading data in the target vector register;

[0130] If the addressing mode is indirect addressing, determining a target address register based on the translation address, and reading data in the target address register;

[0131] The data in the target address register is used as the target vector register address, and based on the target vector register address, the data in the target vector register is read.

[0132] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0133] On the other hand, the present invention further provides a computer program product, the computer program product includes a computer program, the computer program can be stored in a non-transitory computer-readable storage medium, when the computer program is executed by a processor, the computer can execute the vector register addressing method provided by the above methods, the method includes: translating the register address code to obtain a translation address;

[0134] Determining an addressing mode according to the translation address;

[0135] If the addressing mode is direct addressing, determining a target vector register based on the translation address, and reading data in the target vector register;

[0136] If the addressing mode is indirect addressing, determining a target address register based on the translation address, and reading data in the target address register;

[0137] The data in the target address register is used as the target vector register address, and based on the target vector register address, the data in the target vector register is read.

[0138] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, the computer program being implemented when executed by a processor to perform the vector register addressing method provided by the above methods, the method comprising: translating a register address encoding to obtain a translation address;

[0139] Determining an addressing mode according to the translation address;

[0140] If the addressing mode is direct addressing, determining a target vector register based on the translation address, and reading data in the target vector register;

[0141] If the addressing mode is indirect addressing, determining a target address register based on the translation address, and reading data in the target address register;

[0142] The data in the target address register is used as the target vector register address, and based on the target vector register address, the data in the target vector register is read.

[0143] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0144] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A vector register device, characterized in that: include: A decoding unit, an address vector unit and a vector register unit, wherein the decoding unit is connected to the address vector unit and the vector register unit respectively, the address vector unit is connected to the vector register unit, the address vector unit includes an address register group, the address register group includes a plurality of address registers, the vector register unit includes a vector register group, the vector register group includes a plurality of vector registers, and the vector register includes a plurality of vector units; The decoding unit is used to translate the register address code to obtain a translation address, and send the translation address to the address vector unit or the vector register unit; The address vector unit is used to determine the target address register according to the translation address, read the data in the target address register, and send the data in the target address register to the vector register unit; The vector register unit is used to read the data in the target vector register according to the data in the target address register, or read the data in the target vector register according to the translation address; The translation address includes a first translation address representing direct addressing and a second translation address representing indirect addressing; The first translation address includes a vector register index and a unit index; The second translation address includes an address register index; The vector register unit includes an address calculation module, a vector register module and a unit integration module; The address calculation module is used to determine the target vector register according to the vector register index, and generate a mask factor and a shift factor based on the unit index and the unit length in the register address encoding; The vector register module is used to read the data in the target vector register; The unit integration module is used to perform a shift operation and a mask operation on the data in the target vector register based on the mask factor and the shift factor to obtain final vector data.

2. The vector register device according to claim 1, characterized in that: The address calculation module is further used to disassemble each address unit in the target address register to obtain the target vector register address; The vector register module is further used to read data in the corresponding target vector unit according to the target vector register address; The unit integration module is used to perform a splicing operation on the data in the target vector unit to obtain final vector data.

3. The vector register device according to any one of claims 1 to 2, characterized in that: The vector unit is obtained by cutting the vector register based on the target bit width; The address register includes a plurality of address units, and the number of address units in each of the address registers is the same as the number of vector units in each of the vector registers.

4. The vector register device according to claim 3, characterized in that: The number of bits of the address unit satisfies the following formula: ; Where x represents the number of bits in the address unit, N represents the number of vector registers in the vector register group, and M represents the number of vector units in each vector register. Indicates rounding up.

5. A vector register addressing method, characterized in that: include: Translate the register address code to obtain the translation address; Determining an addressing mode according to the translation address; If the addressing mode is direct addressing, determining a target vector register based on the translation address, and reading data in the target vector register; If the addressing mode is indirect addressing, determining a target address register based on the translation address, and reading data in the target address register; Using the data in the target address register as the target vector register address, and reading the data in the target vector register based on the target vector register address; The translation address includes a first translation address representing direct addressing and a second translation address representing indirect addressing; The first translation address includes a vector register index and a unit index; The second translation address includes an address register index; The step of determining a target vector register based on the translation address and reading data in the target vector register comprises: determining the target vector register according to the vector register index in the translation address, and generating a mask factor and a shift factor based on the unit index in the translation address and the unit length in the register address encoding; Reading data in the target vector register; Based on the mask factor and the shift factor, a shift operation and a mask operation are performed on the data in the target vector register to obtain final vector data.

6. The vector register addressing method according to claim 5, characterized in that: The step of using the data in the target address register as the target vector register address and reading the data in the target vector register based on the target vector register address includes: Disassembling each address unit in the target address register to obtain a target vector register address; Read data in the corresponding target vector unit according to the target vector register address; The data in the target vector unit is concatenated to obtain final vector data.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the vector register addressing method according to any one of claims 5 to 6 is implemented.

Citation Information

Patent Citations

  • Processor, chip, electronic equipment and data processing method

    CN114942831A

  • Irregular memory access processing method and device and electronic equipment

    CN116700798A