Memory access instruction processing method and processor

By splitting the vector memory fetch instruction into micro-operation and merging scalar memory fetch instruction, the problem of unreasonable resource allocation in the existing technology is solved, and efficient data access and resource sharing are achieved.

CN119621149BActive Publication Date: 2025-05-16BEIJING VCORE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510162028.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-16
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The prior art is difficult to allocate resources for vector memory access and scalar memory access reasonably, making it difficult to achieve efficient data access and data transmission.

Method used

By splitting the vector memory access instruction into multiple vector memory access micro operations, and combining the scalar memory access instruction with the unexecuted scalar memory access instruction into one scalar memory access micro operation, the memory access width determined by the number of memory banks and bit width of the data cache is processed.

Benefits of technology

It realizes the allocation of resources for vector memory access and scalar memory access more reasonably when data access is retrieved, improves the efficiency and resource utilization of the processor, and supports maximum resource sharing of vector memory access and scalar memory access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621149B_ABST
    Figure CN119621149B_ABST
Patent Text Reader

Abstract

The present invention provides a memory access instruction processing method and a processor, which relate to the field of data processing technology. The method comprises: splitting a target vector memory access instruction into multiple target vector memory access micro-operations based on memory access width, merging a target scalar memory access instruction and at least one unexecuted scalar memory access instruction into a target scalar memory access micro-operation; executing each target vector memory access micro-operation, and / or executing a target scalar memory access micro-operation; the memory access width is determined based on the number of data cache storage bodies in a data cache in a processor and the bit width of the data cache storage body. The memory access instruction processing method and the processor provided by the present invention can more reasonably allocate resources for vector memory access and scalar memory access during data memory access through the splitting of vector memory access instructions and the merging of scalar memory access instructions, thereby realizing maximum resource sharing of vector memory access and scalar memory access.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a memory access instruction processing method and a processor. Background Art

[0002] With the rapid development of artificial intelligence and big data technologies, users have put forward higher and higher requirements on the performance of artificial intelligence algorithms and data analysis applications, and therefore the requirements on computer processing capabilities are also getting higher and higher. Vector processing technology has become a hot technology that is of common concern to academia and industry because it can process multiple data with one instruction, provide data-level parallelism, and thus greatly improve data processing efficiency.

[0003] The computing performance of the processor is closely related to the efficiency of computing data supply. The data access bit width of vector memory access is relatively large, usually reaching 128 bits, 256 bits or 512 bits, which occupies more processor resources and has a greater impact on the processor's main frequency, chip area and power consumption. The data access bit width of scalar memory access is relatively small, usually 8 bits, 16 bits, 32 bits or 64 bits, which occupies fewer processor resources. The data access bit width of vector memory access and scalar memory access is quite different.

[0004] In the related art, when a processor that can support both vector memory access and scalar memory access performs data memory access, it is difficult to reasonably allocate resources for vector memory access and scalar memory access, and thus it is difficult to achieve efficient data access and data transmission. Therefore, how to more reasonably allocate resources for vector memory access and scalar memory access when processing memory access instructions, so as to achieve the maximum resource sharing between vector memory access and scalar memory access, is a technical problem to be solved in this field. Summary of the invention

[0005] The present invention provides a memory access instruction processing method and a processor, which are used to solve the defect that a processor capable of simultaneously supporting vector memory access and scalar memory access in the prior art is difficult to reasonably allocate resources for vector memory access and scalar memory access when performing data memory access, and to achieve more reasonable allocation of resources for vector memory access and scalar memory access when performing data memory access.

[0006] The present invention provides a memory access instruction processing method, comprising the following steps.

[0007] For the target vector memory access instruction currently to be dispatched in the vector memory access instruction dispatch queue, based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction, the target vector memory access instruction is split into multiple target vector memory access micro-operations; for the target scalar memory access instruction currently to be dispatched in the scalar memory access instruction dispatch queue, based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, the target scalar memory access instruction and at least one unexecuted scalar memory access instruction are merged into a target scalar memory access micro-operation; each of the target vector memory access micro-operations is executed, and / or the target scalar memory access micro-operation is executed; wherein the memory access width is determined based on the number of data cache storage bodies in the data cache in the processor and the bit width of the data cache storage body.

[0008] According to a memory access instruction processing method provided by the present invention, the target vector memory access instruction is split into multiple target vector memory access micro-operations based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction, including: when it is determined based on the vector instruction opcode of the target vector memory access instruction that the target vector memory access instruction is a vector store instruction and the target vector memory access instruction is for the same data cache storage line in the data cache, based on the memory access width and the configuration information of the vector control and status register of the target vector memory access instruction, splitting the store address and store data of the target vector memory access instruction into multiple target vector memory access micro-operations; The method comprises the steps of: obtaining the storage address and storage data corresponding to each of the target vector memory access micro-operations; executing each of the target vector memory access micro-operations, including: transmitting the storage address corresponding to each of the target vector memory access micro-operations to the storage instruction address pipeline, so that the storage address corresponding to each of the target vector memory access micro-operations can access the translation backup address buffer through the storage instruction address pipeline, and then writing back the storage address corresponding to each of the target vector memory access micro-operations; directly writing back the storage data corresponding to each of the target vector memory access micro-operations after transmitting; and determining that the target vector memory access instruction has completed writing back when it is determined that the storage address and storage data corresponding to each of the target vector memory access micro-operations are all written back.

[0009] According to a memory access instruction processing method provided by the present invention, the target vector memory access instruction is split into multiple target vector memory access micro-operations based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction, including: when it is determined based on the vector instruction opcode of the target vector memory access instruction that the target vector memory access instruction is a vector fetch instruction and the target vector memory access instruction is for the same data cache storage line in the data cache, based on the memory access width and the configuration information of the vector control and status register of the target vector memory access instruction, fetching the target vector memory access instruction The data address is split to obtain the data fetch address corresponding to each of the target vector memory access micro-operations; the execution of each of the target vector memory access micro-operations includes: transmitting the data fetch address corresponding to each of the target vector memory access micro-operations to the data fetch instruction address pipeline, so that the data fetch address corresponding to each of the target vector memory access micro-operations can access the data cache and the translation backup address buffer through the data fetch instruction address pipeline, and then write back from each of the target vector memory access micro-operations to return the data fetch data; when it is determined that all of the target vector memory access micro-operations are written back, splicing to obtain the result data of the target vector memory access instruction, and then determining that the target vector memory access instruction completes the write back.

[0010] According to a memory access instruction processing method provided by the present invention, the target scalar memory access instruction and at least one unexecuted scalar memory access instruction are merged into a target scalar memory access micro-operation based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, including: in the case where it is determined that the target scalar memory access instruction is a scalar data fetch instruction based on the scalar instruction opcode of the target scalar memory access instruction, the scalar data fetch instruction corresponding to the target scalar memory access instruction is a scalar data fetch instruction corresponding to the target scalar memory access instruction based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, and the scalar data fetch instruction corresponding to the target scalar memory access instruction is an unexecuted scalar data fetch instruction, and the The scalar fetch instruction corresponding to the target scalar memory access instruction corresponds to the same data cache storage row as the target scalar memory access instruction; the target scalar memory access instruction and the scalar fetch instruction corresponding to the target scalar memory access instruction are merged into one of the target scalar memory access micro-operations; the execution of the target scalar memory access micro-operation includes: emitting the target scalar memory access micro-operation to the fetch instruction pipeline, so that after the target scalar memory access micro-operation accesses the data cache through the fetch instruction pipeline and the translation backup address buffer to access the data cache, the target scalar memory access instruction and the scalar fetch instruction corresponding to the target scalar memory access instruction are written back respectively.

[0011] According to a memory access instruction processing method provided by the present invention, based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, the target scalar memory access instruction and at least one unexecuted scalar memory access instruction are merged into a target scalar memory access micro-operation, including: in the case where it is determined based on the scalar instruction opcode of the target scalar memory access instruction that the target scalar memory access instruction is a scalar store instruction, the store address of the target scalar memory access instruction is transmitted to the store instruction address pipeline, so that the store address of the target scalar memory access instruction can access the translation backup address buffer through the store instruction address pipeline, and then the store address of the target scalar memory access instruction is transmitted to the translation backup address buffer through the store instruction address pipeline. The address and storage data are written into a scalar store instruction write-merge buffer; based on the scalar instruction opcode of the target scalar memory access instruction, a scalar store instruction corresponding to the target scalar memory access instruction is determined in the scalar store instruction write-merge buffer, the scalar store instruction corresponding to the target scalar memory access instruction is a scalar store instruction that has not been written to the data cache, and the scalar store instruction corresponding to the target scalar memory access instruction corresponds to the same data cache storage row as the target scalar memory access instruction; the target scalar memory access instruction and the scalar store instruction corresponding to the target scalar memory access instruction are merged into one of the target scalar memory access micro-operations in the scalar store instruction write-merge buffer.

[0012] According to a memory access instruction processing method provided by the present invention, the memory address and memory data of the target vector memory access instruction are split based on the memory access width and the configuration information of the vector control and status register of the target vector memory access instruction to obtain the memory address and memory data corresponding to each of the target vector memory access micro-operations, including: obtaining the bit width of the target vector memory access instruction based on the configuration information of the vector control and status register of the target vector memory access instruction; when the vector memory access instruction address is aligned according to the memory access width, calculating the quotient of the bit width of the target vector memory access instruction and the memory access width as a first quantity, and when the vector memory access instruction address is not aligned according to the memory access width, calculating the quotient of the bit width of the target vector memory access instruction and the memory access width plus 1 as the first quantity; splitting the memory address of the target vector memory access instruction into the first number of memory address groups, respectively serving as the memory address corresponding to each of the target vector memory access micro-operations, and splitting the memory data of the target vector memory access instruction into the first number of memory data groups, respectively serving as the memory data corresponding to each of the target vector memory access micro-operations.

[0013] According to a memory access instruction processing method provided by the present invention, the data access address of the target vector memory access instruction is split based on the memory access width and the configuration information of the vector control and status register of the target vector memory access instruction to obtain the data access address corresponding to each of the target vector memory access micro-operations, including: obtaining the bit width of the target vector memory access instruction based on the configuration information of the vector control and status register of the target vector memory access instruction; when the vector memory access instruction address is aligned according to the memory access width, calculating the quotient of the bit width of the target vector memory access instruction and the memory access width as the second number; when the vector memory access instruction address is not aligned according to the memory access width, calculating the quotient of the bit width of the target vector memory access instruction and the memory access width plus 1 as the second number; splitting the data access address of the target vector memory access instruction into the second number of data access address groups, respectively serving as the data access addresses corresponding to each of the target vector memory access micro-operations.

[0014] According to a memory access instruction processing method provided by the present invention, the scalar data fetch instruction corresponding to the target scalar memory access instruction is determined based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, including: obtaining the bit width of the target scalar memory access instruction and the target data cache storage line corresponding to the target scalar memory access instruction based on the scalar instruction opcode of the target scalar memory access instruction; querying whether there is a scalar data fetch instruction corresponding to the target data cache storage line; and when it is determined that there is a scalar data fetch instruction corresponding to the target data cache storage line, determining the scalar memory access instruction of the target data cache storage line as the scalar data fetch instruction corresponding to the target scalar memory access instruction, and the sum of the bit width of the target scalar memory access instruction and the bit width of the scalar data fetch instruction corresponding to the target scalar memory access instruction is not greater than the memory access width.

[0015] According to a memory access instruction processing method provided by the present invention, the scalar instruction opcode of the target scalar memory access instruction is used to determine the scalar store instruction corresponding to the target scalar memory access instruction in the scalar store instruction write-merge buffer, including: based on the scalar instruction opcode of the target scalar memory access instruction, obtaining the bit width of the target scalar memory access instruction and the target data cache storage line corresponding to the target scalar memory access instruction; querying in the scalar store instruction write-merge buffer whether there is a scalar store instruction corresponding to the target data cache storage line; when it is determined that there is a scalar store instruction corresponding to the target data cache storage line in the scalar store instruction write-merge buffer, determining the scalar store instruction corresponding to the target data cache storage line in the scalar store instruction write-merge buffer as the scalar store instruction corresponding to the target scalar memory access instruction, and the sum of the bit width of the target scalar memory access instruction and the bit width of the scalar store instruction corresponding to the target scalar memory access instruction is not greater than the bit width of the target data cache storage line.

[0016] The present invention also provides a processor, comprising: a vector memory access instruction splitting unit, for splitting the target vector memory access instruction currently to be dispatched in the vector memory access instruction dispatch queue into multiple target vector memory access micro-operations based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction; a scalar memory access instruction merging unit, for merging the target scalar memory access instruction currently to be dispatched in the scalar memory access instruction dispatch queue with at least one unexecuted scalar memory access instruction into one target scalar memory access micro-operation based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction; a memory access execution unit, for executing each of the target vector memory access micro-operations, and / or executing the target scalar memory access micro-operation; wherein the memory access width is determined based on the number of data cache storage bodies in the data cache of the processor and the bit width of the data cache storage body.

[0017] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the memory access instruction processing method described in any one of the above is implemented.

[0018] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned memory access instruction processing methods.

[0019] The memory access instruction processing method and processor provided by the present invention, for a target vector memory access instruction currently to be dispatched in a vector memory access instruction dispatch queue, splits the target vector memory access instruction into multiple target vector memory access micro-operations based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction, and then executes each target vector memory access micro-operation to complete the target vector memory access instruction; for a target scalar memory access instruction currently to be dispatched in a scalar memory access instruction dispatch queue, splits the target scalar memory access instruction into a plurality of target vector memory access micro-operations based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction, and then executes each target vector memory access micro-operation to complete the target vector memory access instruction; After the scalar memory access instructions are merged into a target scalar memory access micro-operation, the above target scalar memory access micro-operation is executed to complete the above target scalar memory access instruction. The vector memory access instructions can be split and the scalar memory access instructions can be merged based on the memory access width determined according to the number of data cache storage bodies in the data cache in the processor and the bit width of the data cache storage body. Furthermore, through the splitting of vector memory access instructions and the merging of scalar memory access instructions, resources can be more reasonably allocated to vector memory access and scalar memory access during data memory access, thereby realizing the maximum resource sharing of vector memory access and scalar memory access, which can improve the efficiency and resource utilization of the processor and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0021] Figure 1 It is a field distribution diagram of the vector type in the related technology.

[0022] Figure 2 This is one of the flow charts of the memory access instruction processing method provided by the present invention.

[0023] Figure 3 This is one of the flow charts of the memory access instruction processing method provided by the present invention.

[0024] Figure 4 It is a structural schematic diagram of the processor provided by the present invention. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0026] In the description of the invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0027] In the description of the present application, the terms "first", "second", etc. are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are usually a class, and the number of objects is not limited. For example, the first object can be one or more. In addition, in the description of the present application, "and / or" represents at least one of the connected objects, and the character " / " generally represents that the front and back associated objects are in an "or" relationship.

[0028] It should be noted that RISC-V (Reduced Instruction Set Computing-Five, the fifth generation of reduced instruction set) is currently a widely used open source reduced instruction set. The vector extension instruction set specification of the RISC-V instruction set, riscv-v-spec (RISC-V Vector Extension, RVV), is an important extension of the RISC-V architecture, which aims to improve the parallel processing performance of the processor. RVV can support vector operations of various bit widths and allow dynamic adjustment of vector length to flexibly adapt to different application requirements.

[0029] The RISC-V vector instruction set RVV contains a vector control and status register (CSR) and a vector register. The vector CSR is used to store configuration information, including vector length and operation bit width, and the vector register is used to store vectorized operation data.

[0030] RVV specifies that 32 vector registers (denoted as v0-v31, all with a bit width of VLEN) and 7 non-privileged CSRs (vstart, vxsat, vxrm, vcsr, vl, vtype, and vlenb, all with a bit width of XLEN) are extended on the basis of the basic spec (Specification). Among them, VLEN refers to the length of a vector register (Vector Length, VLEN); XLEN refers to the length of a vector control and status register.

[0031] In the configuration information stored in the vector CSR, vtype (vector type) is used to indicate the default type of the vector register content. The vector type also determines the arrangement of elements in each vector register and how to group multiple vector registers. Figure 1 This is a field distribution diagram of the vector type in the related technology. The fields included in the vector type are as follows: Figure 1 shown.

[0032] It should be noted that Figure 1 The vlmul[2:0] in the _DML_PAGE_STRING indicates vector register grouping. Multiple vector registers can be grouped together so that a single vector instruction can operate on multiple vector registers.

[0033] LMUL points to the number of logical lanes in the vector register (Lane Multiplication Factor). vlmul represents a signed number. The value of LMUL can be a fractional number in addition to an integer, which reduces the number of bits used in a single vector register. LMUL=2 vlmul[2:0] . Fractional grouping is mainly used for mixed width instruction operations. Therefore, the value of LMUL can be 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8. The corresponding relationship between LMUL and vlmul is shown in Table 1:

[0034] Table 1 Correspondence between LMUL and vlmul

[0035]

[0036] vsew[2:0] indicates the width of each element in the vector register. The vector register is considered to be divided into VLEN / SEW standard width elements. SEW refers to the standard width of a single vector element (Select Element Width, SEW).

[0037] The corresponding relationship between SEW and vsew is shown in Table 2:

[0038] Table 2 Correspondence between SEW and vsew

[0039]

[0040] Figure 1 In the , reserved means that the field is reserved, that is, it is not used in the current version of the instruction set. Reserved fields are usually used for future extensions or compatibility considerations.

[0041] There are two very important CSR registers in RVV, namely vl and vtype. vl is used to configure the number of elements to be processed, and vtype is used to configure SEW, LMUL and others.

[0042] vill: When set, it means an illegal vtype value is set. In this case, the other bits are all 0. This vtype will report an exception, and vill should also be set to 0.

[0043] vma: vector mask agnostic. Vector instructions can use masks to specify active elements. If this bit is set, it means that those inactive elements can keep their original values ​​or be overwritten with all 1s.

[0044] vta: vector tail is not sensed. Elements beyond the vl range are called tail elements (more accurate description is provided later). By default, these elements maintain their original values. However, if vta is set to 1 here, the tail element in the destination vector register group can maintain its original value or be fully overwritten by 1 (not maintaining the original value).

[0045] If vma and vta keep the default value 0, it is the undisturbed strategy, that is, the unchanged strategy. This strategy means that when changing the register, the old value must be read out first, merged, and then written back.

[0046] In order to improve data processing efficiency and flexibility, the RISC-V vector extension instruction set (RVV) has designed a series of flexible memory access instructions. These instructions allow data to be efficiently loaded from memory to vector registers, and data to be stored from vector registers back to memory. RVV memory access instructions mainly include: Strided Loads and Stores. This type of instruction allows data to be loaded from memory at fixed intervals (stride lengths), or data to be stored in fixed intervals in memory. It is suitable for processing spaced data sets (such as a column of a matrix).

[0047] Indexed Loads and Stores, Indexed memory instructions allow the use of an index value in another vector to specify the load or store location of each element. This provides great flexibility in handling irregular data structures, for example, when the data is distributed in non-contiguous locations in memory.

[0048] Unit-stride Loads and Stores: These instructions are used to load or store data continuously, from one address to the next without any gaps. This is the most direct form of memory access and is suitable for efficient processing of continuous data structures.

[0049] Masked Loads and Stores, RVV supports conditional loads and stores using masks. This means that elements in a vector can be selectively loaded or stored based on the bits in the mask vector. This is very useful for handling conditional operations or sparse data sets.

[0050] Through these flexible memory access instructions, RVV can effectively support various data access patterns, thereby optimizing performance and simplifying the programming model. These instructions make RVV very suitable for handling a wide range of application scenarios, including high-performance computing, data analysis, machine learning, etc., where data access patterns may be very diverse.

[0051] The number of vector elements operated by these memory access instructions is VLEN LMUL / SEW. VLEN indicates the length of a vector source register corresponding to the target vector instruction; LMUL indicates the number of logical channels of the vector source register; SEW indicates the standard width of the vector elements in the vector source register. For example, if the length of a vector source register corresponding to a vector instruction is 256 bits, the number of logical channels of the vector source register is 8, and the standard width of the vector elements in the vector source register is 16 bits, the number of vector elements is 128.

[0052] Scalar memory access instructions generally include four types: byte fetch (lb), halfword fetch (lh), word fetch (lw), and doubleword fetch (ld). Scalar storage instructions include four types: byte store (sb), halfword store (sh), word store (sw), and doubleword store (sd). Byte, halfword, word, and doubleword are 8, 16, 32, and 64-bit data, respectively.

[0053] The execution granularity of vector memory access instructions is based on vector elements. The number of elements contained in each instruction is determined by the vector control and status registers and the execution type. The execution granularity of scalar memory access instructions is determined by the instruction type.

[0054] Both vector memory access instructions and scalar memory access instructions access the data cache. The width of the data cache line is generally 256 bits (bit) or 512 bits (bit), and is divided into multiple different data cache bodies. Each data cache body is generally 64 bits and can be accessed in parallel. Since the address comparison and data transfer related to the storage instruction and the load instruction are required after the data cache is accessed, considering the main frequency of the processor, the data bit width of a command operation cannot be too wide.

[0055] Vector memory access refers to the way of accessing data in memory in batches through vector instruction sets. Vector instruction sets allow multiple data elements to be processed simultaneously. These data elements are organized into vector form and stored in vector registers.

[0056] Scalar memory access refers to the way of accessing a single data element in memory through a scalar instruction set. The scalar instruction set can only process one data element at a time.

[0057] Correspondingly, the data access bit width of vector memory access is larger, while the data access bit width of scalar memory access is smaller. However, in the related art, when a processor that can support both vector memory access and scalar memory access performs data memory access, it is difficult to reasonably allocate resources for vector memory access and scalar memory access, and thus it is difficult to achieve efficient data access and data transmission. Therefore, how to more reasonably allocate resources for vector memory access and scalar memory access when processing memory access instructions, so as to achieve the maximum resource sharing between vector memory access and scalar memory access, is a technical problem that needs to be solved in this field.

[0058] Combine the following Figure 2-Figure 3 The memory access instruction processing method provided by the present invention is described.

[0059] Figure 2 This is one of the flow charts of the memory access instruction processing method provided by the present invention. The memory access instruction processing method provided by the present invention is applied to a processor. Figure 2 As shown, the method includes the following: Step 201, for the target vector memory access instruction currently to be dispatched in the vector memory access instruction dispatch queue, based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction, the target vector memory access instruction is split into multiple target vector memory access micro-operations; for the target scalar memory access instruction currently to be dispatched in the scalar memory access instruction dispatch queue, based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, the target scalar memory access instruction and at least one unexecuted scalar memory access instruction are merged into one target scalar memory access micro-operation.

[0060] The memory access width is determined based on the number of data cache memory banks in the data cache in the processor and the bit width of the data cache memory banks.

[0061] It should be noted that the execution subject of the embodiment of the present invention is a processor, and when executing the memory access instruction processing method provided by the present invention, the processor can process vector memory access instructions and scalar memory access instructions.

[0062] The processor is provided with a data cache storage row. The data cache storage row may include N data cache storage bodies. If the bit width of the data cache storage body is M bits, the bit width of the data cache storage row is M×N bits. Wherein, N represents a positive integer greater than zero; M represents a positive integer greater than zero.

[0063] In the embodiment of the present invention, the memory access width may be determined based on the bit width (M×N bits) of the data cache storage row and the bit width (M bits) of the data cache storage bank.

[0064] Specifically, in the embodiment of the present invention, the memory access width can be set to M×K bits, and ensure that . Wherein, K represents a positive integer greater than zero.

[0065] For example, when the data cache storage row includes N=8 data cache storage bodies and the bit width of the data cache storage body is M=64 bits, the bit width of the data cache storage row is 64×8=512 bits, and the memory access width can be set to a multiple of 64 not greater than 512, for example, the memory access width can be set to 2×64=128 bits, or 4×64=256 bits.

[0066] The above processor may be configured with a vector memory access instruction splitting unit, a scalar memory access instruction merging unit and a memory access execution unit.

[0067] It should be noted that the memory access execution unit in the embodiment of the present invention includes a store instruction address pipeline, a fetch instruction pipeline, and a scalar store instruction write-merge buffer, etc. The store instruction address pipeline and the fetch instruction pipeline are shared by vector memory access instructions and scalar memory access instructions.

[0068] It should be noted that after the vector memory access instruction enters the dispatch stage (Dispatch) from the mainstream pipeline, it will enter the vector memory access instruction dispatch queue and start executing the vector memory access instruction.

[0069] In the embodiment of the present invention, the vector memory access instruction to be dispatched currently in the vector memory access instruction dispatch queue is determined as the target vector memory access instruction.

[0070] The vector memory access instruction splitting component can split the target vector memory access instruction into multiple target vector memory access micro-operations based on the above-mentioned memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction.

[0071] It should be noted that the vector instruction opcodes of vector memory access instructions are used to identify different types of vector memory access operations, such as load (i.e., fetching data) and store (i.e., storing data), as well as whether the data operated by the instruction is a byte (8 bits), a halfword (16 bits), a word (32 bits), or a doubleword (64 bits). These opcodes are usually associated with a specific instruction format and determine the behavior of the instruction.

[0072] It should be noted that the RISC-V vector extension introduces a series of vector control and status registers to configure the execution environment of vector instructions, manage the status information of vector processing, and provide fine control over vector operations. These registers usually include the vector length register (VLEN), the vector type register (VTYPE), the vector start register (VSTART), etc., which together determine the behavior and performance of vector instructions. The configuration information of the vector control and status registers can be used to indicate the vector length and type, the starting position of the vector operation, the execution environment of the vector operation, and exception and interrupt handling.

[0073] It should be noted that after the scalar memory access instruction enters the dispatch stage (Dispatch) from the mainstream pipeline, it will enter the scalar memory access instruction dispatch queue and start executing the scalar memory access instruction.

[0074] In the embodiment of the present invention, the scalar memory access instruction currently to be dispatched in the scalar memory access instruction dispatch queue is determined as the target scalar memory access instruction.

[0075] The scalar memory access instruction merging unit may merge the target scalar memory access instruction and at least one unexecuted scalar memory access instruction into a target scalar memory access micro-operation based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction.

[0076] It should be noted that the scalar instruction opcode of the scalar memory access instruction is used to identify different types of scalar memory access operations. Scalar memory access instructions generally include two types of operations: load (i.e., fetching data) and store (i.e., storing data).

[0077] Step 202: Execute each target vector memory access micro-operation, and / or execute each target scalar memory access micro-operation.

[0078] Specifically, after acquiring a plurality of target vector memory access micro-operations, the memory access execution unit may execute each target vector memory access micro-operation, thereby completing the target vector memory access instruction.

[0079] After acquiring the target scalar memory access micro-operation, the memory access execution unit may also execute the target scalar memory access micro-operation, thereby completing the target scalar memory access instruction.

[0080] In an embodiment of the present invention, for a target vector memory access instruction currently to be dispatched in a vector memory access instruction dispatch queue, the target vector memory access instruction is split into a plurality of target vector memory access micro-operations based on a memory access width, a vector instruction opcode of the target vector memory access instruction, and configuration information of a vector control and status register of the target vector memory access instruction, and each target vector memory access micro-operation is executed to complete the target vector memory access instruction. For a target scalar memory access instruction currently to be dispatched in a scalar memory access instruction dispatch queue, the target scalar memory access instruction is separated from at least one unexecuted scalar memory access instruction based on the scalar instruction opcode of the target scalar memory access instruction. After the instructions are merged into a target scalar memory access micro-operation, the above target scalar memory access micro-operation is executed to complete the above target scalar memory access instruction. The vector memory access instruction can be split and the scalar memory access instruction can be merged based on the memory access width determined according to the number of data cache storage bodies in the data cache in the processor and the bit width of the data cache storage body. Furthermore, through the splitting of vector memory access instructions and the merging of scalar memory access instructions, resources can be more reasonably allocated to vector memory access and scalar memory access during data memory access, thereby realizing the maximum resource sharing of vector memory access and scalar memory access, which can improve the efficiency and resource utilization of the processor and has broad application prospects.

[0081] As an optional embodiment, based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status registers of the target vector memory access instruction, the target vector memory access instruction is split into multiple target vector memory access micro-operations, including: when the target vector memory access instruction is determined to be a vector store instruction based on the vector instruction opcode of the target vector memory access instruction and the target vector memory access instruction is for the same data cache storage line in the data cache, based on the memory access width and the configuration information of the vector control and status registers of the target vector memory access instruction, the store address and store data of the target vector memory access instruction are split to obtain the store address and store data corresponding to each target vector memory access micro-operation.

[0082] Specifically, the vector memory access instruction splitting unit may include a vector memory access instruction fetch and storage address splitting unit and a vector memory access instruction storage data splitting unit.

[0083] Figure 3 This is the second flow chart of the memory access instruction processing method provided by the present invention. Figure 3As shown, for the target vector memory access instruction currently to be dispatched in the vector memory access instruction dispatch queue, it can be determined whether the target vector memory access instruction is a vector store instruction or a vector fetch instruction based on the vector instruction opcode of the target vector memory access instruction.

[0084] When it is determined based on the vector instruction opcode of the target vector memory access instruction that the target vector memory access instruction is a vector store instruction and the target vector memory access instruction is for the same data cache storage row in the data cache, the target vector memory access instruction can be respectively transmitted to the vector memory access instruction fetch and store address splitting unit and the vector memory access instruction store data splitting unit.

[0085] The vector memory access instruction data fetching and storage address splitting component can split the storage address of the target vector memory access instruction based on the memory access width and the configuration information of the vector control and status register of the target vector memory access instruction to obtain the storage address corresponding to each target vector memory access micro-operation.

[0086] The vector memory access instruction storage data splitting component can split the storage data of the target vector memory access instruction based on the memory access width and the configuration information of the vector control and status register of the target vector memory access instruction to obtain the storage data corresponding to each target vector memory access micro-operation.

[0087] It should be noted that there is a corresponding relationship between the storage address and the storage data corresponding to each target vector memory access micro-operation.

[0088] As an optional embodiment, based on the memory access width and the configuration information of the vector control and status registers of the target vector memory access instruction, the storage address and storage data of the target vector memory access instruction are split to obtain the storage address and storage data corresponding to each target vector memory access micro-operation, including: obtaining the bit width of the target vector memory access instruction based on the configuration information of the vector control and status registers of the target vector memory access instruction.

[0089] When the vector memory access instruction address is aligned according to the memory access width, the quotient of the bit width of the target vector memory access instruction and the memory access width is calculated as the first quantity. When the vector memory access instruction address is not aligned according to the memory access width, the quotient of the bit width of the target vector memory access instruction and the memory access width plus 1 is calculated as the first quantity.

[0090] Because the address of the vector memory access instruction is not aligned according to the memory access width, the bit width of the first micro-operation operation split out cannot occupy a full memory access width, so the split memory access bit operation will be 1 more than the quotient of the bit width of the target vector memory access instruction and the memory access width.

[0091] The address of the vector memory access instruction is aligned according to the memory access width, which means that the memory access address of the vector memory access instruction is aligned on the boundary of the memory access width. For example, if the memory access width is 64 bits (8 bytes), then the memory access instruction address is aligned according to the memory access width, that is, the lower three bits of the memory access address should be all 0. For example, if the memory access width is 128 bits (16 bytes), then the memory access instruction address is aligned according to the memory access width, that is, the lower four bits of the memory access address should be all 0. For vector memory access addresses that do not meet the above rules, it is called that the vector memory access instruction address is not aligned according to the memory access width.

[0092] The storage address of the target vector memory access instruction is split into a first number of storage address groups, which are respectively used as the storage addresses corresponding to each target vector memory access micro-operation; the storage data of the target vector memory access instruction is split into a first number of storage data groups, which are respectively used as the storage data corresponding to each target vector memory access micro-operation.

[0093] For example, when the bit width of the data cache storage row is 512 bits and the bit width of the data cache storage body is 64 bits, the memory access width can be determined to be 128 bits. When the bit width of the target vector memory access instruction is 256 bits and the address of the vector memory access instruction is aligned according to the memory access width, the first number can be determined to be 2, and the target vector memory access instruction can be split into two 128-bit target vector memory access micro-operations.

[0094] Accordingly, the vector memory access instruction data fetching and storage address splitting component can split the storage address of the target vector memory access instruction into two storage address groups, which are used as storage addresses corresponding to two target vector memory access micro-operations respectively.

[0095] The vector memory access instruction storage data can split the storage data of the target vector memory access instruction into two storage data groups, which are respectively used as storage data corresponding to two target vector memory access micro-operations.

[0096] Executing each target vector memory access micro-operation includes: transmitting the storage address corresponding to each target vector memory access micro-operation to the storage instruction address pipeline, so that the storage address corresponding to each target vector memory access micro-operation can access the translation backup address buffer through the storage instruction address pipeline, and then write back the storage address corresponding to each target vector memory access micro-operation.

[0097] Write back the stored data corresponding to each target vector memory access micro-operation directly after it is emitted;

[0098] When it is determined that the storage address and storage data corresponding to each target vector memory access micro-operation are all written back, it is determined that the target vector memory access instruction completes writing back.

[0099] like Figure 3As shown, in the case where the target vector memory access instruction is a vector store instruction, after the vector memory access instruction data fetching and store address splitting component obtains the store address corresponding to each target vector memory access micro-operation, the store address corresponding to each target vector memory access micro-operation can be transmitted to the store address reservation station shared by the vector memory access instruction and the scalar memory access instruction, and then the store address corresponding to each target vector memory access micro-operation can be transmitted to the store address emission queue, so that the store address corresponding to each target vector memory access micro-operation enters the store instruction address pipeline via the above-mentioned store address emission queue.

[0100] It should be noted that the storage address issue queue in the embodiment of the present invention is shared by vector memory access instructions and scalar memory access instructions.

[0101] After the vector memory access instruction storage data obtains the storage data corresponding to each target vector memory access micro-operation, the storage data corresponding to each target vector memory access micro-operation can be transmitted to the storage data reservation station shared by the vector memory access instruction and the scalar memory access instruction, and then the storage data corresponding to each target vector memory access micro-operation can be transmitted to the storage data transmission queue.

[0102] It should be noted that the memory data emission queue in the embodiment of the present invention is shared by vector memory access instructions and scalar memory access instructions.

[0103] The storage address corresponding to each target vector memory access micro-operation accesses the translation lookaside buffer (TLB) through the storage instruction address pipeline, and the storage address corresponding to each target vector memory access micro-operation is written back, and the storage data corresponding to each target vector memory access micro-operation is directly written back after being transmitted.

[0104] It should be noted that since the vector store instruction is submitted and the data cache is written, the split store data is directly written back. After the store address and store data corresponding to each target vector memory access micro-operation after the target vector memory access instruction is split are all written back, it is determined that the target vector memory access instruction has completed writing back.

[0105] After determining that the target vector memory access instruction completes writing back, the target vector memory access instruction can be submitted in order, and the target vector memory access instruction is written into the data cache storage line after submission.

[0106] As an optional embodiment, based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status registers of the target vector memory access instruction, the target vector memory access instruction is split into multiple target vector memory access micro-operations, including: when the target vector memory access instruction is determined to be a vector fetch instruction based on the vector instruction opcode of the target vector memory access instruction and the target vector memory access instruction is for the same data cache storage line in the data cache, based on the memory access width and the configuration information of the vector control and status registers of the target vector memory access instruction, the fetch address of the target vector memory access instruction is split to obtain the fetch address corresponding to each target vector memory access micro-operation.

[0107] like Figure 3 As shown, based on the vector instruction opcode of the above-mentioned target vector memory access instruction, when it is determined that the above-mentioned target vector memory access instruction is a vector fetch instruction and the target vector memory access instruction is for the same data cache storage row in the data cache, the above-mentioned target vector memory access instruction can be transmitted to the vector memory access instruction fetch and store address splitting component.

[0108] The vector memory access instruction fetch and storage address splitting component can split the fetch address of the target vector memory access instruction based on the memory access width and the configuration information of the vector control and status register of the target vector memory access instruction to obtain the fetch address corresponding to each target vector memory access micro-operation.

[0109] As an optional embodiment, based on the memory access width and the configuration information of the vector control and status registers of the target vector memory access instruction, the data access address of the target vector memory access instruction is split to obtain the data access address corresponding to each target vector memory access micro-operation, including: obtaining the bit width of the target vector memory access instruction based on the configuration information of the vector control and status registers of the target vector memory access instruction.

[0110] When the vector memory access instruction address is aligned according to the memory access width, the quotient of the bit width of the target vector memory access instruction and the memory access width is calculated as the second quantity. When the vector memory access instruction address is not aligned according to the memory access width, the quotient of the bit width of the target vector memory access instruction and the memory access width plus 1 is calculated as the second quantity.

[0111] The data fetch address of the target vector memory access instruction is split into a second number of data fetch address groups, which are respectively used as the data fetch address corresponding to each target vector memory access micro-operation.

[0112] For example, when the bit width of the data cache storage row is 512 bits and the bit width of the data cache storage body is 64 bits, the memory access width can be determined to be 128 bits. When the bit width of the target vector memory access instruction is 256 bits and the vector memory access instruction address is aligned according to the memory access width, the second number can be determined to be 2, and the target vector memory access instruction can be split into two 128-bit target vector memory access micro-operations.

[0113] Accordingly, the vector memory access instruction fetch and store address splitting component can split the fetch address of the target vector memory access instruction into two fetch address groups, which are used as fetch addresses corresponding to two target vector memory access micro-operations respectively.

[0114] Executing each target vector memory access micro-operation includes: transmitting the data access address corresponding to each target vector memory access micro-operation to the data access instruction address pipeline, so that the data access address corresponding to each target vector memory access micro-operation can access the data cache and the translation backup address buffer through the data access instruction address pipeline, and then write back each target vector memory access micro-operation to return the data access data.

[0115] When it is determined that all target vector memory access micro-operations are written back, the result data of the target vector memory access instruction is obtained by splicing, and then it is determined that the target vector memory access instruction completes writing back.

[0116] like Figure 3 As shown, in the case where the target vector memory access instruction is a vector fetch instruction, after the vector memory access instruction fetch and storage address splitting component obtains the fetch address corresponding to each target vector memory access micro-operation, the fetch address corresponding to each target vector memory access micro-operation can be transmitted to the fetch address reservation station shared by the vector memory access instruction and the scalar memory access instruction, and then the fetch address corresponding to each target vector memory access micro-operation can be transmitted to the fetch address transmission queue, and enter the fetch instruction address pipeline through the above-mentioned fetch address queue.

[0117] The storage address corresponding to each target vector memory access micro-operation accesses the data cache (DCache) and the high-speed translation lookaside buffer (TLB) through the storage instruction address pipeline, writes back each target vector memory access micro-operation, and returns the fetched data.

[0118] After all data are written back to the target vector memory access instruction, the data fetch instruction corresponding to each target vector memory access micro-operation after the splitting has been written back and spliced ​​into the result data of the target vector memory access instruction, it is determined that the target vector memory access instruction has completed writing back.

[0119] As an optional embodiment, based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, the target scalar memory access instruction and at least one unexecuted scalar memory access instruction are merged into a target scalar memory access micro-operation, including: when it is determined that the target scalar memory access instruction is a scalar store instruction based on the scalar instruction opcode of the target scalar memory access instruction, the store address of the target scalar memory access instruction is emitted to the store instruction address pipeline, so that the store address of the target scalar memory access instruction can access the translation backup address buffer through the store instruction address pipeline, and then the store address and store data of the target scalar memory access instruction are written into the scalar store instruction write merge buffer.

[0120] Based on the scalar instruction opcode of the target scalar memory access instruction, a scalar store instruction corresponding to the target scalar memory access instruction is determined in the scalar store instruction write merging buffer, the scalar store instruction corresponding to the target scalar memory access instruction is a scalar store instruction that does not write data cache, and the scalar store instruction corresponding to the target scalar memory access instruction and the target scalar memory access instruction correspond to the same data cache storage row.

[0121] like Figure 3 As shown, when it is determined that the target scalar memory access instruction is a scalar store instruction based on the scalar instruction opcode of the scalar vector memory access instruction, the store data corresponding to the target scalar memory access instruction can be added to the store data emission queue, so that the store data corresponding to the target scalar memory access instruction enters the scalar store instruction write merge buffer via the store data emission queue.

[0122] When it is determined that the target scalar memory access instruction is a scalar store instruction based on the scalar instruction opcode of the scalar vector memory access instruction, the storage address corresponding to the target scalar memory access instruction can also be added to the storage address emission queue, so that the storage address corresponding to the target scalar memory access instruction enters the storage instruction address pipeline via the storage address emission queue.

[0123] The storage address corresponding to the target scalar memory access instruction accesses the translation lookaside buffer (TLB) through the storage instruction address pipeline and is written into the scalar storage instruction write-merge buffer.

[0124] Since the scalar store instruction is written into the data cache after submission, the store data corresponding to the scalar store instruction is directly written back, and the store address and store data corresponding to the scalar store instruction are written back before submission. Therefore, the submitted target scalar store instruction is merged with at least one scalar store instruction that has not been written into the data cache in the scalar store instruction write merge buffer (Store Buffer) into a target scalar store micro-operation.

[0125] As an optional embodiment, based on the scalar instruction opcode of the target scalar memory access instruction, a scalar store instruction corresponding to the target scalar memory access instruction is determined in a scalar store instruction write merge buffer, including: based on the scalar instruction opcode of the target scalar memory access instruction, obtaining the bit width of the target scalar memory access instruction and the target data cache storage line corresponding to the target scalar memory access instruction.

[0126] A query is made in the scalar store instruction write-merge buffer whether there is a scalar store instruction corresponding to the target data cache storage line.

[0127] When it is determined that there is a scalar store instruction corresponding to a target data cache storage line in the scalar store instruction write-merge buffer, the scalar store instruction corresponding to the target data cache storage line in the scalar store instruction write-merge buffer is determined as the scalar store instruction corresponding to the target scalar memory access instruction, and the sum of the bit width of the target scalar memory access instruction and the bit width of the scalar store instruction corresponding to the target scalar memory access instruction is not greater than the bit width of the target data cache storage line.

[0128] For example, when the bit width of the data cache storage row is 512 bits and the bit width of the target scalar memory access instruction is 64 bits, 7 scalar store instructions with a bit width of 64 bits can be determined as scalar store instructions corresponding to the target scalar memory access instruction.

[0129] Accordingly, a query is made in the scalar store instruction write-merge buffer whether there is a scalar store instruction corresponding to the target data cache storage line.

[0130] When it is determined that there are scalar store instructions corresponding to the target data cache storage lines in the scalar store instruction write-merge buffer, at most 7 scalar store instructions corresponding to the target data cache storage lines and having a bit width of 64 bits in the scalar store instruction write-merge buffer are determined as scalar store instructions corresponding to the target scalar memory access instruction.

[0131] It should be noted that, when it is determined that there is no scalar store instruction corresponding to the target data cache storage line in the scalar store instruction write-merge buffer, the target scalar memory access instruction can be directly executed.

[0132] A target scalar memory access instruction and a scalar memory access instruction corresponding to the target scalar memory access instruction are merged into a target scalar memory access micro-operation in a scalar memory access instruction write-merge buffer.

[0133] After determining the scalar store instruction corresponding to the target scalar memory access instruction in the scalar store instruction write-merge buffer, the target scalar memory access instruction and the scalar store instruction corresponding to the target scalar memory access instruction may be merged into one target scalar memory access micro-operation.

[0134] After determining the scalar store instruction corresponding to the target scalar memory access instruction, the target scalar memory access instruction and the scalar store instruction corresponding to the target scalar memory access instruction can be write-merged according to the target data cache storage line in the scalar store instruction write-merge buffer, thereby completing the target scalar memory access instruction and the scalar store instruction corresponding to the target scalar memory access instruction write data cache.

[0135] As an optional embodiment, based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, the target scalar memory access instruction and at least one unexecuted scalar memory access instruction are merged into a target scalar memory access micro-operation, including: when it is determined that the target scalar memory access instruction is a scalar data fetch instruction based on the scalar instruction opcode of the target scalar memory access instruction, the scalar data fetch instruction corresponding to the target scalar memory access instruction is determined based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, the scalar data fetch instruction corresponding to the target scalar memory access instruction is an unexecuted scalar data fetch instruction, and the scalar data fetch instruction corresponding to the target scalar memory access instruction and the target scalar memory access instruction correspond to the same data cache storage row.

[0136] like Figure 3 As shown, for the target scalar memory access instruction currently to be dispatched in the scalar memory access instruction dispatch queue, it can be determined whether the target scalar memory access instruction is a scalar store instruction or a scalar fetch instruction based on the scalar instruction opcode of the target scalar memory access instruction.

[0137] When it is determined that the target scalar memory access instruction is a scalar data fetch instruction based on the scalar instruction opcode of the scalar memory access instruction, the target scalar memory access instruction may be emitted to a scalar memory access instruction merging unit.

[0138] The scalar memory access instruction merging component can determine the scalar data fetch instruction that corresponds to the same data cache storage line as the target scalar memory access instruction and has not been executed as the scalar data fetch instruction corresponding to the target scalar memory access instruction based on the above-mentioned memory access width and the scalar instruction opcode of the target scalar memory access instruction.

[0139] As an optional embodiment, based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, a scalar data fetch instruction corresponding to the target scalar memory access instruction is determined, including: based on the scalar instruction opcode of the target scalar memory access instruction, obtaining the bit width of the target scalar memory access instruction and the target data cache storage line corresponding to the target scalar memory access instruction.

[0140] It is queried whether there is a scalar fetch instruction corresponding to the target data cache storage line.

[0141] When it is determined that there is a scalar access instruction corresponding to the target data cache storage line, the scalar memory access instruction of the target data cache storage line is determined as the scalar access instruction corresponding to the target scalar memory access instruction, and the sum of the bit width of the target scalar memory access instruction and the bit width of the scalar access instruction corresponding to the target scalar memory access instruction is not greater than the memory access width.

[0142] For example, when the bit width of the data cache storage row is 512 bits and the bit width of the data cache storage body is 64 bits, a scalar data fetch instruction with a bit width of 64 bits can be determined as the scalar data fetch instruction corresponding to the target scalar memory access instruction.

[0143] Accordingly, it is queried whether there is a scalar fetch instruction corresponding to the target data cache storage line.

[0144] When it is determined that there is a scalar fetch instruction corresponding to the target data cache storage line, at most one scalar fetch instruction corresponding to the target data cache storage line and having a bit width of 64 can be determined as the scalar fetch instruction corresponding to the target scalar memory access instruction.

[0145] The target scalar memory access instruction and the scalar data fetch instruction corresponding to the target scalar memory access instruction are merged into a target scalar memory access micro-operation.

[0146] After determining the scalar data fetch instruction corresponding to the target scalar memory access instruction, the scalar memory access instruction merging component can merge the target scalar memory access instruction and the scalar data fetch instruction corresponding to the target scalar memory access instruction into one target scalar memory access micro-operation.

[0147] It should be noted that, when it is determined that there is no corresponding scalar data fetch instruction that can be merged with the target scalar memory access instruction, the target scalar memory access instruction can be directly executed.

[0148] Executing a target scalar memory access micro-operation includes: emitting the target scalar memory access micro-operation to a fetch instruction pipeline, so that after the target scalar memory access micro-operation accesses a data cache through the fetch instruction pipeline and accesses the data cache through a translation backup address buffer, the target scalar memory access instruction and the scalar fetch instruction corresponding to the target scalar memory access instruction are written back respectively.

[0149] like Figure 3 As shown, in the case where the target scalar memory access instruction is a scalar fetch instruction, the scalar memory access instruction merging component merges the target scalar memory access instruction and the scalar fetch instruction corresponding to the target scalar memory access instruction into a target scalar memory access micro-operation, and then the above target scalar memory access micro-operation can be emitted to the fetch instruction pipeline.

[0150] The target scalar memory access micro-operation accesses the data cache (DCache) and the translation lookaside buffer (TLB) through the fetch instruction pipeline. The fetch instruction pipeline has multiple write-back ports. After the target scalar memory access micro-operation accesses the data cache, the target scalar memory access instruction and the scalar fetch instruction corresponding to the target scalar memory access instruction are written back respectively.

[0151] The memory access instruction processing method provided by the present invention uniformly processes vector memory access instructions and scalar memory access instructions inside a processor, thereby improving the processing efficiency and resource utilization of the processor, fully ensuring the efficiency of large data bit width in vector memory access instruction processing, while avoiding the waste of memory access bandwidth in scalar memory access instruction processing, and fully considering the characteristics of the data cache storage body, designing a reasonable memory access width for data cache, splitting vector memory access instructions according to the above memory access width, and merging scalar memory access instructions according to the above memory access width, thereby realizing maximum resource sharing of vector memory access and scalar memory access.

[0152] Figure 4 Schematic diagram of the structure of the processor provided by the present invention. Figure 4 The processor provided by the present invention is described, and the processor described below and the memory access instruction processing method provided by the present invention described above can be referred to each other. Figure 4 As shown, the processor includes: a vector memory access instruction splitting unit 401, a scalar memory access instruction merging unit 402 and a memory access execution unit 403.

[0153] The vector memory access instruction splitting component 401 is used for splitting the target vector memory access instruction to be dispatched in the vector memory access instruction dispatch queue into a plurality of target vector memory access micro-operations based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction.

[0154] A scalar memory access instruction merging unit 402 is used to merge the target scalar memory access instruction to be dispatched currently in the scalar memory access instruction dispatch queue with at least one unexecuted scalar memory access instruction into a target scalar memory access micro-operation based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction;

[0155] The memory access execution unit 403 is used to execute each target vector memory access micro-operation and / or execute a target scalar memory access micro-operation;

[0156] The memory access width is determined based on the number of data cache memory banks in the data cache in the processor and the bit width of the data cache memory banks.

[0157] The processor in the embodiment of the present invention, for the target vector memory access instruction currently to be dispatched in the vector memory access instruction dispatch queue, splits the target vector memory access instruction into multiple target vector memory access micro-operations based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction, and then executes each target vector memory access micro-operation to complete the target vector memory access instruction. For the target scalar memory access instruction currently to be dispatched in the scalar memory access instruction dispatch queue, based on the scalar instruction opcode of the target scalar memory access instruction, the target scalar memory access instruction is combined with at least one unexecuted scalar memory access instruction. After the memory access instructions are merged into a target scalar memory access micro-operation, the above target scalar memory access micro-operation is executed to complete the above target scalar memory access instruction. The vector memory access instructions can be split and the scalar memory access instructions can be merged based on the memory access width determined according to the number of data cache storage bodies in the data cache in the processor and the bit width of the data cache storage body. Furthermore, through the splitting of vector memory access instructions and the merging of scalar memory access instructions, resources can be more reasonably allocated to vector memory access and scalar memory access during data memory access, thereby realizing the maximum resource sharing of vector memory access and scalar memory access, which can improve the efficiency and resource utilization of the processor and has broad application prospects.

[0158] The present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the memory access instruction processing method provided by the above methods, the method including: for the target vector memory access instruction currently to be dispatched in the vector memory access instruction dispatch queue, based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction, split the target vector memory access instruction into multiple target vector memory access micro-operations; for the target scalar memory access instruction currently to be dispatched in the scalar memory access instruction dispatch queue, based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, merge the target scalar memory access instruction and at least one unexecuted scalar memory access instruction into a target scalar memory access micro-operation; execute each target vector memory access micro-operation, and / or execute the target scalar memory access micro-operation; wherein the memory access width is determined based on the number of data cache storage bodies in the data cache in the processor and the bit width of the data cache storage body.

[0159] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the memory access instruction processing method provided by the above-mentioned methods, the method comprising: for the target vector memory access instruction currently to be dispatched in the vector memory access instruction dispatch queue, based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction, splitting the target vector memory access instruction into multiple target vector memory access micro-operations; for the target scalar memory access instruction currently to be dispatched in the scalar memory access instruction dispatch queue, based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, merging the target scalar memory access instruction and at least one unexecuted scalar memory access instruction into one target scalar memory access micro-operation; executing each target vector memory access micro-operation, and / or executing the target scalar memory access micro-operation; wherein the memory access width is determined based on the number of data cache storage bodies in the data cache of the processor and the bit width of the data cache storage body.

[0160] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.

[0161] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A memory access instruction processing method, characterized in that: Applied to a processor, the method comprises: For a target vector memory access instruction to be dispatched currently in a vector memory access instruction dispatch queue, based on a memory access width, a vector instruction opcode of the target vector memory access instruction, and configuration information of a vector control and status register of the target vector memory access instruction, split the target vector memory access instruction into a plurality of target vector memory access micro-operations, For a target scalar memory access instruction to be dispatched currently in the scalar memory access instruction dispatch queue, based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, merging the target scalar memory access instruction and at least one unexecuted scalar memory access instruction into a target scalar memory access micro-operation; Execute each of the target vector memory access micro-operations, and / or execute the target scalar memory access micro-operations; The memory access width is determined based on the number of data cache memory banks in the data cache of the processor and the bit width of the data cache memory bank; the memory access width is a positive integer multiple of the bit width of the data cache memory bank and the memory access width is not greater than the product of the number of the data cache memory banks and the bit width of the data cache memory bank; The method of splitting the target vector memory access instruction into a plurality of target vector memory access micro-operations based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction comprises: When it is determined based on the vector instruction opcode of the target vector memory access instruction that the target vector memory access instruction is a vector store instruction and the target vector memory access instruction is for the same data cache storage row in the data cache, the store address and store data of the target vector memory access instruction are split based on the memory access width and the configuration information of the vector control and status register of the target vector memory access instruction to obtain the store address and store data corresponding to each of the target vector memory access micro-operations.

2. The memory access instruction processing method according to claim 1, characterized in that: The executing each of the target vector memory access micro-operations comprises: The storage address corresponding to each of the target vector memory access micro-operations is transmitted to the storage instruction address pipeline, so that the storage address corresponding to each of the target vector memory access micro-operations can access the translation backup address buffer through the storage instruction address pipeline, and then the storage address corresponding to each of the target vector memory access micro-operations is written back; Writing back directly the stored data corresponding to each target vector memory access micro-operation after transmitting; When it is determined that the storage address and storage data corresponding to each of the target vector memory access micro-operations are all written back, it is determined that the target vector memory access instruction completes writing back.

3. The memory access instruction processing method according to claim 1, characterized in that: The method of splitting the target vector memory access instruction into a plurality of target vector memory access micro-operations based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction comprises: In a case where it is determined based on the vector instruction opcode of the target vector memory access instruction that the target vector memory access instruction is a vector fetch instruction and the target vector memory access instruction is for a same data cache storage line in a data cache, based on the memory access width and configuration information of the vector control and status register of the target vector memory access instruction, the fetch address of the target vector memory access instruction is split to obtain a fetch address corresponding to each of the target vector memory access micro-operations; The executing each of the target vector memory access micro-operations comprises: The data fetch address corresponding to each of the target vector memory access micro-operations is transmitted to the data fetch instruction address pipeline, so that the data fetch address corresponding to each of the target vector memory access micro-operations can access the data cache and the translation backup address buffer through the data fetch instruction address pipeline, and then each of the target vector memory access micro-operations is written back to return the data fetch data; When it is determined that all the target vector memory access micro-operations are written back, the result data of the target vector memory access instruction is obtained by splicing, and then it is determined that the target vector memory access instruction completes writing back.

4. The memory access instruction processing method according to claim 1, characterized in that: The step of combining the target scalar memory access instruction and at least one unexecuted scalar memory access instruction into a target scalar memory access micro-operation based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction comprises: In a case where it is determined based on the scalar instruction opcode of the target scalar memory access instruction that the target scalar memory access instruction is a scalar data fetch instruction, a scalar data fetch instruction corresponding to the target scalar memory access instruction is determined based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction, the scalar data fetch instruction corresponding to the target scalar memory access instruction is an unexecuted scalar data fetch instruction, and the scalar data fetch instruction corresponding to the target scalar memory access instruction corresponds to the same data cache storage line as the target scalar memory access instruction; Merging the target scalar memory access instruction and the scalar data fetch instruction corresponding to the target scalar memory access instruction into one target scalar memory access micro-operation; The executing the target scalar memory access micro-operation includes: The target scalar memory access micro-operation is emitted to the data fetch instruction pipeline, so that after the target scalar memory access micro-operation accesses the data cache through the data fetch instruction pipeline and the translation backup address buffer accesses the data cache, the target scalar memory access instruction and the scalar fetch instruction corresponding to the target scalar memory access instruction are written back respectively.

5. The memory access instruction processing method according to claim 1, characterized in that: The step of combining the target scalar memory access instruction and at least one unexecuted scalar memory access instruction into a target scalar memory access micro-operation based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction comprises: In the case where it is determined based on the scalar instruction opcode of the target scalar memory access instruction that the target scalar memory access instruction is a scalar store instruction, the store address of the target scalar memory access instruction is transmitted to the store instruction address pipeline, so that the store address of the target scalar memory access instruction accesses the translation backup address buffer through the store instruction address pipeline, and then the store address and store data of the target scalar memory access instruction are written into the scalar store instruction write-merge buffer; Based on the scalar instruction opcode of the target scalar memory access instruction, determining in the scalar store instruction write-merge buffer a scalar store instruction corresponding to the target scalar memory access instruction, the scalar store instruction corresponding to the target scalar memory access instruction is a scalar store instruction that does not write data cache, and the scalar store instruction corresponding to the target scalar memory access instruction and the target scalar memory access instruction correspond to the same data cache storage line; The target scalar memory access instruction and the scalar store instruction corresponding to the target scalar memory access instruction are merged into one target scalar memory access micro-operation in the scalar store instruction write-merge buffer.

6. The memory access instruction processing method according to claim 2, characterized in that: The method of splitting the storage address and storage data of the target vector memory access instruction based on the memory access width and the configuration information of the vector control and status register of the target vector memory access instruction to obtain the storage address and storage data corresponding to each of the target vector memory access micro-operations includes: Acquire the bit width of the target vector memory access instruction based on configuration information of a vector control and status register of the target vector memory access instruction; When the vector memory access instruction address is aligned according to the memory access width, the quotient of the bit width of the target vector memory access instruction and the memory access width is calculated as the first quantity; when the vector memory access instruction address is not aligned according to the memory access width, the quotient of the bit width of the target vector memory access instruction and the memory access width plus 1 is calculated as the first quantity; The storage address of the target vector memory access instruction is split into the first number of storage address groups, which are respectively used as the storage addresses corresponding to each of the target vector memory access micro-operations; the storage data of the target vector memory access instruction is split into the first number of storage data groups, which are respectively used as the storage data corresponding to each of the target vector memory access micro-operations.

7. The memory access instruction processing method according to claim 3, characterized in that: The step of splitting the data access address of the target vector memory access instruction based on the memory access width and the configuration information of the vector control and status register of the target vector memory access instruction to obtain the data access address corresponding to each of the target vector memory access micro-operations includes: In the case where the vector memory access instruction address is aligned according to the memory access width, the bit width of the target vector memory access instruction is obtained based on the configuration information of the vector control and status register of the target vector memory access instruction; in the case where the vector memory access instruction address is not aligned according to the memory access width, the quotient of the bit width of the target vector memory access instruction and the memory access width plus 1 is calculated as the second quantity; The data fetch address of the target vector memory access instruction is split into the second number of data fetch address groups, which are respectively used as the data fetch address corresponding to each of the target vector memory access micro-operations.

8. The memory access instruction processing method according to claim 4, characterized in that: The step of determining the scalar data fetch instruction corresponding to the target scalar memory access instruction based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction comprises: Based on a scalar instruction opcode of the target scalar memory access instruction, obtaining a bit width of the target scalar memory access instruction and a target data cache storage line corresponding to the target scalar memory access instruction; querying whether there is a scalar fetch instruction corresponding to the target data cache storage line; When it is determined that there is a scalar data access instruction corresponding to the target data cache storage line, the scalar memory access instruction of the target data cache storage line is determined to be the scalar data access instruction corresponding to the target scalar memory access instruction, and the sum of the bit width of the target scalar memory access instruction and the bit width of the scalar data access instruction corresponding to the target scalar memory access instruction is not greater than the memory access width.

9. The memory access instruction processing method according to claim 5, characterized in that: The step of determining, based on the scalar instruction opcode of the target scalar memory access instruction, the scalar store instruction corresponding to the target scalar memory access instruction in the scalar store instruction write-merge buffer includes: Based on a scalar instruction opcode of the target scalar memory access instruction, obtaining a bit width of the target scalar memory access instruction and a target data cache storage line corresponding to the target scalar memory access instruction; querying in the scalar store instruction write-merge buffer whether there is a scalar store instruction corresponding to the target data cache storage line; When it is determined that there is a scalar store instruction corresponding to the target data cache storage line in the scalar store instruction write-merge buffer, the scalar store instruction corresponding to the target data cache storage line in the scalar store instruction write-merge buffer is determined to be the scalar store instruction corresponding to the target scalar memory access instruction, and the sum of the bit width of the target scalar memory access instruction and the bit width of the scalar store instruction corresponding to the target scalar memory access instruction is not greater than the bit width of the target data cache storage line.

10. A processor, characterized in that: include: a vector memory access instruction splitting component, for a target vector memory access instruction to be dispatched in a vector memory access instruction dispatch queue, based on a memory access width, a vector instruction opcode of the target vector memory access instruction, and configuration information of a vector control and status register of the target vector memory access instruction, splitting the target vector memory access instruction into a plurality of target vector memory access micro-operations; a scalar memory access instruction merging component, configured to merge, for a target scalar memory access instruction currently to be dispatched in a scalar memory access instruction dispatch queue, the target scalar memory access instruction and at least one unexecuted scalar memory access instruction into a target scalar memory access micro-operation based on the memory access width and the scalar instruction opcode of the target scalar memory access instruction; A memory access execution unit, configured to execute each of the target vector memory access micro-operations, and / or execute the target scalar memory access micro-operations; The memory access width is determined based on the number of data cache memory banks in the data cache of the processor and the bit width of the data cache memory bank; the memory access width is a positive integer multiple of the bit width of the data cache memory bank and the memory access width is not greater than the product of the number of the data cache memory banks and the bit width of the data cache memory bank; The vector memory access instruction splitting component splits the target vector memory access instruction into a plurality of target vector memory access micro-operations based on the memory access width, the vector instruction opcode of the target vector memory access instruction, and the configuration information of the vector control and status register of the target vector memory access instruction, including: When it is determined based on the vector instruction opcode of the target vector memory access instruction that the target vector memory access instruction is a vector store instruction and the target vector memory access instruction is for the same data cache storage row in the data cache, the store address and store data of the target vector memory access instruction are split based on the memory access width and the configuration information of the vector control and status register of the target vector memory access instruction to obtain the store address and store data corresponding to each of the target vector memory access micro-operations.

Citation Information

Patent Citations

  • Access optimization compiling method and device for functions

    CN106201641A

  • Memory access method, processor, electronic equipment and readable storage medium

    CN116932202A

  • Method and apparatus for data processing

    US7305540B1