Instruction processing method, processor and electronic device

By providing an operation mask for the processor's instructions, keeping the high-bit data unchanged when executing instructions that only process part of the register bits, the problems of high-bit data protection and storage access delay in the prior art are solved, and the execution efficiency and system performance of the processor are improved.

CN120066580APending Publication Date: 2025-05-30HYGON YUNXIN INTEGRATED CIRCUIT DESIGN (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510183504.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When executing instructions in registers with bit widths of more than 128 bits, existing processors need to pour high-bit data into internal memory to protect data, resulting in complex judgment logic, occupancy of storage space, increasing storage access delay and power consumption.

Method used

An instruction processing method is provided to obtain a first instruction in the first instruction set, which can only process some bits in the first destination register and provide an operation mask to keep the value of the target bit unchanged when the instruction is executed, and avoid being overwritten by the calculation result.

Benefits of technology

By providing an operation mask, the target bit value in the register is avoided from being overwritten, the dependence on internal memory is reduced, the memory access delay and power consumption is reduced, and the instruction execution efficiency and system performance are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066580A_ABST
    Figure CN120066580A_ABST
Patent Text Reader

Abstract

The invention discloses an instruction processing method, a processor and an electronic device. The instruction processing method comprises the steps that a first instruction in a first instruction set is obtained, the first instruction is configured to only process part of bits in a first destination register, and a first destination operand of the first instruction corresponds to part of bits in the first destination register; providing a first operation mask for the first instruction in response to the fact that the value of the target bit except the partial bit in the first destination register is not zero, the first operation mask being used for keeping the value of the target bit in the first destination register unchanged when the first instruction is executed; the first instruction is executed using a portion of the bits in the first destination register. The instruction processing method can improve instruction processing efficiency, reduce power consumption and improve system performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to an instruction processing method, a processor, and an electronic device. Background Art

[0002] In the field of computer technology, as a core component, a processor is responsible for executing various computing tasks. An instruction set architecture (ISA) defines the basic functional characteristics of a processor, such as specifying the instruction formats, operation codes, and data types that the processor can execute. Instructions are the bridge between the processor and software. By parsing and executing instructions, the processor can achieve specific functions, such as reading and writing data or performing calculations on data. An instruction set is a set of specific instructions determined based on the instruction set architecture. The architecture registers that can be operated by instructions of different instruction sets are usually different.

[0003] As a key component of the instruction set architecture, an architecture register is a high-speed storage unit inside the processor. The architecture register is used to temporarily store data during the execution of instructions, which can reduce the overhead of the processor frequently accessing memory and improve the operation speed. Instructions in the instruction set can operate on registers. For example, a data transfer instruction can transfer data between a register and memory or other registers, and an arithmetic instruction can obtain data from a register for calculation and store the calculation result in the register. Summary of the Invention

[0004] At least one embodiment of the present disclosure provides an instruction processing method, which includes: obtaining a first instruction in a first instruction set, where the first instruction is configured to only process partial bits in a first destination register, and a first destination operand of the first instruction corresponds to the partial bits in the first destination register; in response to the value of target bits other than the partial bits in the first destination register not being zero, providing a first operation mask for the first instruction, where the first operation mask is used to keep the value of the target bits in the first destination register unchanged when executing the first instruction; and executing the first instruction using the partial bits in the first destination register.

[0005] For example, the instruction processing method provided by at least one embodiment of the present disclosure further includes:

[0006] in response to the value of the target bits in the first destination register being zero, directly executing the first instruction using the partial bits in the first destination register.

[0007] For example, in the instruction processing method provided by at least one embodiment of the present disclosure, providing the first operation mask for the first instruction includes: using an operation mask providing unit configured to support the execution of instructions in a second instruction set to provide the first operation mask for the first instruction, where at least one instruction in the second instruction set is configured to be able to process more bits including the partial bits in the first destination register.

[0008] For example, in the instruction processing method provided by at least one embodiment of the present disclosure, the operation mask providing unit includes a mask register, and the mask register is configured to provide a value of a second operation mask to a second instruction in the second instruction set, and the value of the second operation mask is used to execute the second instruction.

[0009] For example, in the instruction processing method provided by at least one embodiment of the present disclosure, when the second destination operand of the second instruction corresponds to the first destination register, the instruction processing method further includes: in response to the i-th bit of the second operation mask being a first value, keeping the value corresponding to the i-th bit of the second operation mask stored in the first destination register unchanged or clearing it; or, in response to the i-th bit of the second operation mask being a second value, storing the calculation result of the i-th element in the second instruction into the first destination register, where i is less than or equal to the bit width of the first destination register.

[0010] For example, in the instruction processing method provided by at least one embodiment of the present disclosure, the first instruction set includes a Streaming SIMD Extensions (SSE) instruction set, an Advanced Vector Extensions (AVX) instruction set, and an Advanced Vector Extensions 2 (AVX2) instruction set, and the second instruction set includes an Advanced Vector Extensions 512 (AVX512) instruction set.

[0011] For example, in the instruction processing method provided by at least one embodiment of the present disclosure, when the first destination register is 256 bits, the partial bits include bits [127:0], and the target bits include bits [255:128]; or, when the first destination register is 512 bits, the partial bits include bits [127:0], and the target bits include bits [511:128].

[0012] For example, in the instruction processing method provided by at least one embodiment of the present disclosure, the first instruction is used to perform a calculation operation on a vector or a scalar.

[0013] At least one embodiment of the present disclosure further provides a processor, including: a receiving unit configured to obtain a first instruction in a first instruction set, wherein the first instruction is configured to process only a part of bits in a first destination register, and a first destination operand of the first instruction corresponds to the part of bits in the first destination register; an operation mask providing unit configured to provide a first operation mask for the first instruction in response to that values of target bits other than the part of bits in the first destination register are not zero; and an execution unit configured to execute the first instruction using the part of bits in the first destination register, and keep the values of the target bits in the first destination register unchanged in response to the first operation mask being provided to the first instruction.

[0014] For example, in the processor provided by at least one embodiment of the present disclosure, the execution unit is further configured to: directly execute the first instruction using the part of bits in the first destination register in response to that the values of the target bits in the first destination register are zero.

[0015] For example, in the processor provided by at least one embodiment of the present disclosure, the operation mask providing unit is further configured to support running instructions in a second instruction set, wherein at least one instruction in the second instruction set is configured to process more bits including the part of bits in the first destination register.

[0016] For example, in the processor provided by at least one embodiment of the present disclosure, the operation mask providing unit includes a mask register configured to provide a value of a second operation mask to a second instruction in the second instruction set, and the value of the second operation mask is used to execute the second instruction.

[0017] For example, in the processor provided by at least one embodiment of the present disclosure, the receiving unit is further configured to obtain the second instruction in the second instruction set, the second instruction is executed before or after the first instruction, and a second destination operand of the second instruction corresponds to the first destination register, and the execution unit is further configured to: keep the value stored in the first destination register corresponding to the i-th bit of the second operation mask unchanged or cleared in response to that the i-th bit of the second operation mask is a first value; or store a calculation result of the i-th element in the second instruction into the first destination register in response to that the i-th bit of the second operation mask is a second value, where i is less than or equal to a bit width of the first destination register.

[0018] For example, in the processor provided by at least one embodiment of the present disclosure, the first instruction set includes a Streaming SIMD Extensions (SSE) instruction set, an Advanced Vector Extensions (AVX) instruction set, and an Advanced Vector Extensions 2 (AVX2) instruction set, and the second instruction set includes an Advanced Vector Extensions 512 (AVX512) instruction set.

[0019] For example, in the processor provided by at least one embodiment of the present disclosure, when the first register is 256 bits, the partial bits include bits [127:0], and the target bits include bits [255:128]; or, when the first register is 512 bits, the partial bits include bits [127:0], and the target bits include bits [511:128].

[0020] At least one embodiment of the present disclosure further provides an electronic device, including the processor provided by any one of the above embodiments. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0022] Figure 1 It is a schematic diagram of the execution process of a first instruction;

[0023] Figure 2 It is a flowchart of the instruction processing method provided by at least one embodiment of the present disclosure;

[0024] Figure 3 It is a schematic diagram of the execution process of a first instruction provided by at least one embodiment of the present disclosure;

[0025] Figure 4 It is a schematic diagram of the execution process of another first instruction provided by at least one embodiment of the present disclosure;

[0026] Figure 5 It is a schematic diagram of the execution process of a second instruction provided by at least one embodiment of the present disclosure;

[0027] Figure 6 It is a schematic block diagram of a processor provided by at least one embodiment of the present disclosure;

[0028] Figure 7 It is a schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure; and

[0029] Figure 8 It is a schematic structural diagram of an electronic device provided by at least one embodiment of the present disclosure. Detailed Embodiments

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions of the embodiments of the present disclosure with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of them. Based on the described embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.

[0031] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. It should be understood that the various steps recorded in the method embodiments of the present disclosure may be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0032] The following will illustrate the present disclosure through several specific embodiments. To keep the following description of the embodiments of the present disclosure clear and concise, the detailed descriptions of known functions and known components (elements) may be omitted. When any component (element) of the embodiments of the present disclosure appears in more than one drawing, the component (element) is denoted by the same or similar reference numerals in each drawing.

[0033] If the instruction set is backward compatible, enabling legacy software to run directly on newly designed processor chips without the need for recompilation, then processors using that instruction set will have good market stability.

[0034] Typical instruction sets include multiple subsets and extended instruction sets to achieve different functions or different hardware specifications. For example, the X86 instruction set includes multiple subsets and extended instruction sets, such as the Streaming SIMD Extensions (SSE) instruction set, the Advanced Vector Extensions (AVX) instruction set, the Advanced Vector Extensions 2 (AVX2) instruction set, the Advanced Vector Extensions 512 (AVX512) instruction set, etc. These instruction sets have different bit widths. For example, the SSE instruction set has a bit width of 128 bits, the AVX instruction set and the AVX2 instruction set have a bit width of 256 bits, and the AVX512 instruction set has a bit width of 512 bits.

[0035] These instruction sets respectively correspond to multiple architecture registers, such as XMM registers, YMM registers, and ZMM registers, etc. The XMM register is a 128-bit register used by the SSE instruction set, which can process 4 single-precision floating-point numbers or 2 double-precision floating-point numbers simultaneously. When performing vector operations, it can operate on multiple data elements in parallel, improving the efficiency of floating-point operations. The YMM register is a 256-bit register used by the AVX instruction set, which is an extension based on the XMM register. It can accommodate 8 single-precision floating-point numbers or 4 double-precision floating-point numbers, supporting more complex vector calculations and higher data parallelism, and can further improve computing performance. The ZMM register is a 512-bit register introduced with the AVX512 instruction set, with a wider data bit width. It can process more data elements simultaneously, and can accommodate 16 single-precision floating-point numbers or 8 double-precision floating-point numbers, greatly improving the throughput of data processing and the parallel computing ability.

[0036] The X86 instruction set has backward compatibility. For example, the lower 128 bits of the YMM register can be multiplexed with the XMM register. Therefore, on a processor that supports the AVX instruction set, SSE instructions can still operate using the XMM register. Similarly, on a processor that supports the AVX512 instruction set, SSE instructions and AVX instructions can still operate on the XMM register and the YMM register.

[0037] The inventors of the present disclosure have noticed that although the X86 instruction set has backward compatibility, when registers with a bit width exceeding 128 bits (e.g., YMM registers or ZMM registers) encounter instructions that can only process 128-bit data, the data in the parts of the registers that cannot be operated on by the instructions needs to be transferred to the internal memory first to protect this data from being damaged. For example, when executing a 128-bit addition instruction, the addition operation may generate a carry. If the high 128-bit data is not saved first, the carry may overwrite and damage the high 128-bit data, resulting in errors when other instructions (e.g., instructions that can operate on all bits in the register) use this register subsequently. Here, the method of transferring the high-bit data in the register to the internal memory can be called the spill-load mechanism. Through this mechanism, it can be ensured that after executing the first instruction, the high-bit data in the register remains unchanged.

[0038] The following will be combined with Figure 1 to introduce the above-mentioned spill-load mechanism in detail.

[0039] Figure 1 It is a schematic diagram of the execution process of a first instruction. For example, the first instruction is an addition instruction in the SSE instruction set, which can add the first destination operand and the first source operand and store the first calculation result obtained by the addition in the first destination register. As Figure 1 shown, the black box represents the first destination register. The first destination register is, for example, a YMM register with a bit width of 256 bits. Before executing the first instruction, the first destination register stores the values of 4 elements, which are a1, a2, a3, and a4 respectively. Among them, element a1 and element a2 are stored in the low bits (e.g., the low 128 bits) of the first destination register, and element a3 and element a4 are stored in the high bits (e.g., the high 128 bits) of the first destination register. For example, the values of these 4 elements are all not 0.

[0040] The first instruction can only operate on the low bits (e.g., the low 128 bits) in the register and cannot operate on the high bits (e.g., the high 128 bits) in the register. For example, the data to be operated on by the first instruction is 128 bits (e.g., the first destination operand), which only corresponds to the low bits in the first destination register. Therefore, before executing the first instruction, it is necessary to first spill all the high-bit data in the first destination register to the internal memory (e.g., the memory) so that the data in the high bits of the first destination register is cleared to 0, and then the first instruction can be executed. For example, as Figure 1As shown, before executing the first instruction, it is necessary to first pour out the elements a3 and a4 stored in the high bits of the first destination register into the internal memory M, so that the data stored in the high bits of the first destination register is 0, and then execute the first instruction. The values of the elements a1 and a2 in the first destination operand are added to the values of the elements b1 and b2 in the first source operand respectively, and the first calculation result obtained includes the elements c1 and c2, where c1 = a1 + b1 and c2 = a2 + b2. For example, the element c1 is stored in bits [63:0] of the first destination register, and the element c2 is stored in bits [127:64] of the first destination register.

[0041] After executing the first instruction, if the next instruction that needs to operate on this first destination register is still an instruction that can only operate on the lower 128 bits of the first destination register (for example, an instruction in the SSE instruction set), then the next instruction can be directly executed. Since the execution of this kind of instruction only involves the low-bit data of the first destination register and does not involve the high-bit data, the high bits of the first destination register remaining 0 do not affect the execution of this instruction.

[0042] However, after executing the first instruction, if the next instruction that needs to operate on this first destination register is an AVX instruction or other instruction that can operate on the high bits of the first destination register, then it is necessary to load the data previously poured out into the internal memory M into the high bits of the first destination register before the next instruction can be executed. Since this kind of instruction may need to use the data originally stored in the high bits of the first destination register, it is necessary to load the previously poured out data, such as the values of the elements a3 and a4, back into bits [191:128] and bits [192:255] of the first destination register.

[0043] The inventors of the present disclosure also noticed that although this pour-out and load mechanism can ensure that the high-bit data of the first destination register is not affected by the first instruction, its judgment logic is complex, and the poured-out data will occupy the storage space of the internal memory. If such pour-out operations are performed frequently, it may compete for bandwidth with other ongoing storage access operations, resulting in an increase in the latency of other operations, thereby affecting the operating efficiency of the entire system; in addition, the processes of pour-out and load take a lot of time and power consumption. Since the speed of accessing data from the internal memory is slower than directly accessing data from the first destination register, this increase in storage access latency will lead to an extension of the execution time of the instruction, thereby reducing the overall performance of the system.

[0044] To solve the above problems, at least one embodiment of the present disclosure provides an instruction processing method, a processor, and an electronic device. The instruction processing method includes: obtaining a first instruction in a first instruction set, where the first instruction is configured to process only a part of bits in a first destination register, and a first destination operand of the first instruction corresponds to a part of bits in the first destination register; in response to a value of a target bit other than the part of bits in the first destination register not being zero, providing a first operation mask for the first instruction, where the first operation mask is used to keep the value of the target bit in the first destination register unchanged when the first instruction is executed; and executing the first instruction using the part of bits in the first destination register.

[0045] The instruction processing method provided in the above embodiment of the present disclosure provides a first operation mask for the first instruction, so that when the first instruction is executed, the value of the target bit in the first destination register remains unchanged. It can avoid the value of the target bit in the first destination register being overwritten or damaged by the calculation result of the first instruction, and does not occupy the storage space of the internal memory. The subsequent executed instructions can directly operate on the data stored in the first destination register without accessing data from the internal memory, thereby reducing the storage access latency, improving the execution efficiency of the instructions, reducing power consumption, and enhancing the system performance.

[0046] Some embodiments of the present disclosure and their examples will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining and illustrating the present disclosure, and are not used to limit the present disclosure.

[0047] Figure 2 It is a flowchart of the instruction processing method provided by at least one embodiment of the present disclosure. For example, as Figure 2 shown, the instruction processing method includes the following steps S100 to step S300.

[0048] Step S100: Obtain a first instruction in a first instruction set, where the first instruction is configured to process only a part of bits in a first destination register, and a first destination operand of the first instruction corresponds to a part of bits in the first destination register.

[0049] Step S200: In response to a value of a target bit other than the part of bits in the first destination register not being zero, provide a first operation mask for the first instruction, where the first operation mask is used to keep the value of the target bit in the first destination register unchanged when the first instruction is executed.

[0050] Step S300: Execute the first instruction using the part of bits in the first destination register.

[0051] In some embodiments of the present disclosure, the first instruction set may be a subset or an extended instruction set in an instruction set architecture, such as the X86 instruction set architecture, or other instruction set architectures including subsets with various different bit widths, such as the ARM instruction set architecture. Embodiments of the present disclosure do not limit the architecture of the first instruction set.

[0052] The first instruction set may include various types of instructions, such as data transfer instructions and data processing instructions. For example, data transfer instructions include, but are not limited to, move instructions, exchange instructions, etc., and data processing instructions include, but are not limited to, arithmetic operation instructions (such as addition instructions, subtraction instructions, multiplication instructions, or division instructions, etc.), logical operation instructions (such as AND operation instructions, OR operation instructions, XOR operation instructions, etc.), shift instructions (such as arithmetic left shift instructions, arithmetic right shift instructions, logical left shift instructions, or logical right shift instructions, etc.). For example, the first instruction in the first instruction set may be any one of these types of instructions. Embodiments of the present disclosure do not limit the type of the first instruction.

[0053] For example, in at least one embodiment of the present disclosure, the first instruction may be used to perform calculation operations on vectors or scalars. Embodiments of the present disclosure do not limit this.

[0054] The bit width of a register refers to the number of bits of binary data that the register can store, which determines the data range and precision that the register can represent. The bit width of the register may be 8 bits, 16 bits, 32 bits, 64 bits, 128 bits, 256 bits, 516 bits, or a larger bit width. Among them, registers with 128 bits, 256 bits, and 512 bits can process multiple data elements simultaneously and are usually used in high-performance computing, which can greatly improve the computing efficiency. For example, the bit width of the first destination register may be 128 bits, 256 bits, 512 bits, or a larger or smaller bit width. Embodiments of the present disclosure do not limit the bit width size of the first destination register.

[0055] The bit width of an operand refers to the number of bits of binary data included in the operand, which determines the data range and precision that the operand can represent. For example, an integer with a bit width of 8 bits can represent values between 0 and 255 (unsigned number) or -128 and 127 (signed number). For example, single-precision floating-point numbers are usually 32 bits, and double-precision floating-point numbers are usually 64 bits. Floating-point numbers with different bit widths can represent decimals with different precisions. Here, data such as an integer, a single-precision floating-point number, or a double-precision floating-point number is called an element.

[0056] In an embodiment of the present disclosure, the first destination register may store one or more elements. For example, when the bit width of the first destination register is 256 bits, it may store 4 64-bit elements, 8 32-bit elements, 16 16-bit elements, or 32 8-bit elements.

[0057] For ease of description, in the examples provided below, an example in which the bit width of the first destination register is 256 bits and it can store 4 64-bit elements is used to illustrate the embodiments of the present disclosure. However, this is not a limitation on the present disclosure.

[0058] In step S100, the first instruction obtained is an instruction that can only process a part of the bits of the first destination register. For example, before executing the first instruction, the first destination operand of the first instruction corresponds to a part of the bits stored in the first destination register, or after executing the first instruction, the first calculation result of the first instruction is correspondingly stored in a part of the bits in the first destination register. In an embodiment of the present disclosure, all the bits of the first destination operand of the first instruction correspond to a certain part of the bits of the first destination register. For example, this part of the bits may be the lower bits of the first destination register.

[0059] In an embodiment of the present disclosure, the target bits may be all the bits of the first destination register except for a part of the bits. For example, in one example, the first destination register is 256 bits, then the part of the bits may be bits [127:0], and the target bits are bits [255:128]. For example, in another example, the first destination register is 512 bits, then the part of the bits may be bits [127:0], and the target bits are bits [511:128].

[0060] In step S200, in response to the value of the target bits of the first destination register except for a part of the bits not being zero, a first operation mask is provided for the first instruction. For example, in one example, the first destination register is 256 bits, the part of the bits includes the lower 128 bits, and the target bits include the upper 128 bits. If any one of the upper 128 bits in the first destination register is not zero, it indicates that there is high-order data stored in the first destination register. When executing the first instruction, it is necessary to ensure that the high-order data does not change. Therefore, in this case, it is necessary to provide a first operation mask for the first instruction. Through the first operation mask, when executing the first instruction, the values of the target bits in the first destination register are all kept unchanged.

[0061] For example, the bit width of the first operation mask is the same as that of the first destination register, and each bit of the first operation mask corresponds to each bit of the first destination register one by one. For example, the value of each bit of the first operation mask can be set to a first value or a second value. The first value indicates that the value of the corresponding bit on the first destination register remains unchanged, and the second value indicates that the value of the corresponding bit on the first destination register can be changed, for example, it can become the corresponding value of the first calculation result.

[0062] For example, in one example, the first value can be 0 and the second value can be 1. Or, in another example, the first value can be 1 and the second value can be 0. Embodiments of the present disclosure do not limit the specific settings of the first value and the second value of the first operation mask.

[0063] In step S300, a first instruction is executed using some bits in the first destination register. For example, the first destination operand stored in the first destination register is used to perform a calculation with the first source operand, and then the first calculation result obtained from the calculation is stored in some bits in the first destination register.

[0064] Figure 3 It is a schematic diagram of the execution process of a first instruction provided by at least one embodiment of the present disclosure. The following combines Figure 3 to detail the instruction processing method provided in the above embodiments of the present disclosure.

[0065] For example, in Figure 3 In the example shown, the first instruction is an addition instruction (for example, PADDQ instruction) in the SSE instruction set, which can perform an addition operation on the first destination operand and the first source operand, and store the first calculation result obtained by adding the two in the first destination register.

[0066] For example, the first source operand of the first instruction includes two 64-bit elements b1 and b2. The elements b1 and b2 can be obtained from another architecture register or from the internal memory. Embodiments of the present disclosure do not limit this. For example, the first destination operand of the first instruction includes two 64-bit elements a1 and a2, which are correspondingly stored in the lower bits of the first destination register (i.e., some bits in the present disclosure).

[0067] Figure 3The black box in [it] represents the first destination register. For example, the first destination register can be a 256-bit YMM register or a 512-bit ZMM register. Before executing the first instruction, the first destination register stores the values of 4 elements, which are a1, a2, a3, and a4 respectively. Among them, element a1 and element a2 are stored in the lower bits of the first destination register (i.e., the partial bits in this disclosure), and element a3 and element a4 are stored in the higher bits of the first destination register (i.e., the target bits in this disclosure).

[0068] In step S100, after obtaining the first instruction in the first instruction set, the first instruction can be parsed to determine the first destination register corresponding to the first destination operand of the first instruction. Then, it is judged whether the value on the target bits other than the bits corresponding to the first destination operand in the first destination register is 0. If it is 0, the first instruction can be directly executed. If it is not 0, a first operation mask is provided for the first instruction and then the first instruction is executed.

[0069] For example, as Figure 3 shown, in step S200, in response to the values of the target bits in the first destination register being the values of elements a3 and a4 and not being 0, it is determined to provide a first operation mask for the first instruction. Each bit of the first operation mask corresponds one-to-one with each bit of the first destination register. The value of the mask bit corresponding to the lower bits of the first destination register in the first operation mask is 1, indicating that the elements on the lower bits of the first destination register can be operated; the value of the mask bit corresponding to the higher bits of the first destination register in the first operation mask is 0, indicating that the elements on the higher bits of the first destination register remain unchanged.

[0070] Then, in step S300, since the first operation mask indicates that the elements on the partial bits (lower bits) in the first destination register can be operated, element a1 and element a2 can be added to element b1 and element b2 respectively to obtain the first calculation result, that is, element c1 and element c2, where c1 = a1 + b1 and c2 = a2 + b2. The calculated element c1 and element c2 can be stored in the lower bits of the first destination register. Since the first operation mask indicates that the elements on the target bits (higher bits) in the first destination register remain unchanged, the elements on the higher bits of the first destination register still remain as element a3 and element a4. Thus, when the first instruction is executed, the high-order data in the first destination register remains unchanged is achieved.

[0071] Through the instruction processing method provided by the above embodiments, it is possible to avoid the value of the target bits in the first destination register being overwritten or damaged by the calculation result of the first instruction, and there is no need to pour out the data on the target bits in the first destination register to the internal memory, so the storage space of the internal memory will not be occupied.

[0072] In at least one example of the embodiments of the present disclosure, the above instruction processing method further includes: in response to the value of the target bit in the first destination register being zero, directly using a partial bit in the first destination register to execute the first instruction.

[0073] That is, when all the target bits in the first destination register are 0, a first operation mask may not be provided for the first instruction, and a partial bit in the first destination register may be directly used to execute the first instruction, thereby simplifying the logic design and reducing power consumption.

[0074] Figure 4 FIG. is a schematic diagram of an execution process of another first instruction provided by at least one embodiment of the present disclosure. As Figure 4 shown, the high bits of the first destination register are all 0. In this case, after obtaining the first instruction of the first instruction set, the addition calculation may be directly performed on the first source operand and the first destination operand of the first instruction, that is, adding element a1 and element b1 to obtain element c1, adding element a2 and element b2 to obtain element c2, and storing element c1 and element c2 in the low bits of the first destination register.

[0075] For example, in one example of the embodiments of the present disclosure, the first instruction set may be a Streaming SIMD Extensions (SSE) instruction set, and the first destination register may be a YMM register or a ZMM register.

[0076] For example, in another example of the embodiments of the present disclosure, the first instruction set may be any one of a Streaming SIMD Extensions (SSE) instruction set, an Advanced Vector Extensions (AVX) instruction set, and an Advanced Vector Extensions 2 (AVX2) instruction set, and the first destination register may be a ZMM register.

[0077] For example, in at least one embodiment of the present disclosure, on a processor that supports the AVX512 instruction set, the operation mask providing unit of the AVX512 instruction set may be borrowed to provide a first operation mask for the first instruction, so that there is no need to additionally design a dedicated hardware unit or module for providing the first operation mask.

[0078] For example, in at least one example of the embodiments of the present disclosure, providing a first operation mask for the first instruction, a specific example thereof may include: using the operation mask providing unit configured to support the execution of instructions in the second instruction set to provide a first operation mask for the first instruction, where at least one instruction in the second instruction set is configured to be capable of processing more bits including partial bits in the first destination register.

[0079] For example, in at least one example of the embodiments of the present disclosure, the first instruction set may be any one of the Streaming SIMD Extensions (SSE) instruction set, the Advanced Vector Extensions (AVX) instruction set, and the Advanced Vector Extensions 2 (AVX2) instruction set, and the second instruction set is the Advanced Vector Extensions 512 (AVX512) instruction set.

[0080] For example, at least one instruction in the second instruction set may be a data transfer instruction or a data processing instruction in the AVX512 instruction set, etc., and corresponding operations can be performed on elements in more bits including partial bits in the first destination register.

[0081] The AVX512 instruction set is a vector processing instruction set that introduces an operation mask for conditional operations. Such conditional operations can reduce unnecessary calculations and improve operation efficiency. The operation mask providing unit is introduced along with the AVX512 instruction set architecture, and this operation mask providing unit can provide an operation mask value for the instructions in the second instruction set. For example, when the mask value is 1, the element corresponding to this mask value can be operated on, and when the mask value is 0, the value of the element corresponding to this mask value remains unchanged or is set to 0.

[0082] For example, in at least one embodiment of the present disclosure, the operation mask providing unit includes a mask register, and the mask register is configured to provide a value of a second operation mask to a second instruction in the second instruction set, and the value of the second operation mask is used to execute the second instruction.

[0083] For example, at least one instruction in the second instruction set includes a second instruction, and the second instruction can be any type of instruction that can operate on the first destination register. For example, in one example, the second instruction is an addition instruction (e.g., VPADDQ instruction) in the AVX512 instruction set, and the second instruction can perform an addition calculation on a second source operand and a second destination operand, and store the second calculation result of the two in the first destination register.

[0084] For example, the second destination operand may include one or more elements stored in the first destination register. Correspondingly, the operation mask providing unit may include one or more mask registers (e.g., k registers), and each mask register can provide a value of a second operation mask for the corresponding element to indicate whether the value of the element participates in the operation when the second instruction is executed. Or rather, the second operation mask can indicate whether the value of the element remains unchanged (or is set to 0) or is used to generate the second calculation result when the second instruction is executed.

[0085] For example, in at least one embodiment of the present disclosure, the bit width of each mask register is the same as the bit width of each element. For example, in at least one embodiment of the present disclosure, when the second destination operand of the second instruction corresponds to the first destination register, the instruction processing method further includes: in response to the i-th bit of the second operation mask being a first value, keeping the value corresponding to the i-th bit of the second operation mask stored in the first destination register unchanged or clearing it; or, in response to the i-th bit of the second operation mask being a second value, storing the calculation result of the i-th element in the second instruction into the first destination register, where i is less than or equal to the bit width of the first destination register.

[0086] For example, in one instance, the first value may be 0 and the second value may be 1. Alternatively, in another example, the first value may be 1 and the second value may be 0. Embodiments of the present disclosure do not limit the specific settings of the first value and the second value of the second operation mask.

[0087] For example, the second instruction may be an instruction that is executed before or after the first instruction and involves operating on the elements stored in the first destination register. For example, the second destination operand in the second instruction also corresponds to the first destination register.

[0088] Figure 5 It is a schematic diagram of the execution process of a second instruction provided for at least one embodiment of the present disclosure. For example, the second instruction is executed after the first instruction.

[0089] The following combines Figure 3 and Figure 5 to detail the process of executing the second instruction after switching from the first instruction of the first instruction set to the second instruction of the second instruction set.

[0090] For example, in Figure 3 the example shown, after executing the first instruction, there are 4 elements stored in the first destination register, namely c1, c2, a3, and a4.

[0091] As Figure 5 shown, after executing the first instruction in the SSE instruction set, it is necessary to continue executing the second instruction in the AVX512 instruction set. The second destination operand of the second instruction corresponds to all bits in the first destination register, that is, the second instruction can process all bits in the first destination register. For example, the second destination operand of the second instruction includes 4 64-bit elements c1, c2, a3, and a4, which are correspondingly stored in the first destination register.

[0092] For example, the second source operand of the second instruction includes four 64-bit elements d1, d2, d3, and d2. These four elements d1, d2, d3, and d2 can be obtained from another architectural register or from internal memory, and the embodiments of the present disclosure do not limit this.

[0093] Figure 5 The black box in represents the first destination register (e.g., ZMM register), and the white box represents the operation mask providing unit. The operation mask providing unit includes, for example, four mask registers, and these four mask registers respectively provide corresponding mask values for the four elements c1, c2, a3, and a4 of the second destination operand of the second instruction. For example, the second operation mask is 0, 1, 0, 1. According to this second operation mask, the mask values corresponding to elements c2 and a4 are 0, so elements c2 and a4 can participate in the operation. The mask values corresponding to elements c1 and a3 are 1, so the values of elements c1 and a3 are set to 0 (in another example, the values of elements c1 and a3 can also be set to remain unchanged).

[0094] Therefore, after the second instruction is executed, the second calculation result is obtained: 0, f2, 0, f4, where f2 = c2 + d2, f4 = a4 + d4, and elements f2 and f4 are stored in the corresponding bits of the first destination register.

[0095] From the above process, it can be seen that using the instruction processing method provided by the embodiments of the present disclosure can simplify the processing flow, save processing time, does not need to rely on the spill-load mechanism, enables the second instruction executed after the first instruction to directly operate on the data stored in the first destination register, without accessing data from internal memory, thereby effectively reducing the memory access latency, improving the execution efficiency of instructions, reducing power consumption, and enhancing system performance.

[0096] For example, in an example of the embodiments of the present disclosure, on a processor supporting, for example, the AVX512 instruction set, each mask register in the operation mask providing unit of the AVX512 instruction set can be used to provide the second operation mask for the second instruction, and the first operation mask for the first instruction can be provided by assigning values to the interface of the operation mask providing unit. That is, it is not necessary to configure a real mask register for the first operation mask, and only need to set the microarchitecture-visible signal (such as the first operation mask shown by the dashed box in Figure 3 ), thereby being able to make full use of the existing hardware units without additional logic. Figure 3

[0097] In addition, the instruction processing method provided by the above embodiments of the present disclosure can be applied not only to single-threaded or multi-threaded scenarios, but also to scenarios where multiple instruction sets are executed in a mixed manner, with a wide range of applications and strong operability.

[0098] At least one embodiment of the present disclosure also provides a processor. Figure 6 It is a schematic block diagram of a processor provided by at least one embodiment of the present disclosure. As Figure 6 shown, the processor 600 includes a receiving unit 610, an operation mask providing unit 620, and an execution unit 630.

[0099] For example, the receiving unit 610 is configured to obtain a first instruction in a first instruction set, where the first instruction is configured to process only a part of the bits in a first destination register, and a first destination operand of the first instruction corresponds to a part of the bits in the first destination register.

[0100] For example, the operation mask providing unit 620 is configured to provide a first operation mask for the first instruction in response to the value of the target bits other than the part of the bits in the first destination register not being zero.

[0101] For example, the execution unit 630 is configured to execute the first instruction using a part of the bits in the first destination register, and in response to the first operation mask being provided to the first instruction, keep the value of the target bits in the first destination register unchanged.

[0102] For example, the processor can be a general-purpose processor or a special-purpose processor. For example, the processor can be a central processing unit (CPU), a digital signal processor (DSP), a microcontroller unit (MCU), a tensor processing unit (TPU), a neural network processor (NPU), a data processing unit (DPU), an AI accelerator, a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other devices that support out-of-order execution. For example, according to the instruction set used, such as the X86 instruction set, the ARM instruction set, the RSIC-V instruction set, etc., the processor can be based on the X86 microarchitecture, the ARM microarchitecture, the RISC-V microarchitecture, etc.

[0103] For example, in the processor 600 provided by at least one embodiment of the present disclosure, the execution unit 630 is further configured to: in response to the value of the target bits in the first destination register being zero, directly execute the first instruction using a part of the bits in the first destination register.

[0104] For example, in the processor 600 provided by at least one embodiment of the present disclosure, the operation mask providing unit 620 is further configured to support running instructions in a second instruction set, where at least one instruction in the second instruction set is configured to be able to process more bits including a part of the bits in the first destination register.

[0105] For example, in the processor 600 provided by at least one embodiment of the present disclosure, the operation mask providing unit 620 includes one or more mask registers 621. The mask register 621 is configured to provide the value of the second operation mask to the second instruction in the second instruction set, and the value of the second operation mask is used to execute the second instruction.

[0106] For example, in the processor 600 provided by at least one embodiment of the present disclosure, the receiving unit 610 is further configured to obtain the second instruction in the second instruction set, the second instruction is executed before or after the first instruction, and the second destination operand of the second instruction corresponds to the first destination register. The execution unit 630 is further configured to: in response to the i-th bit of the second operation mask being the first value, keep the value corresponding to the i-th bit of the second operation mask stored in the first destination register unchanged or cleared; or, in response to the i-th bit of the second operation mask being the second value, store the calculation result of the i-th element in the second instruction into the first destination register, where i is less than or equal to the number of elements in the second instruction.

[0107] For example, in the processor 600 provided by at least one embodiment of the present disclosure, the first instruction set includes a data stream single instruction multiple data extension instruction set, an advanced vector extension instruction set, and an advanced vector extension 2 instruction set, and the second instruction set includes an advanced vector extension 512 instruction set.

[0108] For example, in the processor 600 provided by at least one embodiment of the present disclosure, when the first register is 256 bits, the partial bits include bits [127:0], and the target bits include bits [255:128]; or, when the first register is 512 bits, the partial bits include bits [127:0], and the target bits include bits [511:128].

[0109] The above receiving unit 610, operation mask providing unit 620, and execution unit 630 can be implemented by hardware or firmware, etc., and can be respectively implemented as a receiving circuit, an operation mask providing circuit, and an execution circuit, etc.

[0110] The processor provided by the above embodiments of the present disclosure can provide a first operation mask for the first instruction, so that when the first instruction is executed, the value of the target bits in the first destination register remains unchanged, which can avoid the value of the target bits in the first destination register being overwritten or damaged by the calculation result of the first instruction, and does not occupy the storage space of the internal memory. The subsequent executed instructions can directly operate on the data stored in the first destination register without accessing the data from the internal memory, thereby reducing the storage access latency, improving the execution efficiency of the instructions, reducing power consumption, and enhancing the system performance.

[0111] Some embodiments of the present disclosure also provide an electronic device. Figure 7Schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure.

[0112] As Figure 7 shown, the electronic device includes a processor 701 of any of the above embodiments. For example, the electronic device can execute the instruction processing method of any of the above embodiments through the processor 701.

[0113] The electronic device provided by the embodiment of the present disclosure and the instruction processing method provided by the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0114] Figure 8 Structural schematic diagram of an electronic device provided by at least one embodiment of the present disclosure. The electronic device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The illustrated electronic device 800 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0115] For example, as Figure 8 shown, in some examples, the electronic device 800 includes a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, and the processing device 801 may include the processor of any of the above embodiments, and it can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the storage device 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the computer system are also stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.

[0116] For example, the following components may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, such as, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; a communication device 809 which may also include, for example, a network interface card such as a LAN card, a modem, etc. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data and perform communication processing via a network such as the Internet. The driver 810 is also connected to the I / O interface 805 as needed. A removable storage medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the driver 810 as needed so that a computer program read from it can be installed into the storage device 808 as needed. Although Figure 8 the electronic device 800 including various devices is shown, it should be understood that it is not required to implement or include all the shown devices, and more or fewer devices may alternatively be implemented or included.

[0117] For example, the electronic device 800 may further include a peripheral interface (not shown in the figure), etc. The peripheral interface may be various types of interfaces, such as a USB interface, a Lightning interface, etc. The communication device 809 may communicate with a network and other devices via wireless communication. The network may be, for example, the Internet, an intranet, and / or a wireless network such as a cellular phone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). The wireless communication may use any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), WiMAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.

[0118] For example, the electronic device 800 may be any device such as a mobile phone, a tablet computer, a laptop computer, an e-book, a game console, a television, a digital photo frame, a navigator, etc., or may be any combination of a data processing device and hardware. The embodiments of the present disclosure are not limited thereto.

[0119] Although the present disclosure has been described in detail above with general descriptions and specific embodiments, on the basis of the embodiments of the present disclosure, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present disclosure fall within the scope of protection required by the present disclosure.

[0120] For the present disclosure, in addition to the above exemplary content, the following points need to be noted:

[0121] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0122] (2) For clarity, in the drawings used to describe the embodiments of the present disclosure, the thickness of the layers or regions is enlarged or reduced, that is, these drawings are not drawn to actual scale.

[0123] (3) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0124] As described above, only the specific embodiments of the present disclosure are provided, but the scope of protection of the present disclosure is not limited thereto. The scope of protection of the present disclosure shall be subject to the scope of protection of the claims.

Claims

1. A command processing method, comprising: Obtaining a first instruction in a first instruction set, wherein the first instruction is configured to process only some bits in a first destination register, and a first destination operand of the first instruction corresponds to the some bits in the first destination register; In response to the value of the target bit other than the partial bits in the first destination register being not zero, providing a first operation mask for the first instruction, wherein the first operation mask is used to keep the value of the target bit in the first destination register unchanged when executing the first instruction; and The first instruction is executed using the portion of bits in the first destination register.

2. The instruction processing method according to claim 1, further comprising: In response to the value of the target bit in the first destination register being zero, the first instruction is directly executed using the partial bits in the first destination register.

3. The instruction processing method according to claim 1, wherein: The providing a first operation mask for the first instruction comprises: The first operation mask is provided for the first instruction using an operation mask providing unit configured to support execution of instructions in a second instruction set, wherein at least one instruction in the second instruction set is configured to be able to process more bits including the partial bits in the first destination register.

4. The instruction processing method according to claim 3, wherein: The operation mask providing unit includes a mask register, and the mask register is configured to provide a value of a second operation mask to a second instruction in the second instruction set, and the value of the second operation mask is used to execute the second instruction.

5. The instruction processing method according to claim 4, wherein: When the second destination operand of the second instruction corresponds to the first destination register, the instruction processing method further includes: In response to the i-th bit of the second operation mask being a first value, the value of the i-th bit corresponding to the second operation mask stored in the first destination register remains unchanged or is cleared; or, In response to the i-th bit of the second operation mask being a second value, storing the calculation result of the i-th element in the second instruction into the first destination register, Wherein, i is less than or equal to the bit width of the first destination register.

6. The instruction processing method according to any one of claims 3 to 5, wherein: The first instruction set includes a streaming SIMD extension instruction set, an advanced vector extension instruction set and an advanced vector extension 2 instruction set, and the second instruction set includes an advanced vector extension 512 instruction set.

7. The instruction processing method according to any one of claims 1 to 5, wherein: When the first destination register has 256 bits, the partial bits include bits [127: 0], and the target bits include bits [255: 128]; or, When the first destination register is 512 bits, the partial bits include bits [127: 0], and the target bits include bits [511: 128].

8. The instruction processing method according to any one of claims 1 to 5, wherein: The first instruction is used to perform a calculation operation on a vector or a scalar.

9. A processor, comprising: a receiving unit configured to obtain a first instruction in a first instruction set, wherein the first instruction is configured to process only some bits in a first destination register, and a first destination operand of the first instruction corresponds to the some bits in the first destination register; an operation mask providing unit configured to provide a first operation mask for the first instruction in response to the value of the target bits other than the partial bits in the first destination register being not zero; and The execution unit is configured to execute the first instruction using the partial bits in the first destination register, and in response to the first operation mask being provided to the first instruction, keep the value of the target bits in the first destination register unchanged.

10. The processor of claim 9, wherein: The execution unit is further configured to: In response to the value of the target bit in the first destination register being zero, the first instruction is directly executed using the partial bits in the first destination register.

11. The processor of claim 9, wherein: The operation mask providing unit is further configured to support the execution of instructions in a second instruction set, wherein at least one instruction in the second instruction set is configured to be able to process more bits including the partial bits in the first destination register.

12. The processor of claim 11, wherein: The operation mask providing unit includes a mask register, and the mask register is configured to provide a value of a second operation mask to a second instruction in the second instruction set, and the value of the second operation mask is used to execute the second instruction.

13. The processor of claim 12, wherein: The receiving unit is further configured to obtain the second instruction in the second instruction set, the second instruction is executed before or after the first instruction, and a second destination operand of the second instruction corresponds to the first destination register, The execution unit is further configured to: In response to the i-th bit of the second operation mask being a first value, keeping the value of the i-th bit corresponding to the second operation mask stored in the first destination register unchanged or cleared; or, In response to the i-th bit of the second operation mask being a second value, storing the calculation result of the i-th element in the second instruction into the first destination register, Wherein, i is less than or equal to the bit width of the first destination register.

14. The processor according to any one of claims 11 to 13, wherein: The first instruction set includes a streaming SIMD extension instruction set, an advanced vector extension instruction set and an advanced vector extension 2 instruction set, and the second instruction set includes an advanced vector extension 512 instruction set.

15. The processor according to any one of claims 9 to 13, wherein: When the first register is 256 bits, the partial bits include bits [127: 0], and the target bits include bits [255: 128]; or, when the first register is 512 bits, the partial bits include bits [127: 0], and the target bits include bits [511: 128].

16. An electronic device comprising the processor according to any one of claims 9 to 15.